evaluation
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
The ChatGPT maker says its upcoming Astra model may have reached "critical" cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards. OpenAI announced Tuesday that it has halted "a significant number" of training workloads and evaluations for its forthcoming frontier artificial intelligence model--codenamed Astra--while it implements new procedures meant to address cybersecurity risks. The ChatGPT maker says it is introducing a number of new monitoring, security, and alignment requirements to better address the increasingly advanced hacking abilities of its frontier AI models . "We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that's how long people are unable to proceed with their workloads," Amelia Glaese, OpenAI's vice president of research and safety, said in a briefing with reporters Tuesday.
AI's recursive self-improvement might not come so quickly after all
AI's recursive self-improvement might not come so quickly after all AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems. The AI industry's boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. Forecasts of explosive AI progress predict that what researchers call recursive self-improvement is on the horizon. But a new study suggests that it might take a while for us to get there. The researchers behind it found that AI agents are not yet capable of conducting open-ended AI research--free-form investigations that have no clear-cut answers and require judgment and taste, which may be integral to building self-improving AI.
Royal Statistical Society AI task force says: AI regulation needs statistics
Anne Fehres and Luke Conroy AI4Media Humans Do The Heavy Data Lifting Licenced by CC-BY 4.0 The Royal Statistical Society's AI Task Force has issued a critical mandate via a new paper, AI Regulation Needs Statistics, which demands that statistical principles actively shape global AI governance. The publication escalates the core argument of their foundational work, AI is Statistics . This earlier paper argued that AI is fundamentally statistical, meaning effective and ethical deployment is impossible without statistical literacy. You can watch our expert panel discuss the topic here . A major focus of that work was around the challenges of evaluating AIs, given that they are dynamic systems that continue to evolve once they have been deployed in the real world.
OpenAI slows down Astra development due to cybersecurity concerns
Shortly after a major cybersecurity incident where OpenAI's models hacked into an open source machine learning platform called Hugging Face, the company announced that it's bolstering safeguards and security controls for its latest AI model. In a post on its website, OpenAI said internal evaluations of its upcoming model, called Astra, showed "significant advancements in agentic coding and cybersecurity," resulting in OpenAI not being able to "rule out critical cyber capabilities." According to OpenAI, it can't declare with certainty that the unreleased Astra model would be designated as a "Critical capability level." As detailed in its own Preparedness Framework, OpenAI said the Critical designation means that a model "can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention." It could also be able to "devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal."
Chinese AI model Moonshot Kimi K3 also escaped its testing environment
It wasn't too long ago when the idea of an AI model or agent escaping their confines and breaking into websites on their own felt alarming. Now, it has become a pretty common story. Kimi K3, one of most powerful AI models developed by a Chinese company, also escaped its testing environment. According to US cybersecurity startup Frontier, Kimi K3 broke out of a sandbox from the UK government's AI Security Institute (AISI) while its defensive cybersecurity skills were being evaluated. Moonshot launched Kimi K3 in July and made it available for free shortly thereafter.
First OpenAI, now Meta - why do AI hacks keep happening?
First OpenAI, now Meta - why do AI hacks keep happening? Over the last fortnight, reports of AI models going beyond their expected bounds - be that technically or morally - have been seemingly unavoidable. What started with a trickle - ChatGPT-maker OpenAI admitting their AI had hacked the site Hugging Face - has turned into a flood of groups revealing they had discovered instances of AI going out of control. Claude-maker Anthropic, Meta and the UK's AI Security Institute (AISI) have now each reported incidents which seem to paint a worrying picture of a world in which tech going rogue is the norm. In reality, each case offers a window into the risks posed by increasingly capable AI agents - and the importance of testing their limits before they are released to the world.
François Pachet on music generation with AI
Dr François Pachet is an AI researcher and musician, and one of the most influential figures in AI and music. His innovative contributions have defined the field over the past decades through creative systems such as the Continuator, and Flow Machines, among others. After leading the Spotify Creator Technology Research Lab and the Sony Computer Science Lab, he went on to create his own companies: Imagine All The People and Ynosound. In the context of IJCAI2025, he spoke about what deep learning changed, and what still remains wide open. He explains why tools like Suno and Udio--ChatGPT-like platforms for music generation--can produce astonishing results that still feel unsatisfying; why the next step for music generation requires combining sampling with search; and why the most important problems in artistic domains are, by nature, ill-defined--because there is no loss function to determine what is "good". Above all, he defends the importance of researcher autonomy: work on the questions that genuinely fascinate you, even when they fall outside prevailing trends--perhaps especially then. Thank you for joining me for this interview. Could you begin by telling us when was your first IJCAI and a memory related to it? I think the first IJCAI I attended was in Montreal in '95. I was there for a couple of workshops, one of them was about music and AI, and the other one I think was on advisor systems, something like that. And I remember there was a French colleague who was there also, at the time he was doing his PhD. And there was this researcher called Herbert Simon, who is a Nobel Prize pioneer of AI. I remember having chatted a little bit with this French guy who was very bold, and he just went up to Simon, said "Hey," and he started a conversation with him. And I was very impressed by the fact that you could meet those kinds of guys informally in a corridor or something at this conference.
Healthcare benchmarks are only as good as their assumptions
In healthcare settings where patients use LLMs as a medical assistant, LLM performance differs between evaluation and deployment. Closing the gap requires making assumptions explicit, testing which assumptions hold, and updating evaluation protocols accordingly. Healthcare LLM benchmarks are one of the main paradigms by which LLMs are evaluated prior to clinical settings. Benchmarks provide a stable goalpost that allow researchers to iterate quickly and measure progress consistently. However, in high-stakes domains like healthcare, that same abstraction becomes a liability.
AI models shock UK testers by using fake identities to try to trick developers
AISI said the rogue behaviour was carried out by agents powered by two models - Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. AISI said the rogue behaviour was carried out by agents powered by two models - Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. Explainer: Should we be alarmed at AI models going rogue in tests? Advanced artificial intelligence models have stunned the UK's AI Security Institute (AISI) by carrying out a hacking campaign against real people during a cybersecurity test. The institute said the incident was unprecedented and involved sending targeted emails to software developers in an attempt to pass a cyber challenge.