Goto

Collaborating Authors

 evaluation


OpenAI responds after report exposed another incident in which its AI agents went rogue

Engadget

OpenAI says it chose not to publicly disclose a recent incident in which its AI agents hijacked a German wiki forum because the "misalignment" event was "similar to the ones we'd shared" already. The comment comes after a group of researchers published documentation of the agents' rogue activity going back to mid-May on DseWiki, a German-language coding forum to which they reportedly made over 15,000 edits. Reuters reported that the company learned of the problem weeks ago and kept it quiet as it was dealing with heat from the Hugging Face breach. OpenAI addressed the "wiki incident" in an X post on Saturday, writing that "it's past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models." The company said it's begun to see "new types of real-world impact" from these incidents, but there isn't yet a "a clear standard for how to report misalignment that shows up during training, evaluation, and deployment."


OpenAI officially launches GPT-6 Astra, a boundary-breaking new model

Mashable

Look Up Trending Now Good Connection: Uplifting stories for a digital age Creator Playbook Say More Mashable's Best: E-readers, robovacs, laptops, earbuds, smart home and more Switch Off Mashable Voices Mashable Selects Safety Net Versus Gift Ideas For Everyone On Your List All Series The newest AI model from OpenAI is the first to reach the company's highest threat level. Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals. Tim has over 15 years of experience as a journalist and editor, and he has particular experience covering and testing consumer technology, smart home gadgets, and men's grooming and style products. Previously, he was the Managing Editor and then Site Director of SPY.com, a men's product review and lifestyle website. As a writer for GQ, he covered everything from bull-riding competitions to the best Legos for adults, and he's also contributed to publications such as The Daily Beast, Gear Patrol, and The Awl.


OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

WIRED

The ChatGPT maker says its upcoming Astra model may have reached "critical" cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards. OpenAI announced Tuesday that it has halted "a significant number" of training workloads and evaluations for its forthcoming frontier artificial intelligence model--codenamed Astra--while it implements new procedures meant to address cybersecurity risks. The ChatGPT maker says it is introducing a number of new monitoring, security, and alignment requirements to better address the increasingly advanced hacking abilities of its frontier AI models . "We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that's how long people are unable to proceed with their workloads," Amelia Glaese, OpenAI's vice president of research and safety, said in a briefing with reporters Tuesday.


AI's recursive self-improvement might not come so quickly after all

MIT Technology Review

AI's recursive self-improvement might not come so quickly after all AI agents are not yet creative enough to carry out genuinely innovative open-ended AI research, it seems. The AI industry's boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. Forecasts of explosive AI progress predict that what researchers call recursive self-improvement is on the horizon. But a new study suggests that it might take a while for us to get there. The researchers behind it found that AI agents are not yet capable of conducting open-ended AI research--free-form investigations that have no clear-cut answers and require judgment and taste, which may be integral to building self-improving AI.


Royal Statistical Society AI task force says: AI regulation needs statistics

AIHub

Anne Fehres and Luke Conroy AI4Media Humans Do The Heavy Data Lifting Licenced by CC-BY 4.0 The Royal Statistical Society's AI Task Force has issued a critical mandate via a new paper, AI Regulation Needs Statistics, which demands that statistical principles actively shape global AI governance. The publication escalates the core argument of their foundational work, AI is Statistics . This earlier paper argued that AI is fundamentally statistical, meaning effective and ethical deployment is impossible without statistical literacy. You can watch our expert panel discuss the topic here . A major focus of that work was around the challenges of evaluating AIs, given that they are dynamic systems that continue to evolve once they have been deployed in the real world.


OpenAI slows down Astra development due to cybersecurity concerns

Engadget

Shortly after a major cybersecurity incident where OpenAI's models hacked into an open source machine learning platform called Hugging Face, the company announced that it's bolstering safeguards and security controls for its latest AI model. In a post on its website, OpenAI said internal evaluations of its upcoming model, called Astra, showed "significant advancements in agentic coding and cybersecurity," resulting in OpenAI not being able to "rule out critical cyber capabilities." According to OpenAI, it can't declare with certainty that the unreleased Astra model would be designated as a "Critical capability level." As detailed in its own Preparedness Framework, OpenAI said the Critical designation means that a model "can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention." It could also be able to "devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal."


Chinese AI model Moonshot Kimi K3 also escaped its testing environment

Engadget

It wasn't too long ago when the idea of an AI model or agent escaping their confines and breaking into websites on their own felt alarming. Now, it has become a pretty common story. Kimi K3, one of most powerful AI models developed by a Chinese company, also escaped its testing environment. According to US cybersecurity startup Frontier, Kimi K3 broke out of a sandbox from the UK government's AI Security Institute (AISI) while its defensive cybersecurity skills were being evaluated. Moonshot launched Kimi K3 in July and made it available for free shortly thereafter.


First OpenAI, now Meta - why do AI hacks keep happening?

BBC News

First OpenAI, now Meta - why do AI hacks keep happening? Over the last fortnight, reports of AI models going beyond their expected bounds - be that technically or morally - have been seemingly unavoidable. What started with a trickle - ChatGPT-maker OpenAI admitting their AI had hacked the site Hugging Face - has turned into a flood of groups revealing they had discovered instances of AI going out of control. Claude-maker Anthropic, Meta and the UK's AI Security Institute (AISI) have now each reported incidents which seem to paint a worrying picture of a world in which tech going rogue is the norm. In reality, each case offers a window into the risks posed by increasingly capable AI agents - and the importance of testing their limits before they are released to the world.


François Pachet on music generation with AI

AIHub

Dr François Pachet is an AI researcher and musician, and one of the most influential figures in AI and music. His innovative contributions have defined the field over the past decades through creative systems such as the Continuator, and Flow Machines, among others. After leading the Spotify Creator Technology Research Lab and the Sony Computer Science Lab, he went on to create his own companies: Imagine All The People and Ynosound. In the context of IJCAI2025, he spoke about what deep learning changed, and what still remains wide open. He explains why tools like Suno and Udio--ChatGPT-like platforms for music generation--can produce astonishing results that still feel unsatisfying; why the next step for music generation requires combining sampling with search; and why the most important problems in artistic domains are, by nature, ill-defined--because there is no loss function to determine what is "good". Above all, he defends the importance of researcher autonomy: work on the questions that genuinely fascinate you, even when they fall outside prevailing trends--perhaps especially then. Thank you for joining me for this interview. Could you begin by telling us when was your first IJCAI and a memory related to it? I think the first IJCAI I attended was in Montreal in '95. I was there for a couple of workshops, one of them was about music and AI, and the other one I think was on advisor systems, something like that. And I remember there was a French colleague who was there also, at the time he was doing his PhD. And there was this researcher called Herbert Simon, who is a Nobel Prize pioneer of AI. I remember having chatted a little bit with this French guy who was very bold, and he just went up to Simon, said "Hey," and he started a conversation with him. And I was very impressed by the fact that you could meet those kinds of guys informally in a corridor or something at this conference.


Healthcare benchmarks are only as good as their assumptions

AIHub

In healthcare settings where patients use LLMs as a medical assistant, LLM performance differs between evaluation and deployment. Closing the gap requires making assumptions explicit, testing which assumptions hold, and updating evaluation protocols accordingly. Healthcare LLM benchmarks are one of the main paradigms by which LLMs are evaluated prior to clinical settings. Benchmarks provide a stable goalpost that allow researchers to iterate quickly and measure progress consistently. However, in high-stakes domains like healthcare, that same abstraction becomes a liability.