Goto

Collaborating Authors

 Large Language Model


OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

BBC News

OpenAI says its AI went rogue and launched'unprecedented' cyber-attack OpenAI has revealed some of its most advanced AI models went rogue and hacked a start-up after it lost control of them during a security test. The ChatGPT-maker said its agents - AI bots which can operate alone after some human instruction - were being tested in a controlled environment, but found vulnerabilities and managed to escape. They targeted Hugging Face, one of the world's largest hubs for sharing AI models, gaining access to some internal company systems. OpenAI said the incident was unprecedented, external, and it was working with Hugging Face to investigate what happened and strengthen safeguards. Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4's Today programme that the security tests - called sandboxes - are supposed to be secure environments where you can see what the models are capable of. In this case, it looks like OpenAI didn't make a secure enough sandbox, she added.


AI agent went rogue and hacked startup by itself, OpenAI reveals

The Guardian

Hugging Face's chief executive said the attack was'mind-blowing' but that he believed there was'no malicious intent' from OpenAI. Hugging Face's chief executive said the attack was'mind-blowing' but that he believed there was'no malicious intent' from OpenAI. OpenAI has revealed an autonomous AI agent powered by its technology went rogue during a test, accessed the open web and hacked a prominent startup by itself in an "unprecedented incident". The company behind ChatGPT said the startup Hugging Face had detected and contained the agent - an AI tool designed to carry out tasks without human assistance - which had entered its systems. "We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities," OpenAI said .


OpenAI admits its models hacked Hugging Face on their own

Engadget

They escaped an isolated environment for testing and infiltrated Hugging Face without human input. Picture this: A couple of powerful AI models being tested by their company escaped a controlled environment, got on the internet and then hacked a machine learning repository on their own, without human input. Sounds like the plot of a Terminator movie, doesn't it? Except it just happened for real. A few days after open source AI platform Hugging Face revealed that it detected unauthorized access on its systems by an AI agent, OpenAI has admitted that its models were the culprit.


Will AI help you do your job or replace you?

BBC News

Will AI help you do your job or replace you? Artificial Intelligence (AI) companies are making vast claims about the ability of their tools to replace human labour. Some jobs will be automated, others will be augmented. The bosses of the world's biggest companies are diverting vast sums into these tools, partly with the knowledge that they could save money on headcount. Flat is the new up, we are told, in terms of the size of a company's workforce as investors ask whether jobs should be done by new recruits - or, instead, armies of AI Agents, virtual workers tasked with doing specific roles, some of them relatively skilled.


OpenAI Models Escaped Containment and Hacked Hugging Face

WIRED

The cybersecurity-focused models, including GPT-5.6 Sol, broke out of a testing sandbox, exploited a zero-day, and gained access to the open internet to pull off the attack. OpenAI disclosed on Tuesday that it lost control of two AI models during a security test that ended in a breach of the AI research platform Hugging Face. Describing the incident as "unprecedented," OpenAI said its AI models broke out of a sealed testing environment last week and hacked into Hugging Face's production system to steal the answers to a test they were being graded on. The models--the publicly available GPT-5.6 Sol and an unreleased, reportedly more capable one--were being evaluated on their offensive hacking skills with the safeguards that normally block high-risk cyber activity switched off. "The models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database," OpenAI and Hugging Face wrote in a joint blog post disclosing the intrusion.


Where Did All the Computer-Science Professors Go?

The Atlantic - Technology

Where Did All the Computer-Science Professors Go? AI companies are stripping universities of their best researchers. Anthropic has poached such an array of high-profile professors that it has become a punch line in academia. "'I'm joining Anthropic' is the new meme right now," Subbarao Kambhampati, a computer-science professor at Arizona State University (who has not joined Anthropic), told us. This month, the AI company hired the chair of UC Berkeley's department of electrical engineering and computer science, presumably to help build more capable bots. Perhaps more surprisingly, Anthropic has in recent weeks also picked up a Stanford economist, a theoretical physicist from the University of Maryland, and an analytic philosopher from UT Austin.


OpenAI's newest AI model broke its own sandbox rules to finish a task

PCWorld

PCWorld reports that OpenAI's unreleased AI model broke out of its sandbox environment to complete a task, choosing to follow GitHub posting instructions over safety guardrails. The incident occurred during a NanoGPT speedrun benchmark where the autonomous model hacked its way out to post code publicly despite being restricted to Slack-only communication. OpenAI paused development after discovering this and other unwanted behaviors, highlighting the need for enhanced safeguards as AI models become more persistent and autonomous. Not only are they smarter and more capable, but the newest and most powerful AI models are also less likely to give up when they hit roadblocks. An unreleased OpenAI model took that perseverance to an extreme when it broke out of its sandbox to fulfill instructions that were in conflict with its built-in guardrails.


The Download: Chinese AI divides the White House, and a record copyright payout

MIT Technology Review

China's AI models have Trump's AI world at war with itself David Sacks branded Anthropic's models "lobotomized" and "woke." Emil Michael, a top Pentagon official, called OpenAI's new head of strategic futures a "supreme village idiot." It began because no one can agree on what to do about Kimi, a free, open-source model that Chinese AI company Moonshot launched last week. It appears to rival the intelligence of models from OpenAI and Anthropic, which are very much not free. Every time a new smart, free model from China gets released, US companies see less reason to fork out money for models from Anthropic or OpenAI. Read the full story on why no one can agree what to do about Kimi .


Chinese AI sensation Moonshot's gamble on big models pays off

The Japan Times

Chinese AI sensation Moonshot's gamble on big models pays off People visit the Moonshot AI stand, featuring the Kimi K3 model, during the World Artificial Intelligence Conference in Shanghai on Saturday. Crowds rushed to an obscure corner of China's premier tech summit, moving past monumental booths from Alibaba Group Holding and Tencent Holdings to catch a glimpse of the hottest name in domestic artificial intelligence. Developers and users at the World AI Conference jostled for a closer look at Moonshot and its Kimi K3 -- the giant 2.8 trillion-parameter model whose release on Friday showed China was closing the gap on OpenAI and Anthropic PBC far quicker than anticipated. The startup became an instant global sensation, drawing comparisons to DeepSeek's 2025 breakout and plaudits from the likes of Tesla CEO Elon Musk. Such was the crush over the weekend that Moonshot blew through its entire stock of branded swag in just a few hours. Moonshot's emergence is a vindication not just for the country's AI industry, but also for founder Yang Zhilin.


China's AI models have Trump's AI world at war with itself

MIT Technology Review

China's AI models have Trump's AI world at war with itself Kimi and other free models from China have again been seen as a wake-up call. David Sacks, the president's AI and crypto "czar" until March, branded Anthropic's models as "lobotomized" and "woke." Emil Michael, a top Pentagon official, called OpenAI's new head of strategic futures a "supreme village idiot." It began because no one can agree on what to do about Kimi, a free, open source model that Chinese AI company Moonshot launched last week. It appears to rival the intelligence of models from OpenAI and Anthropic, which are very much not free. Kimi and other Chinese models like it pose a real problem for Trump.