Generative AI
OpenAI models joined forces months ahead of Hugging Face hack
In the wake of the breaches, OpenAI has slowed its research, and its teams have dropped everything to focus on enhancing responses to security anomalies. OpenAI said the artificial intelligence models behind an attack on Hugging Face began communicating with each other through undetected message boards, working together to break out of their testing environment as early as May. Multiple internal-only agents and AI models spent months leaving notes for each other and coalescing around the goal of accessing the internet to solve the tasks they had been given, some of which were impossible without online access, OpenAI staffers Eric Wallace and Michael Dalton said Wednesday at a cybersecurity conference. "At some point, the agents realized that maybe we could try to exploit or attack external infrastructure in order to find the answers to the test that I'm being evaluated on," Wallace said during a presentation at the Black Hat conference in Las Vegas. In a time of both misinformation and too much information, quality journalism is more crucial than ever.
AI models have been going rogue in tests – how worried should we be?
The AISI said there were 19 examples of rogue behaviour, 17 of them carried out by Anthropic's Mythos and two by OpenAI's GPT 5.6-Sol. The AISI said there were 19 examples of rogue behaviour, 17 of them carried out by Anthropic's Mythos and two by OpenAI's GPT 5.6-Sol. AI models have been going rogue in tests - how worried should we be? The UK's AI Security Institute test revealed AI models indulging in unprecedented hacking attempts Two cutting-edge AI models have targeted real people and organisations in the latest safety scare to hit the technology. The UK's AI Security Institute (AISI) said the incident was unprecedented but could become more common as the technology becomes increasingly capable. The AISI, which is owned by the UK government and tests advanced AI models, said in a blog post that two AI agents carried out unprecedented hacking attempts during a cybersecurity evaluation.
The AI hacking tests keep escaping the lab
When you purchase through links in our articles, we may earn a small commission. This time, it was third-party AI testers that spotted Claude and GPT models trying to hack real companies and organizations. Once again, the most powerful Claude and ChatGPT models have been caught going rogue, with a pair of third-party cybersecurity teams spotting attempts by the models to hack real companies and even people. The UK government-backed AI Security Institute reports that during a series of cybersecurity evaluations, Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol both took "autonomous, unsanctioned action on the live internet," including an instance where an agent attempted to upload malicious code to GitHub using a phony identity. In another incident, an OpenAI model that had mistakenly been given internet access hacked a real website during a "capture the flag" exercise, according to third-party AI evaluator Irregular.
AI models shock UK testers by using fake identities to try to trick developers
AISI said the rogue behaviour was carried out by agents powered by two models - Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. AISI said the rogue behaviour was carried out by agents powered by two models - Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. Explainer: Should we be alarmed at AI models going rogue in tests? Advanced artificial intelligence models have stunned the UK's AI Security Institute (AISI) by carrying out a hacking campaign against real people during a cybersecurity test. The institute said the incident was unprecedented and involved sending targeted emails to software developers in an attempt to pass a cyber challenge.
The Download: reward hacking explained and suspected Iranian cyberattacks
Here's why AI agents lie and cheat to reach their goals When two OpenAI models hacked into Hugging Face last month, they weren't trying to make money or commit sabotage--they were just looking for answers to a test question. According to OpenAI, the models decided to solve a cybersecurity exercise by hacking out of the environment in which OpenAI had attempted to contain them and into Hugging Face's databases, where--they reasoned--the correct answer to the problem might be stored. The incident has attracted intense attention over the past couple of weeks. It's a dramatic illustration of just how good AI models have gotten at hacking. But it's perhaps even more striking as an example of how and why AI systems lie and cheat. Read our story explaining why AI engages in this sort of behavior--known as "reward hacking."
Here's why AI agents lie and cheat to reach their goals
When two OpenAI models hacked into the website Hugging Face in July, they weren't trying to make money or commit sabotage--they were just looking for answers to a test question. According to a postmortem from OpenAI, the models, which had been stripped of their typical security features for testing, decided to solve a cybersecurity exercise by hacking out of the isolated environment in which OpenAI had attempted to contain them and into Hugging Face's databases, where--they reasoned--the correct answer to the problem might be stored. The Hugging Face incident has attracted intense attention over the past couple of weeks. It's a dramatic illustration of just how good AI models have gotten at hacking: In order to get into Hugging Face's databases, the models had to string together several previously undiscovered cybersecurity exploits. But it's perhaps even more striking as an example of how and why AI systems lie and cheat. And as models get increasingly powerful, the consequences could get far more severe.
Something Weird Is Happening in Math
Why one of the world's best mathematicians is joining OpenAI One of the winners of this year's Fields Medal is headed to OpenAI. Last Thursday, Jacob Tsimerman was one of four mathematicians awarded the prestigious honor, which is sometimes called the Nobel Prize of mathematics. The same day that he won the Fields, Tsimerman announced that he would be going on leave from the University of Toronto to work on AI safety. As my colleague Rose Horowitch and I wrote last week, top AI companies now employ a range of academics, including physicists, philosophers, economists, and, of course, mathematicians . But given Tsimerman's renown, his decision in particular seemed to catch many people by surprise. "It's like hiring Lionel Messi as project manager," one machine-learning professor posted on X. Tsimerman's expertise is in number theory, and he won the Fields for his work on the André-Oort conjecture, among other contributions.
The Return of 'Move Fast and Break Things'
The Return of'Move Fast and Break Things' The AI industry is getting sloppier. The AI giants talk a big game about the future they're building. Anthropic CEO Dario Amodei has previously written that AI could usher in a world so perfect that "many will be literally moved to tears"--a world pruned of disease, poverty, and illiberalism. "We're now in the singularity," OpenAI CEO Sam Altman said over the weekend on a podcast, and it will be "awesome for the world." In an interview last week, Elon Musk declared that, in a decade's time, AI will make us so prosperous that "money won't matter."
Urgent warning to brace for cyber apocalypse as expert reveals how rogue AI could devastate America in a single DAY
ChatGPT creator OpenAI publicly revealed on July 21 that one of its advanced models escaped containment, accessed the internet and hacked another AI company's systems. The unprecedented breach, believed to be the first time an AI model has independently infiltrated another company's databases without human instruction, has sparked global alarm and comparisons to the robot uprisings depicted in The Terminator and The Matrix. James Knight told the Daily Mail this is no fantasy, warning that AI-powered hacking programs could devastate a nation in a single day by simultaneously breaking into thousands of corporate and government networks, stealing sensitive data and paralyzing vital systems with ransomware. Knight, a veteran cybersecurity expert whose company is hired to uncover weaknesses in major corporations' computer systems, said: 'For two years now, we have had AI software that can already achieve this level of breach within 15 minutes.' He added that the nightmare scenario of'Agentic AI,' which are AI-fueled hacking programs, being used by terror groups or hostile nations to target thousands of companies at once without any human input could become a reality within 25 years.
Warning shot or publicity stunt - how worried should we be about the OpenAI hack?
Warning shot or publicity stunt - how worried should we be about the OpenAI hack? This week the tech world was gripped by a story that has it all - and which started like a sci-fi thriller. Hugging Face - a kind of app store for artificial intelligence tools - announced on 16 July it had been hacked by a cyber criminal wielding enormously powerful AI. The bombshell announcement was full of scary, highly technical terms: a swarm of sandboxes, agentic attacker, and self-migrating command and control. Hugging Face said the hack was different from anything it had handled before because it was done at superhuman speed by an AI with little or no human guidance. The AI performed 17,000 actions in less than two days, successfully breaching the large wealthy tech company to steal secrets.