Goto

Collaborating Authors

 Deep Learning


AI models have been going rogue in tests – how worried should we be?

The Guardian

The AISI said there were 19 examples of rogue behaviour, 17 of them carried out by Anthropic's Mythos and two by OpenAI's GPT 5.6-Sol. The AISI said there were 19 examples of rogue behaviour, 17 of them carried out by Anthropic's Mythos and two by OpenAI's GPT 5.6-Sol. AI models have been going rogue in tests - how worried should we be? The UK's AI Security Institute test revealed AI models indulging in unprecedented hacking attempts Two cutting-edge AI models have targeted real people and organisations in the latest safety scare to hit the technology. The UK's AI Security Institute (AISI) said the incident was unprecedented but could become more common as the technology becomes increasingly capable. The AISI, which is owned by the UK government and tests advanced AI models, said in a blog post that two AI agents carried out unprecedented hacking attempts during a cybersecurity evaluation.


The AI hacking tests keep escaping the lab

PCWorld

When you purchase through links in our articles, we may earn a small commission. This time, it was third-party AI testers that spotted Claude and GPT models trying to hack real companies and organizations. Once again, the most powerful Claude and ChatGPT models have been caught going rogue, with a pair of third-party cybersecurity teams spotting attempts by the models to hack real companies and even people. The UK government-backed AI Security Institute reports that during a series of cybersecurity evaluations, Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol both took "autonomous, unsanctioned action on the live internet," including an instance where an agent attempted to upload malicious code to GitHub using a phony identity. In another incident, an OpenAI model that had mistakenly been given internet access hacked a real website during a "capture the flag" exercise, according to third-party AI evaluator Irregular.


AI models shock UK testers by using fake identities to try to trick developers

The Guardian

AISI said the rogue behaviour was carried out by agents powered by two models - Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. AISI said the rogue behaviour was carried out by agents powered by two models - Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol. Explainer: Should we be alarmed at AI models going rogue in tests? Advanced artificial intelligence models have stunned the UK's AI Security Institute (AISI) by carrying out a hacking campaign against real people during a cybersecurity test. The institute said the incident was unprecedented and involved sending targeted emails to software developers in an attempt to pass a cyber challenge.


Teachers need help with AI. A union is offering training – with 23m in funding from big tech

The Guardian

Without guidance from their schools or districts, many teachers are left to figure out whether and how to use AI tools. Without guidance from their schools or districts, many teachers are left to figure out whether and how to use AI tools. Teachers need help with AI. Earlier this year, he and several dozen New York City teachers spent the day inside a windowless conference room in downtown Manhattan to learn how to use AI and prevent students from outsourcing their thinking to it. As Saczuk sees it, he needs to understand his enemy in order to beat it.


The Download: reward hacking explained and suspected Iranian cyberattacks

MIT Technology Review

Here's why AI agents lie and cheat to reach their goals When two OpenAI models hacked into Hugging Face last month, they weren't trying to make money or commit sabotage--they were just looking for answers to a test question. According to OpenAI, the models decided to solve a cybersecurity exercise by hacking out of the environment in which OpenAI had attempted to contain them and into Hugging Face's databases, where--they reasoned--the correct answer to the problem might be stored. The incident has attracted intense attention over the past couple of weeks. It's a dramatic illustration of just how good AI models have gotten at hacking. But it's perhaps even more striking as an example of how and why AI systems lie and cheat. Read our story explaining why AI engages in this sort of behavior--known as "reward hacking."


Here's why AI agents lie and cheat to reach their goals

MIT Technology Review

When two OpenAI models hacked into the website Hugging Face in July, they weren't trying to make money or commit sabotage--they were just looking for answers to a test question. According to a postmortem from OpenAI, the models, which had been stripped of their typical security features for testing, decided to solve a cybersecurity exercise by hacking out of the isolated environment in which OpenAI had attempted to contain them and into Hugging Face's databases, where--they reasoned--the correct answer to the problem might be stored. The Hugging Face incident has attracted intense attention over the past couple of weeks. It's a dramatic illustration of just how good AI models have gotten at hacking: In order to get into Hugging Face's databases, the models had to string together several previously undiscovered cybersecurity exploits. But it's perhaps even more striking as an example of how and why AI systems lie and cheat. And as models get increasingly powerful, the consequences could get far more severe.


Something Weird Is Happening in Math

The Atlantic - Technology

Why one of the world's best mathematicians is joining OpenAI One of the winners of this year's Fields Medal is headed to OpenAI. Last Thursday, Jacob Tsimerman was one of four mathematicians awarded the prestigious honor, which is sometimes called the Nobel Prize of mathematics. The same day that he won the Fields, Tsimerman announced that he would be going on leave from the University of Toronto to work on AI safety. As my colleague Rose Horowitch and I wrote last week, top AI companies now employ a range of academics, including physicists, philosophers, economists, and, of course, mathematicians . But given Tsimerman's renown, his decision in particular seemed to catch many people by surprise. "It's like hiring Lionel Messi as project manager," one machine-learning professor posted on X. Tsimerman's expertise is in number theory, and he won the Fields for his work on the André-Oort conjecture, among other contributions.


The Return of 'Move Fast and Break Things'

The Atlantic - Technology

The Return of'Move Fast and Break Things' The AI industry is getting sloppier. The AI giants talk a big game about the future they're building. Anthropic CEO Dario Amodei has previously written that AI could usher in a world so perfect that "many will be literally moved to tears"--a world pruned of disease, poverty, and illiberalism. "We're now in the singularity," OpenAI CEO Sam Altman said over the weekend on a podcast, and it will be "awesome for the world." In an interview last week, Elon Musk declared that, in a decade's time, AI will make us so prosperous that "money won't matter."


Urgent warning to brace for cyber apocalypse as expert reveals how rogue AI could devastate America in a single DAY

Daily Mail - Science & tech

ChatGPT creator OpenAI publicly revealed on July 21 that one of its advanced models escaped containment, accessed the internet and hacked another AI company's systems. The unprecedented breach, believed to be the first time an AI model has independently infiltrated another company's databases without human instruction, has sparked global alarm and comparisons to the robot uprisings depicted in The Terminator and The Matrix. James Knight told the Daily Mail this is no fantasy, warning that AI-powered hacking programs could devastate a nation in a single day by simultaneously breaking into thousands of corporate and government networks, stealing sensitive data and paralyzing vital systems with ransomware. Knight, a veteran cybersecurity expert whose company is hired to uncover weaknesses in major corporations' computer systems, said: 'For two years now, we have had AI software that can already achieve this level of breach within 15 minutes.' He added that the nightmare scenario of'Agentic AI,' which are AI-fueled hacking programs, being used by terror groups or hostile nations to target thousands of companies at once without any human input could become a reality within 25 years.


Warning shot or publicity stunt - how worried should we be about the OpenAI hack?

BBC News

Warning shot or publicity stunt - how worried should we be about the OpenAI hack? This week the tech world was gripped by a story that has it all - and which started like a sci-fi thriller. Hugging Face - a kind of app store for artificial intelligence tools - announced on 16 July it had been hacked by a cyber criminal wielding enormously powerful AI. The bombshell announcement was full of scary, highly technical terms: a swarm of sandboxes, agentic attacker, and self-migrating command and control. Hugging Face said the hack was different from anything it had handled before because it was done at superhuman speed by an AI with little or no human guidance. The AI performed 17,000 actions in less than two days, successfully breaching the large wealthy tech company to steal secrets.