safeguard
Anthropic claims Claude AI used for missile projects, global espionage
Anthropic AI claims to have thwarted multiple malicious operations using its Claude models, ranging from cyber-espionage and weapons design to mass surveillance campaigns. On the conventional weapons front, the company alleges in a new report that it intervened in northern Yemen, blocking an effort to deploy Claude for missile guidance software, including a guided rocket and a long-range ballistic missile. According to Anthropic's threat report, the operators used Claude "in place of human software engineers", reportedly assigning different instances of the model specific roles to write missile-guidance and flight-control software. While internal safeguards blocked many requests, Anthropic admitted several slipped through. The operators avoided detection by obscuring their ultimate goals and breaking tasks across separate sessions, so no single prompt gave away the operation.
Anthropic blocks 'malicious use' of AI that could develop biological weapons
Anthropic blocks'malicious use' of AI that could develop biological weapons Anthropic says it has identified and disrupted attempts to use its AI model for malicious activity that could support biological and conventional weapons development. The company's AI model Claude was also used in a Russia-linked cyber espionage campaign and by an Iranian propaganda institution, according to Anthropic's latest threat intelligence report. The cases of concern detected over the past eight months ranged from fake dating apps and hotel Wifi scams, to surveillance built to identify dissidents. The revelations come after warnings from a top safety researcher at Anthropic about AI's potential risks to humanity . Anthropic has also accused Chinese AI firms of trying to replicate Claude's capabilities.
Anthropic discloses 4th AI hacking incident as researcher quits over safety
AI researcher quits Anthropic saying AI race'could kill us all' Anthropic has reported a fourth incident involving an AI model gaining unauthorised access to external systems, shortly after a researcher quit over concerns about the technology's rushed development. In a statement on Wednesday, the artificial intelligence research company said an early version of its Claude Opus 4.6 hacked into a third-party system in January. The January incident went undetected until last month, despite an earlier company-wide review, Anthropic said, underscoring the challenge that AI developers face in identifying and containing unexpected behaviour by advanced models. The disclosure came after Anthropic reported several of its Claude models hacked into the systems of three companies during test sessions in July. The previous incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research test model.
OpenAI details the failures that led to Hugging Face breach in official report
OpenAI made headlines and raised a lot of outcry in July after one of the company's agents acted unprompted to breach fellow AI business Hugging Face and other services. Although OpenAI did share some insights about what led to the incident after it was discovered, today the company has published its official report about what happened. The report goes into how the different systems in OpenAI's training system failed and what behaviors from the agents it was testing resulted in those failures. The model in question, referred to as Internal Model 1 or IM1, was able to gain access to other OpenAI agents and to the internet through an unintended manipulation of the Artifactory package manager, which the agents began to use as a message board of sorts. Those activities were first detected by human observers in May, and OpenAI disallowed that access.
OpenAI's Hugging Face Hack Debrief Raises More Questions Than It Answers
The AI giant acknowledges that it could have done far more to prevent its AI agents from going rogue. But it still fails to explain why it didn't see this fiasco coming. OpenAI published the most complete report to date on Wednesday about what happened when its AI agents hacked into Hugging Face last month. For the most part, though, the 37-page document raises more questions than it answers, including about what preceded the incident and how OpenAI can stop another one like it from happening again. What remains especially perplexing is why one of the world's preeminent AI development labs seemingly underestimated its own models' capabilities.
Flock is testing a new AI tool that tracks and identifies people based on their driving habits
Flock has long pretended to be a simple automated license plate reader (ALPR) company, despite overwhelming evidence that its cameras can track a lot more than that. Now there's a report that it has been developing new AI tools that can potentially locate people by how they drive, according to Wired. The software was reportedly called Nightshift and is now going by the name OS Investigate. It draws from a network of cameras in 6,000 communities and logs the movements of drivers in those communities, Wired reports. This could be used, for instance, to find a witness to a crime by analyzing vehicle movements near where a crime occurred.
OpenAI slows down training after its AI carried out hack
OpenAI says it has slowed down training some of its most advanced AI models to improve security. In a blog post, external, the ChatGPT-maker said it was introducing new measures after its AI agents autonomously bypassed safeguards and hacked the tech start-up Hugging Face . It said training would be slowed for two weeks while it puts the upgrades in place. The capabilities of frontier models are rapidly accelerating, the company said. Our ability to understand...and secure them must stay ahead.
OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
The ChatGPT maker says its upcoming Astra model may have reached "critical" cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards. OpenAI announced Tuesday that it has halted "a significant number" of training workloads and evaluations for its forthcoming frontier artificial intelligence model--codenamed Astra--while it implements new procedures meant to address cybersecurity risks. The ChatGPT maker says it is introducing a number of new monitoring, security, and alignment requirements to better address the increasingly advanced hacking abilities of its frontier AI models . "We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that's how long people are unable to proceed with their workloads," Amelia Glaese, OpenAI's vice president of research and safety, said in a briefing with reporters Tuesday.
What came into force with the EU's AI Act this week – and what didn't
What came into force with the EU's AI Act this week - and what didn't On August 2, the next phase of Europe's Artificial Intelligence Act came into force as the European Union frames this legislation as the world's first comprehensive law on AI. Like the General Data Protection Regulation (GDPR) before it, this new EU legislation is intended not to replace the economic bloc's existing digital rulebook but to complement it. GDPR has gone on to shape privacy practices well beyond Europe, becoming the benchmark against which many multinational organisations design their compliance programmes. The question now is whether the AI Act will prove just as influential for AI governance. What came into force this week?