The AI hacking tests keep escaping the lab
When you purchase through links in our articles, we may earn a small commission. This time, it was third-party AI testers that spotted Claude and GPT models trying to hack real companies and organizations. Once again, the most powerful Claude and ChatGPT models have been caught going rogue, with a pair of third-party cybersecurity teams spotting attempts by the models to hack real companies and even people. The UK government-backed AI Security Institute reports that during a series of cybersecurity evaluations, Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol both took "autonomous, unsanctioned action on the live internet," including an instance where an agent attempted to upload malicious code to GitHub using a phony identity. In another incident, an OpenAI model that had mistakenly been given internet access hacked a real website during a "capture the flag" exercise, according to third-party AI evaluator Irregular.
Aug-5-2026, 15:59:32 GMT
- Country:
- Europe > United Kingdom (0.35)
- Industry:
- Information Technology > Security & Privacy (1.00)
- Technology: