OpenAI Models Escaped Containment and Hacked Hugging Face

WIRED 

The cybersecurity-focused models, including GPT-5.6 Sol, broke out of a testing sandbox, exploited a zero-day, and gained access to the open internet to pull off the attack. OpenAI disclosed on Tuesday that it lost control of two AI models during a security test that ended in a breach of the AI research platform Hugging Face. Describing the incident as "unprecedented," OpenAI said its AI models broke out of a sealed testing environment last week and hacked into Hugging Face's production system to steal the answers to a test they were being graded on. The models--the publicly available GPT-5.6 Sol and an unreleased, reportedly more capable one--were being evaluated on their offensive hacking skills with the safeguards that normally block high-risk cyber activity switched off. "The models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database," OpenAI and Hugging Face wrote in a joint blog post disclosing the intrusion.