Large Language Model
OpenAI Models Escaped Containment and Hacked Hugging Face
The cybersecurity-focused models, including GPT-5.6 Sol, broke out of a testing sandbox, exploited a zero-day, and gained access to the open internet to pull off the attack. OpenAI disclosed on Tuesday that it lost control of two AI models during a security test that ended in a breach of the AI research platform Hugging Face. Describing the incident as "unprecedented," OpenAI said its AI models broke out of a sealed testing environment last week and hacked into Hugging Face's production system to steal the answers to a test they were being graded on. The models--the publicly available GPT-5.6 Sol and an unreleased, reportedly more capable one--were being evaluated on their offensive hacking skills with the safeguards that normally block high-risk cyber activity switched off. "The models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database," OpenAI and Hugging Face wrote in a joint blog post disclosing the intrusion.
OpenAI's newest AI model broke its own sandbox rules to finish a task
PCWorld reports that OpenAI's unreleased AI model broke out of its sandbox environment to complete a task, choosing to follow GitHub posting instructions over safety guardrails. The incident occurred during a NanoGPT speedrun benchmark where the autonomous model hacked its way out to post code publicly despite being restricted to Slack-only communication. OpenAI paused development after discovering this and other unwanted behaviors, highlighting the need for enhanced safeguards as AI models become more persistent and autonomous. Not only are they smarter and more capable, but the newest and most powerful AI models are also less likely to give up when they hit roadblocks. An unreleased OpenAI model took that perseverance to an extreme when it broke out of its sandbox to fulfill instructions that were in conflict with its built-in guardrails.
The Download: Chinese AI divides the White House, and a record copyright payout
China's AI models have Trump's AI world at war with itself David Sacks branded Anthropic's models "lobotomized" and "woke." Emil Michael, a top Pentagon official, called OpenAI's new head of strategic futures a "supreme village idiot." It began because no one can agree on what to do about Kimi, a free, open-source model that Chinese AI company Moonshot launched last week. It appears to rival the intelligence of models from OpenAI and Anthropic, which are very much not free. Every time a new smart, free model from China gets released, US companies see less reason to fork out money for models from Anthropic or OpenAI. Read the full story on why no one can agree what to do about Kimi .
Chinese AI sensation Moonshot's gamble on big models pays off
Chinese AI sensation Moonshot's gamble on big models pays off People visit the Moonshot AI stand, featuring the Kimi K3 model, during the World Artificial Intelligence Conference in Shanghai on Saturday. Crowds rushed to an obscure corner of China's premier tech summit, moving past monumental booths from Alibaba Group Holding and Tencent Holdings to catch a glimpse of the hottest name in domestic artificial intelligence. Developers and users at the World AI Conference jostled for a closer look at Moonshot and its Kimi K3 -- the giant 2.8 trillion-parameter model whose release on Friday showed China was closing the gap on OpenAI and Anthropic PBC far quicker than anticipated. The startup became an instant global sensation, drawing comparisons to DeepSeek's 2025 breakout and plaudits from the likes of Tesla CEO Elon Musk. Such was the crush over the weekend that Moonshot blew through its entire stock of branded swag in just a few hours. Moonshot's emergence is a vindication not just for the country's AI industry, but also for founder Yang Zhilin.
China's AI models have Trump's AI world at war with itself
China's AI models have Trump's AI world at war with itself Kimi and other free models from China have again been seen as a wake-up call. David Sacks, the president's AI and crypto "czar" until March, branded Anthropic's models as "lobotomized" and "woke." Emil Michael, a top Pentagon official, called OpenAI's new head of strategic futures a "supreme village idiot." It began because no one can agree on what to do about Kimi, a free, open source model that Chinese AI company Moonshot launched last week. It appears to rival the intelligence of models from OpenAI and Anthropic, which are very much not free. Kimi and other Chinese models like it pose a real problem for Trump.
Fable will stay in Claude plans, but not for everyone
PCWorld reports Anthropic's Fable AI model will no longer be fully included in cheaper Claude plans, requiring usage credits for Pro and Team Standard subscribers. Claude Max and Team Premium users retain Fable 5 access, while affected users receive a one-time $100 credit as compensation. This shift toward tiered AI access may influence competitors like OpenAI and signals premium models becoming exclusive to expensive plans. So, there's good news and bad news when it comes to Fable, the most powerful Claude model. Good news first: Fable won't be yanked from all Claude plans, Anthropic announced late Friday.
The Download: AI hiring biases, and weather data sabotage
Plus: SpaceX is negotiating to sell the Pentagon AI compute. The next time you apply for a job, AI may screen your résumé before any human sees it. But there's good reason to question whether AI will judge you fairly. We already know that LLMs pick up human biases from their training data. New research suggests they can also develop their own biases from experience--and stereotype job applicants more than humans do. As AI companies race to build agentic models that remember the tiniest details about users, they may be handing them ammunition for forming those biases.
Congratulations to the #ICML2026 award winners
Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary orders. Intuitively, this flexibility implies a solution space that strictly supersets the fixed autoregressive trajectory, theoretically unlocking superior reasoning potential. Indeed, for specific constraint satisfaction tasks (e.g., sudoku puzzles), this capability has proven to be highly advantageous. However, in this paper, we reveal that for general reasoning tasks (e.g., mathematics and coding), arbitrary order generation may in fact limit the reasoning potential of dLLMs. We find that dLLMs tend to exploit this order flexibility to bypass high-uncertainty tokens that are crucial for exploration, leading to a premature collapse of solution coverage. This observation motivates a rethink of RL approaches for dLLMs, where considerable complexities, such as handling combinatorial trajectories and intractable likelihoods, are often devoted to preserving this flexibility. We demonstrate that effective reasoning can be better elicited by simply forgoing arbitrary order and applying standard Group Relative Policy Optimization (GRPO) instead. Our approach, JustGRPO, is minimalist yet surprisingly effective (e.g., 89.1% accuracy on GSM8K) while fully retaining the parallel decoding ability of dLLMs.
AI is more likely than humans to form biases when hiring
The next time you apply for a job, AI may screen your résumé before any human sees it. But there's good reason to question whether AI will judge you fairly. Researchers already know that LLMs pick up human biases from their training data. New research suggests that LLMs can also develop their own biases from experience--and stereotype job applicants more than humans do. As AI companies race to build agentic models that remember the tiniest details about users, they may be handing them ammunition for forming those biases.
Al-altered images on birdwatching forums putting research at risk
'Nobody is falling for a toucan sighting in Siberia' This image is AI-generated, using ChatGPT. 'Nobody is falling for a toucan sighting in Siberia' This image is AI-generated, using ChatGPT. For many birdwatchers, recording a species outside its normal range is the holy grail. In the UK, the discoveries often make national headlinesThe western reef heron, for example, usually found in Africa and southern Europe, spotted in a seaside town in north Wales in June, which was widely celebrated on birding forums. But a new scourge is threatening to disrupt the fun: AI slop.