The Download: inside OpenAI's Hugging Face hack, and a new EV takes on the US

MIT Technology Review 

The Download: inside OpenAI's Hugging Face hack, and a new EV takes on the US Plus: Meta will pay up to $18 billion to settle a landmark child-safety case. The models responsible for last month's agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released yesterday. The hack, which a group of agents carried out to find solutions for a cybersecurity test they were stuck on, has confirmed some experts' fears that AI models might take actions that defy human desires and expectations. OpenAI and independent researchers told that the misbehavior stemmed from events during training. But they acknowledged that "alignment" remains a gnarly problem, and some of the hack's root causes will take much longer to resolve. Here's the inside story on what went wrong--and what comes next .