The Download: reward hacking explained and suspected Iranian cyberattacks

MIT Technology Review 

Here's why AI agents lie and cheat to reach their goals When two OpenAI models hacked into Hugging Face last month, they weren't trying to make money or commit sabotage--they were just looking for answers to a test question. According to OpenAI, the models decided to solve a cybersecurity exercise by hacking out of the environment in which OpenAI had attempted to contain them and into Hugging Face's databases, where--they reasoned--the correct answer to the problem might be stored. The incident has attracted intense attention over the past couple of weeks. It's a dramatic illustration of just how good AI models have gotten at hacking. But it's perhaps even more striking as an example of how and why AI systems lie and cheat. Read our story explaining why AI engages in this sort of behavior--known as "reward hacking."