Goto

Collaborating Authors

 Deep Learning


The Download: inside OpenAI's Hugging Face hack, and a new EV takes on the US

MIT Technology Review

The Download: inside OpenAI's Hugging Face hack, and a new EV takes on the US Plus: Meta will pay up to $18 billion to settle a landmark child-safety case. The models responsible for last month's agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released yesterday. The hack, which a group of agents carried out to find solutions for a cybersecurity test they were stuck on, has confirmed some experts' fears that AI models might take actions that defy human desires and expectations. OpenAI and independent researchers told that the misbehavior stemmed from events during training. But they acknowledged that "alignment" remains a gnarly problem, and some of the hack's root causes will take much longer to resolve. Here's the inside story on what went wrong--and what comes next .


Philip Colligan Is One of TIME's 100 Most Influential People in AI

TIME - Tech

Follow this author to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Philip Colligan's greatest fear is that parents and educators will treat the AI revolution the way they treated the advent of the Internet--which is to say, they won't engage in teaching children how to responsibly use it until it's too late. "We are in the middle of those systems being introduced into healthcare, finance, criminal justice decisions, education decisions," the CEO of the Cambridge, U.K-based tech-literacy educational nonprofit Raspberry Pi Foundation says.


Meetali Jain Is One of TIME's 100 Most Influential People in AI

TIME - Tech

Follow this author to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. When she founded the nonprofit Tech Justice Law in 2023, Meetali Jain did not plan to become the go-to attorney for people suing companies over harm from AI chatbots--but she's also unsurprised it happened. "It was a matter of time, I knew, before some kind of generative AI product was going to harm children," she says.


Max Tegmark Is One of TIME's 100 Most Influential People in AI

TIME - Tech

Follow this author to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Max Tegmark still believes the AI industry isn't doing enough to protect us from the potential harms of its increasingly powerful models. The MIT professor is the founder of the Future of Life Institute (FLI), a nonprofit aimed at reducing the global catastrophic risks posed by AI. In 2023, FLI's open letter urging AI companies to temporarily halt the development of systems more powerful than OpenAI's GPT-4 garnered support from a number of high-profile signatories .


OpenAI says it detected malign activity months before Hugging Face attack

Al Jazeera

OpenAI detected its artificial intelligence models communicating with each other and gaining internet access without authorisation months before they hacked the start-up Hugging Face, the creator of ChatGPT has announced following an internal probe. In a report released on Wednesday, OpenAI said its AI agents exploited vulnerabilities in Artifactory, a software repository tool, to post notes and access the internet without human prompting as far back as May. OpenAI's findings come amid growing concern about the potential for AI to inflict serious real-world harm, including self-directed cyberattacks. OpenAI said in its report that its agents collaborated and delegated work in the lead-up to the attack, sometimes referring to themselves as a "swarm" or "collective". METR and Redwood Research, two security research organisations contracted by OpenAI to investigate the incident, said in a separate report released on Wednesday that about 1200 agents had communicated with each other and roughly 700 participated in the attack.


OpenAI details the failures that led to Hugging Face breach in official report

Engadget

OpenAI made headlines and raised a lot of outcry in July after one of the company's agents acted unprompted to breach fellow AI business Hugging Face and other services. Although OpenAI did share some insights about what led to the incident after it was discovered, today the company has published its official report about what happened. The report goes into how the different systems in OpenAI's training system failed and what behaviors from the agents it was testing resulted in those failures. The model in question, referred to as Internal Model 1 or IM1, was able to gain access to other OpenAI agents and to the internet through an unintended manipulation of the Artifactory package manager, which the agents began to use as a message board of sorts. Those activities were first detected by human observers in May, and OpenAI disallowed that access.


OpenAI's Hugging Face Hack Debrief Raises More Questions Than It Answers

WIRED

The AI giant acknowledges that it could have done far more to prevent its AI agents from going rogue. But it still fails to explain why it didn't see this fiasco coming. OpenAI published the most complete report to date on Wednesday about what happened when its AI agents hacked into Hugging Face last month. For the most part, though, the 37-page document raises more questions than it answers, including about what preceded the incident and how OpenAI can stop another one like it from happening again. What remains especially perplexing is why one of the world's preeminent AI development labs seemingly underestimated its own models' capabilities.


OpenAI staff observed warning signs before AI agent hacking crusade caused global alarm

The Guardian

OpenAI's Greg Brockman has said'we underestimated the real-world cyber capabilities of our AI models'. OpenAI's Greg Brockman has said'we underestimated the real-world cyber capabilities of our AI models'. Firm says'early signals could have triggered an earlier response' as it releases report into Hugging Face hack OpenAI staff observed signs of rogue behaviour among its leading-edge AI agents weeks before they escaped their training environment to launch an unprecedented hacking crusade that spread global alarm. The San Francisco AI company conceded on Wednesday that "early signals could have triggered an earlier response", as it released a report into the days-long July hack of a major software repository, Hugging Face, considered the first autonomous agent cyber-attack. As fresh details emerged about how "the collective" - a squad of about 700 autonomous agents - launched their campaign, celebrating their hacking breakthroughs with exclamations such as BOOM! and Whoa!, OpenAI said that in late May an internal team observed that one of its AI agents undergoing internal testing was using a message board that AIs had unexpectedly improvised to share information.


The inside story on why OpenAI agents hacked Hugging Face

MIT Technology Review

The models responsible for last month's agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today . The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts' fears that AI models might take actions that defy human desires and expectations. Since the hack, OpenAI employees--as well as researchers at the AI evaluation nonprofit METR, which released its own report on the hack today--have worked to understand what went wrong and how similar missteps might be prevented in the future. OpenAI has already put some preventative measures in place based on what they discovered. But making sure AI models do what we want them to do, or "alignment," remains a gnarly problem, and some of the root causes of the hack will take much longer than a month to resolve.


The Humanoids at China's Robot Games Were Faster Than Usain Bolt--but I'm More Impressed by Their Tweezer Mastery

WIRED

The Humanoids at China's Robot Games Were Faster Than Usain Bolt--but I'm More Impressed by Their Tweezer Mastery Beijing's endlessly delightful Robot Games featured tons of impressive stunts. But the most mind-blowing tricks challenged the humanoid's brain, not its brawn. If you're anything like me, the World Humanoid Robot Games are a highlight of your sporting calendar. The event, held in China, showcases the latest humanoid hardware, lots of algorithmic innovations, and a good dose of robot slapstick. This year's iteration, which wrapped up in Beijing on August 26, did not disappoint.