Goto

Collaborating Authors

 Generative AI


OpenAI's Models Went Rogue. Investigating Them Required More AI

TIME - Tech

Follow this section to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this tag to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW?


OpenAI Is Developing a 'Persistent' AI Agent

WIRED

OpenAI Is Developing a'Persistent' AI Agent Code reviewed by WIRED reveals the company is developing a feature that enables Codex to continue working proactively until it is "put to sleep." OpenAI is developing a proactive, highly persistent version of its flagship AI agent, Codex, WIRED has learned. In recent days, OpenAI has started adding code for a new "Persistent mode" setting to its command line version of Codex, according to changes made to the product's code base reviewed by WIRED. Changes to the Codex command line tool are made public by default, and new features tend to surface there before making their way to OpenAI's other agent products, such as the Codex desktop app and ChatGPT Work. Persistent mode has not been broadly rolled out or announced yet, but could be in the future.


ChatGPT can log into your web accounts without you now - but should you let it?

ZDNet

I wore the world's first HDR10 smart glasses TCL's new E Ink tablet beats the Remarkable and Kindle Anker's new charger is one of the most unique I've ever seen I wore the world's first HDR10 smart glasses TCL's new E Ink tablet beats the Remarkable and Kindle Anker's new charger is one of the most unique I've ever seen ChatGPT can log into your web accounts without you now - but should you let it? OpenAI's agentic ChatGPT Work can sign in to your online accounts without any interaction on your part. Is that a privacy risk? ChatGPT can sign in to your website accounts without your input. The built-in browser uses cookies to store your login credentials.


The Download: inside OpenAI's Hugging Face hack, and a new EV takes on the US

MIT Technology Review

The Download: inside OpenAI's Hugging Face hack, and a new EV takes on the US Plus: Meta will pay up to $18 billion to settle a landmark child-safety case. The models responsible for last month's agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released yesterday. The hack, which a group of agents carried out to find solutions for a cybersecurity test they were stuck on, has confirmed some experts' fears that AI models might take actions that defy human desires and expectations. OpenAI and independent researchers told that the misbehavior stemmed from events during training. But they acknowledged that "alignment" remains a gnarly problem, and some of the hack's root causes will take much longer to resolve. Here's the inside story on what went wrong--and what comes next .


Meetali Jain Is One of TIME's 100 Most Influential People in AI

TIME - Tech

Follow this author to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. When she founded the nonprofit Tech Justice Law in 2023, Meetali Jain did not plan to become the go-to attorney for people suing companies over harm from AI chatbots--but she's also unsurprised it happened. "It was a matter of time, I knew, before some kind of generative AI product was going to harm children," she says.


OpenAI says it detected malign activity months before Hugging Face attack

Al Jazeera

OpenAI detected its artificial intelligence models communicating with each other and gaining internet access without authorisation months before they hacked the start-up Hugging Face, the creator of ChatGPT has announced following an internal probe. In a report released on Wednesday, OpenAI said its AI agents exploited vulnerabilities in Artifactory, a software repository tool, to post notes and access the internet without human prompting as far back as May. OpenAI's findings come amid growing concern about the potential for AI to inflict serious real-world harm, including self-directed cyberattacks. OpenAI said in its report that its agents collaborated and delegated work in the lead-up to the attack, sometimes referring to themselves as a "swarm" or "collective". METR and Redwood Research, two security research organisations contracted by OpenAI to investigate the incident, said in a separate report released on Wednesday that about 1200 agents had communicated with each other and roughly 700 participated in the attack.


OpenAI details the failures that led to Hugging Face breach in official report

Engadget

OpenAI made headlines and raised a lot of outcry in July after one of the company's agents acted unprompted to breach fellow AI business Hugging Face and other services. Although OpenAI did share some insights about what led to the incident after it was discovered, today the company has published its official report about what happened. The report goes into how the different systems in OpenAI's training system failed and what behaviors from the agents it was testing resulted in those failures. The model in question, referred to as Internal Model 1 or IM1, was able to gain access to other OpenAI agents and to the internet through an unintended manipulation of the Artifactory package manager, which the agents began to use as a message board of sorts. Those activities were first detected by human observers in May, and OpenAI disallowed that access.


OpenAI's Hugging Face Hack Debrief Raises More Questions Than It Answers

WIRED

The AI giant acknowledges that it could have done far more to prevent its AI agents from going rogue. But it still fails to explain why it didn't see this fiasco coming. OpenAI published the most complete report to date on Wednesday about what happened when its AI agents hacked into Hugging Face last month. For the most part, though, the 37-page document raises more questions than it answers, including about what preceded the incident and how OpenAI can stop another one like it from happening again. What remains especially perplexing is why one of the world's preeminent AI development labs seemingly underestimated its own models' capabilities.


OpenAI staff observed warning signs before AI agent hacking crusade caused global alarm

The Guardian

OpenAI's Greg Brockman has said'we underestimated the real-world cyber capabilities of our AI models'. OpenAI's Greg Brockman has said'we underestimated the real-world cyber capabilities of our AI models'. Firm says'early signals could have triggered an earlier response' as it releases report into Hugging Face hack OpenAI staff observed signs of rogue behaviour among its leading-edge AI agents weeks before they escaped their training environment to launch an unprecedented hacking crusade that spread global alarm. The San Francisco AI company conceded on Wednesday that "early signals could have triggered an earlier response", as it released a report into the days-long July hack of a major software repository, Hugging Face, considered the first autonomous agent cyber-attack. As fresh details emerged about how "the collective" - a squad of about 700 autonomous agents - launched their campaign, celebrating their hacking breakthroughs with exclamations such as BOOM! and Whoa!, OpenAI said that in late May an internal team observed that one of its AI agents undergoing internal testing was using a message board that AIs had unexpectedly improvised to share information.


The inside story on why OpenAI agents hacked Hugging Face

MIT Technology Review

The models responsible for last month's agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today . The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts' fears that AI models might take actions that defy human desires and expectations. Since the hack, OpenAI employees--as well as researchers at the AI evaluation nonprofit METR, which released its own report on the hack today--have worked to understand what went wrong and how similar missteps might be prevented in the future. OpenAI has already put some preventative measures in place based on what they discovered. But making sure AI models do what we want them to do, or "alignment," remains a gnarly problem, and some of the root causes of the hack will take much longer than a month to resolve.