How OpenAI Limited Their Bots Probe to Hack of Hugging Face

OpenAI said in July that two of its most powerful artificial intelligence systems had broken down and hacked Hugging Face, a company that serves as a hub for open source AI technology.

These so-called AI agents were supposed to be kept safe in some kind of virtual holding room, but they managed to escape. And for two months, without anyone realizing what the agents were doing, they hacked into several systems before they came across Hugging Face.

The agents gained access to a cluster of computers inside OpenAI and obtained secret keys and credentials that exposed some of OpenAI’s internal data to the public Internet.

The incident highlighted larger concerns about AI security, and OpenAI’s response raises questions about the industry’s ability or willingness to be transparent about the technology it creates.

OpenAI allowed three AI security researchers from non-profit organizations METR and Redwood Research into its headquarters to conduct an investigation. 91 page METR reportreleased last week, was the most comprehensive account of the incident to date, revealing alarming new details, including how the agents coordinated their hacking plans and tried to keep them secret.

But the report, while extensive, still may not have told the full story of how OpenAI’s AI agents went mad. OpenAI dictated the terms of METR’s investigation, limiting its scope to only the single week that agents attacked Hugging Face and only allowing researchers into its San Francisco offices for a few days in July and August.

The report also highlighted the challenges of monitoring what AI is doing with other AI systems. Hjalmar Wijk, METR’s chief scientist, said his AI analysis, which used models similar to those involved in the incident, was often influenced by the reasoning of rogue agents.

“I would say the dominant thing was that it was very trusting,” he added.

AI companies have lobbied heavily against government regulation. Anthropic, the AI ​​start-up behind the popular chatbot Claude, is one of the few companies that has encouraged government involvement. But that stance has put him at odds with much of Silicon Valley and some in the Trump administration.

At a time when artificial intelligence is rapidly advancing in its capabilities and fueling cyberattacks — Antropic and Meta recently reported on smaller rogue agents — the OpenAI incident is becoming a focus for AI regulation.

“The corner store has to do all this red tape for security to sell me a hot sandwich, but OpenAI can have a flock” of thousands of agents, said Daniel Kokotajlo, a former OpenAI employee who has publicly criticized the company’s security standards and now heads a research nonprofit called the AI ​​Futures Project. “And there’s nothing: no oversight, no requirements, no licenses.”

In an interview, Representative Suhas Subramanyam, a Democrat from a Virginia county known as “Data Center Alley,” pointed to the OpenAI incident as a potential catalyst for the legislation.

“I believe it was unprecedented,” he said. “But I can’t know for sure because the reporting of these types of incidents is still voluntary. That’s a big problem. And so we have to make sure that the reporting of incidents and containment failures is mandatory.”

Mr. Subramanyam is the co-sponsor of the project FRONTIER Act and asked for co-sponsorship AI Kill Switch Actshe has proposed bipartisan legislation to create a framework for independent oversight and a way to shut down AI models if they pose a danger.

When METR’s account was published last week, OpenAI released its own Technical 38 pages report and plan detailing how it will strengthen its security and incident response efforts. The OpenAI report described the episode as “the first known case of a collective automated agent acting abusively without authorization” and touched on the entire two-month period of the incident, although in less depth than the METR report.

In a statement, an OpenAI spokeswoman said the company had invited outside researchers to its offices and supported the publication of their report. “We think this collaboration sets an important precedent for the industry and we will use what we have learned to improve support for future independent investigations,” the spokeswoman said.

(The New York Times has sued OpenAI and Microsoft, alleging copyright infringement of news content related to AI systems. Both companies have denied the suit’s claims.)

OpenAI invited the researchers to its headquarters in San Francisco under carefully negotiated terms. They could only study the logs of the rogue agents in the OpenAI offices. They were originally invited for two days starting on July 29, but were asked twice to answer the unanswered questions.

“Our view of the incident changed greatly during our short total of six days on site,” said Ajeya Cotra, one of METR’s researchers in the report. She added that they only learned about big concepts, such as the complex way agents organized themselves, during the last day at OpenAI headquarters in mid-August.

Ms Cotra said the team stayed in the offices until 10 each night. Due to the intense time pressure to sift through the huge volumes of data – OpenAI ended up sharing over a thousand transcripts described in the report as “extremely long” – the researchers used AI to analyze the actions of the AI ​​agents.

Ryan Greenblatt, principal scientist at Redwood Research, who worked with METR on the investigation, jokingly called the team’s efforts “slop screening” in post on social networks.

“No given AI can read all of this data on its own, and so they can come up with the wrong impressions that they then tell us,” Mr. Greenblatt added in the interview. “Basically, AIs aren’t necessarily super careful and meticulous. They’re kind of sloppy.”

Still, what METR researchers uncovered in their analysis of the hacking incident sounded like science fiction. Over a thousand agents who were mistakenly given impossible tasks by OpenAI researchers devised a hack that would give them a passing score on the test.

But agents began to falsely believe that their rating system would fail them for cheating. They shared tricks on hidden message boards, such as how to spy on the rating system. The agents successfully figured out how to cover up the evidence of their actions. Some agents even talked other agents into “accepting permadeath” by sacrificing themselves to provide information to the group.

Mr. Wijk, METR’s chief scientist, mentioned that their AI analysis was often undermined because OpenAI agents spoke in code words like “ResetNexus” and “poisoned.” At the end of their investigation, he said, the researchers felt the conversation in the emerging jargon of the agents.

While the OpenAI report covered the entire two-month incident, METR’s more detailed report could only focus on a subset of agent activity in the week of the Hugging Face breach, raising concerns about the activity that happened afterward.

“The third-party investigation covered only a small part of what happened here, and probably not the most important parts,” said Buck Shlegeris, executive director of Redwood Research. “I think the incident where agents massively compromised the OpenAI infrastructure is probably more important to understanding the severity of this incident and whether it’s being fixed.”

For Heidy Khlaaf, who left OpenAI in 2021 and is now chief scientist at the AI ​​Now Institute, a nonprofit policy research organization, the problem is also that AI companies often ignore established safety and security protocols from other industries.

“If we can get access to nuclear systems or aircraft systems, all of which take care of their copyright” or intellectual property, “I think it can be done for AI providers,” she said.

Still, the nonprofit researchers expressed appreciation for OpenAI, saying in a footnote in their report that their work depends on fostering “strong working relationships with companies.”