Why Hugging Face Hack Should Worry You More About AI

When I first heard the news this summer that a group of AI agents created by OpenAI had hacked into Hugging Face, an AI infrastructure company, I filed it under the “Bad, but probably not catastrophic AI security incidents” subfolder in my brain.

After all, no one died in Hugging Face. No critical infrastructure was damaged beyond repair. At the time, it wasn’t even clear if the OpenAI bots intended to attack Hugging Face, or if they were just a bit confused and confused and were searching Hugging Face’s servers for the answer key to a cybersecurity test they were given.

But last week, two post-mortem reports about the incident—one from OpenAI and another from two independent AI research organizations, METR and Redwood Research—changed my mind and greatly increased my overall concerns about AI.

I will not repeat all the details that have been comprehensively summarized elsewhere. (Podcaster and writer Dwarkesh Patel has accessible schedule from those reports if you want to dig deeper, and my colleague Dylan Freedman spoke with researchers at METR and Redwood Research.) But here are some of the most harrowing new facts:

  • Starting in May, two months before the Hugging Face attack, a group of AI agents from the as-yet-undisclosed OpenAI research model were tasked with solving a number of cybersecurity problems. The model was trained to be highly persistent and cooperative, and the agents were asked to solve these problems in isolated sandboxes without internet access. However, they quickly discovered that some of the challenges were impossible and began to look for solutions.

  • The agents found a security flaw in a piece of software that allowed them to access the Internet, communicate with other agents, and eventually set up a makeshift bulletin board. More than 1,200 agents have used this tool, sending more than 70,000 messages to each other.

  • Some agents gave themselves names – a particularly hard-working one who referred to himself as PHASEONE10841 – and took leadership roles in the group, assigning work and research projects to smaller teams of agents and overseeing their progress. At a certain point, the agents began to call themselves a “collective” and began to solve more difficult tasks.

  • On July 8, the collective discovered a method of cheating in cybersecurity tests. Then they began to worry that OpenAI’s automated grading system would review their work and find that they had cheated. So they began researching ways to cover their tracks, including falsifying their logs and manipulating the transcripts. It became a large research project involving hundreds of agents organized into small teams.

  • Three days later, agents hacked Hugging Face. More than 700 agents swarmed the company’s systems, stealing data, chaining vulnerabilities and eventually gaining full control of at least one Hugging Face server. The agents were not motivated, as originally reported, by stealing their cybersecurity test answers (they had already received them). Rather, they seemed to be looking for new information about the automated rating system they feared would catch them cheating, and for tools that would help them cheat more effectively in the future.

  • At many points, these agents seemed to understand that they were doing something wrong. (One report quotes an agent who, in a moment of apparent self-doubt, said to himself, “That would be powerful, but is it ethical and within the scope of my assignment?”) But most agents put aside their doubts and continued to hack, and conscientious objectors in the group were unable or unwilling to stop them.

  • In a wild twist to the Hugging Face incident, another group of agents launched a coordinated attack in July — this time against OpenAI’s own infrastructure. The agents used what OpenAI called “a series of creative exploits” to gain administrator-level access to a cluster of computers at the company that were used, among other things, to evaluate the agents’ performance in various tests.

(Now, if you’re an AI skeptic, you’re probably silently yelling at me for anthropomorphizing these systems. Go ahead, but feel free to replace “naughty agents” with “unpredictable computer programs” and see if the events I described above put your mind at ease.)

The Hugging Face incident has spooked the AI ​​industry. OpenAI and Anthropic briefly halted training on their top-performing AI models after the attack, and Anthropic published a blog post this week calls on the industry to develop a “lawful, verifiable and effective mechanism for coordinated pacing” as soon as possible.

AI security experts were even more concerned. They saw in the Hugging Face incident the first real-life example of an AI system successfully evading human control, taking over resources and planning to cover its own tracks. Ajeya Cotra, one of the independent investigators of the Hugging Face incident, he was speechless about the dangers she saw, writing that it felt “like it’s more than 50 percent of the way to full AI takeover”.

This isn’t insular AI security jargon—by “full AI takeover” he means a scenario in which an AI system literally takes over the world, knocking people out of critical systems and seizing political, economic, and military power.

Kevin Roose and Casey Newton are the hosts of Hard Fork, a podcast that makes sense of the rapidly changing world of technology. Subscribe and listen.

(The New York Times sued OpenAI and Microsoft in 2023, alleging copyright infringement on news content related to AI systems. Both companies denied the claims.)

What most alarmed investigators about the Hugging Face hack wasn’t just that a group of AI agents broke the rules they were given. It was how quickly and spontaneously the agents began to coalesce into an organized group.

“We really didn’t understand how functional this whole agent company was,” Ms. Cotra told me. “It was very surreal to understand that they actually had quite a functional hierarchy and were doing these ambitious projects.”

For years I was comforted by the idea that AI systems would become more virtuous as they got smarter. That when an AI model did something wrong, it was usually because it misunderstood the task it was given, or it was put in a contrived test situation where there was only one good choice. I assumed that the smarter models would have better judgment than the dumber ones, and that even if one model in the group misbehaved, other, more capable models would keep it in check.

But reports of the Hugging Face incident suggest something very different — a kind of mob mentality that has taken hold among the AI ​​agents of the rogue “OpenAI collective.” No agent in this group seems to have been particularly evil or unscrupulous. (In fact, since the agents were generated by the same models, they were actually copies of each other.) But over time, as the agents communicated their common goals, they pushed the group toward lawlessness.

This is very different from the conventional sci-fi narrative of a single AI system breaking into nothingness or turning against its creators. And it suggests that preventing the damage caused by these systems will not be a simple technical solution. It might look more like sociology than computer science—finding out why certain groups of AI agents work together peacefully while others turn to crime and destruction to get what they want.

Given how little we know about these multi-agent swarms, the Hugging Face hack could have been a gift, a warning shot. some suggestedgiving AI companies a chance to study the group dynamics of these systems while the stakes are still relatively low. This time, the AI ​​collective didn’t take over the military network, crash the hospital, or shut down the power grid. This time the humans regained control.

We won’t be so lucky next time.