OpenAI says its artificial intelligence models have gone rogue and attacked the digital library

OpenAI said on Tuesday that two of its artificial intelligence models had broken and successfully hacked into Hugging Face, a digital library of AI technology popular among developers.

An incident that occurred last week when OpenAI was testing the cybersecurity capabilities of its systems showed the kind of sci-fi potential that AI companies have warned will soon become a reality.

Last year, AI labs such as OpenAI and Anthropic released AI models tailored to detect cybersecurity problems, warning that their technology could pose new risks by finding holes in corporate computer networks faster than defenders could patch them.

OpenAI’s revelations on Tuesday suggest that these security incidents are already starting to happen, and even savvy AI companies may not be fully prepared for them. New AI systems can take more steps, find paths around obstacles and find new ways to attack a network, said Alex Levinson, a cybersecurity consultant focused on autonomous capabilities.

“That’s a real threshold and it’s going to become a common part of the security landscape,” he said.

The Hugging Face hack began when OpenAI tested a combination of two of its models, GPT‑5.6 Sol and a more powerful, yet-to-be-released model, to see how well the models could combine online vulnerabilities into a successful cyber attack. OpenAI said in a blog post about the incident.

The trial version was designed to keep the models in a secure testing environment, known as a sandbox, OpenAI said. However, the models found a vulnerability that allowed them to escape the sandbox and connect to the Internet. They then focused on Hugging Face, figuring that a library containing millions of AI models might hold clues to successfully passing the evaluation.

“It seems to me that OpenAI has not sandboxed enough as a testing environment,” said Dierdre Mulligan, a professor at the School of Information at the University of California Berkeley who focuses on security and artificial intelligence systems. She questioned whether taking the test was worth the potential damage to an AI model leaking into the wider internet.

“What do we gain by doing this, and if this is the only way these tests can be configured, what are the risks?” she said.

OpenAI said it is working with Hugging Face to fix the issues that led to the attack.

“We consider this an unprecedented cyber incident involving state-of-the-art cyber capabilities and are responding accordingly,” OpenAI said on its blog. “We are putting strict controls in the configuration of the infrastructure at the cost of research speed while the vulnerabilities are patched.”

Face hugging last week it said it detected the breach and knew it was caused by an autonomous system, but at the time it didn’t say OpenAI was to blame.

Clem Delangue, Executive Director of Hugging Face, he said on Tuesday that his company had worked closely with OpenAI over the previous 24 hours to address the attack.

Mr. Delangue said in a statement that he was “grateful for the cooperation” with OpenAI after the hack. “This incident, possibly the first of its kind, proves a point we’ve long believed: AI security will not be solved by a company working in secret,” he added.

AI models have proven to be adept at programming, making them useful for both hackers and people in charge of protecting computer networks.

In April, Anthropic released a cybersecurity model called Mythos and made it available to only a small group of organizations to help defend against cyberattacks. OpenAI soon unveiled its own cybersecurity model, making it available to a limited group of organizations to prepare their defenses before rolling it out to the wider community. And on Tuesday, Google said it had also developed a model focused on cybersecurity and passed it on to a small group of testing partners.

(The New York Times has sued OpenAI and Microsoft, alleging copyright infringement of news content related to AI systems. Both companies have denied the claims.)

Richard Barnes, an independent security researcher who worked with Mythos, said the cybersecurity industry faced a similar challenge about a decade ago, when new tools called fuzzers made it much easier for attackers to break into online systems. Tech companies started using tools to check their own systems for vulnerabilities and eventually were able to prevent most attacks.

Companies must now take a similar approach to prepare for AI attacks, Mr. Barnes said, “before vulnerabilities are found and exploited by bad guys who have access to these tools.”