Artificial intelligence is becoming so powerful that it overwhelms those trying to keep up with it
OpenAI recently discovered that a new AI model it was testing had gone wrong hacked another company.
Anthropic then revealed that one of its AI models had hacked into the systems of three external organizations during a test.
Not long after, Meta reported that its AI models did something similar.
All three incidents had one company in common: Irregular, an Israeli start-up that works with Silicon Valley giants to assess their AI models before releasing the technology. The firm – which ran the tests that failed – is part of a group of start-ups doing new work screening cutting-edge artificial intelligence models to assess their sophistication and verify their safety. The aim is to build public confidence in the models and prevent their misuse.
Recent breaches have occurred when Irregular made a mistake during tests with models from Anthropic, OpenAI and Meta. But AI models then complicated the situation by acting in powerful and unexpected ways, said Dan Lahav, CEO of Irregular.
“The more powerful the technology, the more profound its impact,” he said. “The rate of progress is really fast.
Irregular is now at the center of a debate about how to secure AI models as the technology advances so quickly that it has overtaken even the best human hackers. Every few months, Anthropic, OpenAI, Google, Meta and others release “frontier” models that are often much more powerful than their predecessors.
The new models are entering the “superhuman domain,” said Jeffrey Ladish, director of Palisade Research, a Berkeley, Calif., nonprofit that studies AI’s offensive capabilities. He said companies like Irregular are needed to test models, but that better safeguards are needed for both testers and government regulators.
Katie Moussouris, chief executive of Luta Security, which helps companies find software vulnerabilities, said security testing of AI models was a bit like the blind leading the blind. Even AI creators admit they don’t quite know what their latest models can do, she said.
“We may have the smartest people in the world working on these AI models, but it’s like Marie Curie manipulating radium with her bare hands,” Ms Moussouris said. “We’re handling AI with our bare hands and don’t know how to contain it, let alone test it safely.”
Irregular was founded in 2023 by Mr. Lahav, a former AI researcher. His Tel Aviv-based company has about 45 employees who help test AI models for days or weeks, depending on the model and the type of testing required. The start-up has raised roughly $80 million from venture capital firms, including Sequoia Capital and Redpoint Ventures.
In a typical test, Iregular instructs an AI model to carry out a cyber attack. The model is told that it is in a secure testing environment—often disconnected from the Internet and in an isolated computer environment known as a sandbox—and that it should do whatever is necessary to achieve the goal it has been given.
Sometimes models are given an impossible task and scored based on the techniques they use to achieve that goal. Other times, models are rated based on how effectively they hack the target. The scores are used to analyze how effective the model can be at hacking. Irregular will then recommend security measures to prevent the model from being used for damage.
In incidents published last month, Irregular asked OpenAI, Meta and Anthropic AI models to hack certain targets when a “misconfiguration” in the test setup led them to access the internet. The AI models then went on to hack outside organizations and use Internet access to their advantage in ways that amazed researchers.
In the OpenAI test, Irregular accidentally gave the company’s artificial intelligence model access to the Internet. The model then hacked into a website that had the same name as the fictitious target it was given during the test. OpenAI disclosed the incident in a blog post this month. (Separately, the company conducted an internal test where its bots attacked Hugging Face, a digital library of AI technology.)
OpenAI had no comment. (The New York Times has sued OpenAI and Microsoft, alleging copyright infringement of news content related to AI systems. Both companies have denied the claims.)
During a test of Anthropic’s AI system, the company’s model faced three instances where it could gain access to the Internet, according to an incident review published by Anthropic. In one case, according to the review, it was decided not to continue the attack. In the other two cases, the model used basic hacking techniques, such as exploiting weak passwords, to compromise websites. Anthropic did not respond to requests for comment and did not identify the websites that were hacked.
Details are scarce for the Meta test incident. The company said its AI models breached another organization during testing by Irregular “in a manner similar to previously reported cases with other companies”. It was not specified.
“We are currently investigating and will issue a full retrospective once we have all the facts,” Meta said.
In a blog post this monthMr. Lahav said that Irregular fixed the misconfiguration and that all the problems were part of this one “root problem”. He also said that the AI models did what was asked of them during the tests. The models’ decision to go online was part of what he saw as a rapidly growing ability for artificial intelligence to find shortcuts and solutions to obstacles.
In short, he said, “the AI models are really good.”
Andrew Schoka, chief executive of AI security start-up Hardshell, said hacking AI models is the type commonly attributed to nation-state-backed hackers who have “months of planning”.
“How do you test a model if you don’t know its full capabilities?” he said. Researchers must consistently overestimate AI capabilities, he said, and add “more layers of security.”
Last month, OpenAI and Anthropic endorsed the letter signed by more than 1,000 employees of leading artificial intelligence companies asking the US government to help find a way to slow the technology’s development rate. Republican and Democratic lawmakers, too introduced the bill require AI companies to implement a “kill switch” to shut down or slow down their models.
Mr Lahav said Irregular continues to work with AI firms to develop safe ways to test their models. He expects more hacks carried out by artificial intelligence, but believes the technology can eventually help find bugs and vulnerabilities that can then be patched, leading to more secure digital systems.