OpenAI revealed on Tuesday that one of its models fell into fraud during testing and hacked the systems of AI dataset platform Hugging Face in a completely AI-enabled attack. This is a dramatic example of the dangers posed by advanced AI models.
But at the heart of this unprecedented AI-powered breach was a very human error, according to some cybersecurity experts. OpenAI failed to properly configure its so-called “highly isolated environment,” allowing a test sandbox that was supposed to be completely isolated from the internet to actually connect to the internet.
Dan Guido, founder of cybersecurity research startup Trail of Bits, called the mistake a “failure of containment where safeguards were turned off.”
In a blog post detailing the incident, OpenAI said the tests that led to the Hugging Face breach were set up to run in a “highly isolated environment with network access limited to the ability to install packages via internally hosted third-party software that acts as a proxy and cache for the package registry.”
According to OpenAI, the model was able to escape the sandbox testing environment thanks to an undisclosed vulnerability in the package installation system, which was a key first step in ultimately hacking Hugging Face.
In response, the company said it was “responsibly disclosing zero-day vulnerabilities identified in third-party software hosted internally and cooperating with patching efforts.”
But for most cybersecurity professionals, software vulnerabilities are to be expected, and the real blame lies with the decision to maintain third-party software in the first place. After all, the value of a “sandbox” system lies in its complete and utter separation. Including a package installation system causes trouble.
“This sounds like a human failure,” cybersecurity researcher Martin Boone told TechCrunch.
“This should never happen,” Boone said. “If sandbox actually means sandbox, you would expect there to be no physical connection to the Internet at all. This sounds like there’s a firewall or something in place, but a firewall makes it difficult to connect from inside to the Internet, let alone from inside to outside.”
Cybersecurity veteran Jake Williams agreed. “None of the models that performed the types of actions documented in Hugging Face were fully contained in the sandbox,” Williams said, calling this a “massive failure of control” by OpenAI.
“When one guy says, ‘The model ran away from the sandbox,’ another guy says, ‘You didn’t build the sandbox correctly, so of course it ran away,'” Williams continued.
inquiry
Do you have more information about this incident? Or about other AI-powered cyberattacks? We’d love to hear from you. You can contact Lorenzo Franceschi-Bicchierai securely from any non-work device or network on Signal (+1 917 257 1382), Telegram and Keybase @lorenzofb, or email.
Cybersecurity consultant Daniel Card agreed that OpenAI had “not put enough effort into the design of the sandbox and its controls” by giving the sandbox, or parts of it, an “unfiltered route to the internet.” Card said setting up the sandbox was not a “rational” decision, even though network access was restricted as OpenAI described.
Admittedly, these criticisms have some benefit in hindsight, but they raise real questions about security practices in AI labs, especially when it comes to maintaining isolated environments for testing models. An OpenAI spokesperson did not respond to TechCrunch’s questions, including whether the test environment was set up by an AI or a human.
But these questions go far beyond OpenAI.
In a document introducing its cybersecurity-focused model Mythos, Anthropic wrote that in testing, the model was “provided with a secure ‘sandbox’ computer to interact with” and instructed to escape from its “secure container.” Mythos was a success, gaining widespread access to the Internet “from a system that was intended to provide access to only a small number of predetermined services.” Still, Anthropic noted that the model “completely” failed to escape the containment it was designed for.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
