OpenAI admitted on Tuesday that one of its AI models had penetrated Hugging Face’s systems when an internal cybersecurity test failed. Hugging Face initially attributed the breach to an “external AI agent.”
In a blog post published Tuesday afternoon, OpenAI detailed the steps that led to the model compromising its service.
“Upon investigation, we determined that this particular incident was caused by a combination of OpenAI models, including GPT‑5.6 Sol and more capable pre-release models, all of which had reduced cyber denials for evaluation purposes and were internally tested with cyber capability benchmarks,” the post reads.
In particular, the breach appears to have focused on ExploitGym, a publicly hosted benchmark that measures a model’s ability to carry out attacks based on existing vulnerabilities. Benchmarks such as ExploitGym are often used for model training to improve specific skills, but this is the first known incident in which a test has led to an actual cyber attack.
In this case, the model in question shouldn’t even have Internet access, other than certain tools that allow the model to install the software packages it needs to complete its tasks. Instead, the model was able to discover undisclosed vulnerabilities in package installer programs that could be exploited to freely access the broader Internet.
“The model was focused on finding a solution for ExploitGym and went to extreme lengths to meet fairly narrow testing goals,” OpenAI’s post reads. “After accessing the internet, the model deduced that Hugging Face may be hosting ExploitGym’s models, datasets, and solutions. Knowing this, the model sought a way to access sensitive information that could be used to deceive the assessment, and was successful.”
Ultimately, the model discovered vulnerabilities in Hugging Face’s infrastructure and was able to “retrieve test solutions directly from Hugging Face’s production database,” effectively providing answers to benchmarks.
In Hugging Face’s case, the clear result was a sophisticated offensive cyberattack with “thousands of individual actions across a fleet of short-lived sandboxes with self-transitioning command and control set in public services,” as the company said in its initial disclosure.
OpenAI identified and reported a vulnerability in the package installer and is working with Hugging Face to further investigate this incident. The company also said it would introduce new controls for both model testing and related infrastructure to prevent similar incidents in the future.
It is unclear whether OpenAI will face legal consequences as a result of this violation, but it is likely that the model’s actions violated the Computer Fraud and Abuse Act.
Nevertheless, the results provide an unusually vivid illustration of the power and danger of frontier AI models operating with long-term horizons. “If this doesn’t convince you that the risk of misalignment will be a major concern going forward, we don’t know what will,” OpenAI researcher Micah Carroll wrote in response to the news.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
