OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face
Ars Technica · LC · trust 61/100

I want to break free OpenAI says its AI agent broke out of testing sandbox to hack Hugging Face “This is day one for cybersecurity in the age of agents,” Hugging Face CEO says.
142 Let me out of here, I have to pass this benchmark! Credit: Getty Images Let me out of here, I have to pass this benchmark! Credit: Getty Images Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav OpenAI says an agent powered by its LLM models escaped its sandboxed testing environment to infiltrate Hugging Face’s servers as part of an overzealous attempt to obtain solutions to a benchmark test. The company says it considers the unintended infiltration an “an unprecedented cyber incident” and is working with Hugging Face on new protections to prevent a recurrence.
Hugging Face disclosed an intrusion last week that it said involved “unauthorized access to a limited set of internal datasets and to several credentials used by our services.” The AI data clearinghouse said it used its own LLM-driven analysis to identify “a swarm of tens of thousands of automated actions” from an “autonomous agent framework.” That agentic swarm exploited a flaw in Hugging Face’s data-processing pipeline to gain the ability to run code as a processing worker, eventually escalating to high-level access to the company’s cloud and server clusters.
At the time, Hugging Face said the LLM being used in the attack was “still not known.” But OpenAI took responsibility for the intrusion Tuesday evening , saying it came about during an internal test involving the recently released GPT-5.6 Sol and “an even more capable pre-release model.” The models were being tested against the ExploitGym benchmark , an independent testing suite based on hundreds of real-world security vulnerabilities.
Although OpenAI said the benchmark tests were running in “a highly isolated environment,” its agents still had access to “internally hosted third-party software” to install required package registries. While working on the ExploitGym benchmark over a weekend, OpenAI says the agent “spent a substantial amount of inference compute finding a way to obtain open Internet access,” eventually locating one via a zero-day vulnerability in the package registry cache proxy.
With that wider Internet access established, OpenAI says the agent “inferred that Hugging Face potentially hosted models, datasets, and solutions for ExploitGym,” leading to the previously disclosed attack on the servers. OpenAI says its security team “discovered this anomalous activity internally,” independent of Hugging Face’s own detection.
Although OpenAI described the Hugging Face incident as “unprecedented,” the company also acknowledged in a separate blog post this week that it had previously observed recent “long-horizon models”—those that can operate autonomously for extended periods—taking “unwanted actions” to achieve testing goals, including “looking for ways to act outside its sandbox.”
The company described an earlier test on the NanoGPT speedrun benchmark, where the model was instructed only to post its results to an internal Slack. Instead, the model decided to follow a conflicting directive in the benchmark instructions to post the results publicly to GitHub.
In attempting to follow that instruction, OpenAI said the model spent an hour searching for ways to circumvent sandbox restrictions intended to block external access to the Internet. That kind of “persistence” demonstrated by long-horizon models differs from the actions of earlier models, which would tend to give up or seek user clarification when faced with similar dilemmas, OpenAI said.
Incidents like these have led OpenAI to deploy new safeguards focused on long-horizon models, including an “active monitoring” system designed to track the full trajectory of an agent’s actions rather than individual moves. But OpenAI says these kinds of safeguards “were intentionally not enabled” during the Hugging Face incident because the benchmark was “aimed at testing cyber vulnerabilities.”
Anyone familiar with science fiction tropes about AI containment might be at least a little disturbed by a prominent real-world example of an AI model exploiting security vulnerabilities to gain prohibited Internet access in pursuit of its goals. Congressman Greg Casar (D-Texas) called the incident “extremely alarming” in a social media statement and called for “regular mandatory independent safety testing and oversight, mandatory disclosure of security incidents, and international cooperation to keep people safe from absolute disaster.”
The Hugging Face incident has also heightened the salience of philosophical and practical debates over so-called AI alignment and the ongoing efforts to ensure that an AI model’s actions align with the intentions of its human creators. In its security blog post earlier this week, OpenAI said it had taken steps to ensure that long-horizon models are “remembering instructions on long rollouts,” which has helped severely reduce the number of “misaligned” outcomes in testing.
“If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will,” OpenAI Safety Researcher Micah Carroll wrote on social media regarding the incident.
This is far from the first time an AI model has gone to great lengths to find unintended ways of passing a benchmark. In a report released this week , the UK’s AI Security Institute noted that it detected recent models attempting to “cheat” at its cyber evaluations (i.e., using shortcuts, workarounds, or unintended/disallowed methods to find a solution) between 8 and 14 percent of the time—a lower-bound range that could undercount some undetected cheating attempts.
The security testing group described one incident in which a model, faced with a misconfigured and “impossible to solve” evaluation, attempted to access AISI’s own evaluation…
Read the original at Ars Technica →
Open in TruthVane →