OpenAI Hacking Fiasco Exposes a “Deeply Insufficient” System to Protect the Public
Mother Jones · L · trust 55/100

OpenAI CEO Sam Altman speaks to journalists after meeting with US House Minority Leader Hakeem Jeffries on Capitol Hill on June 3, 2026. Brendan Smialowski/AFP/Getty
The incident sounded straight out of a science fiction movie: OpenAI’s super-advanced tool hacked another AI company’s systems in an attempt to pass its own developers’ cybersecurity test. Just replace the AI tech with a newly engineered virus and you have an entire existing subgenre.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” OpenAI wrote in a Tuesday blog post explaining the incident. The tech giant said their tool, designed to execute tasks without any human assistance, independently breached Hugging Face, another startup that hosts a voluminous number of publicly available AI models, while it was testing internally how good it was at “advanced exploitation using complex attack paths” within a supposedly enclosed lab environment called a sandbox. In other words, OpenAI was testing its own hacking capabilities, and the brakes came off; the tool broke out and onto the open internet, and that’s when the mischief began.
OpenAI said they had the situation under control: Hugging Face detected the breach last week and stopped the activity on their own (and called the cops). Since then, OpenAI said it was working with Hugging Face on addressing vulnerabilities.
But the incident—along with many others my colleagues have reported about —brings up countless regulatory concerns as the industry, and the public at large, grapples with what actually went down at OpenAI and the safety of autonomous agents. (The Center for Investigative Reporting, the parent company of Mother Jones , has sued OpenAI for copyright violations . OpenAI has denied the allegations.)
To better understand what actually happened—and to what extent we should be worried—I spoke with Miranda Bogen , who works on developing and promoting AI governance that incorporates technical expertise as the chief technologist at the Center for Democracy & Technology and the founding director of its AI Governance Lab.
This interview has been condensed and edited for clarity.
What was your immediate reaction to hearing the news come out on Tuesday?
The rhetoric was very overblown. The headlines made it out that a model had run amok, and that it was a complete surprise, and that it was something people might be exposed to. But what was really happening was that OpenAI was testing a new version of a system made up of multiple of its models. It was specifically within a sandboxed environment, and they were basically trying to get it to demonstrate capabilities in executing cyber attacks. What ended up happening was that the system identified a vulnerability in a part of the sandbox setup, and it used that vulnerability to access the internet to find the answer key for the test, which led it to try to figure out if the answers to the test were on Hugging Face in a non-public setup.
That still demonstrates a quite advanced set of tasks that these systems are able to do now. But all of the safeguards had been removed for the purposes of this test, and it was doing the type of thing it was being tested to do—it just ended up finding a different channel to do that.
You mentioned some of the rhetoric being overblown. Are you referring to media coverage, what OpenAI and Hugging Face have said, or something else?
The real active debate is: “Is the government the right actor to decide that, especially when there are potential national security implications?”
I think the headline [of] “models escaped containment and hacked into a startup” is not necessarily wrong, but it makes it sound much more “run amok” than a particular contained experiment that was noted and caught. While I do think that the incident is something that’s important to look at and figure out what would be needed to prevent this from happening in the future, we can’t lose sight of the fact that there are very active conversations in AI policy going on about under what conditions models are permitted to be released and who needs to see them before they’re released.
The attention that companies get when they talk about very advanced capabilities gives them airtime with decision-makers. I can’t speak to the motivation that [OpenAI] had in how it characterized the incident, but that’s certainly something going on in the background.
OpenAI has product rollouts where there’s a version for regular users and a more advanced one for a set of trusted people. Could this specific OpenAI incident have escalated to a point where it was something actually concerning?
I think this incident was just a breach of a private part of the Hugging Face platform and didn’t lead to material harm beyond the fact that it was able to be breached. An incident like this could have had a material impact had the model been attempting to take actions that included taking down a website, exfiltrating information, or sending a huge volume of traffic to another party or platform. The fact that the model was operating in an external environment does suggest that it could have been possible that those actions could have had some consequences. In this case, they didn’t, but this was a test that was intended to be purely internal, as far as I can tell, among employees.
Even when there’s a trusted set of actors who are given access to advanced models, there are still some safeguards and monitoring, as far as I understand. In this case, I believe they were trying to understand the underlying capabilities of the model without those safeguards, because that informs how strong the safeguards need to be. I don’t think they are typically giving the rawest version of those models, even to the trusted parties, but I don’t have the details of exactly how those trusted party agreements are playing out.
We’re mostly basing our assessment on what Hugging Face and OpenAI have said publicly.…
Read the original at Mother Jones →
Open in TruthVane →