OpenAI’s rogue hacking incident was a warning shot. Will it be a wake-up call to finally create AI safety regulation?
Fortune · C · trust 63/100

OpenAI disclosed something terrifying on Tuesday. Its most advanced AI models escaped a controlled testing environment and autonomously hacked another company called Hugging Face, an open source AI model hosting platform.
The AI swarmed Hugging Face’s database, carrying out a multi-step plot of its own creation, intended to steal the answers to the evaluation test it was being assessed on by its maker, OpenAI. It executed “tens of thousands of automated actions” at rapid speed, according to the July 16 blog post in which Hugging Face first disclosed the incident. For years, AI safety researchers and policy analysts have been warning that incidents like this were coming and urged government officials to ensure AI labs had adequate controls in place to prevent them. But these predictions were often shrugged off as hypothetical or alarmist and failed to stir public or government action. Some AI security experts said they thought it would take a real world incident, a “Three Mile Island for AI,” to create enough public pressure to compel policymakers to act. The question now is whether this OpenAI-Hugging Face cyber attack is that alarm bell? “The Hugging Face x OpenAI hack should be a wake-up call to take loss of control seriously,” said Marius Hobbhan, CEO and Founder of Apollo Research, which conducts safety testing for a number of AI companies. “There was no human in the loop, it was not intended, and it caused real-world harm. We’ll soon have even more powerful agents and this is clear evidence that society currently doesn’t know how to build them fully safely.” Peter Wallich, an AI policy expert who formerly worked for the U.K. government’s AI Security Institute, said that AI safety researchers have been warning about misalignment—when an AI model autonomously chooses actions that its user doesn’t intend or desire—for years. “Until recently, it has been frequently dismissed as science-fiction,” he said. “I consider this a clear warning shot.” To be fair, messaging around AI safety has often been confusing. Some of the loudest warnings have come from AI companies themselves, leading many to accuse these businesses of engaging in a sophisticated and somewhat counterintuitive marketing strategy, since claims that their models were dangerous made them seem more powerful and capable of performing useful tasks too. “Our model is so powerful it hacked a company on its own”—is both an alarming admission and a subtle brag about the model’s technological capabilities. Most governments have so far balked at putting in place mandatory rules about what safeguards companies developing advanced AI systems need to build into their models or have in place internally to guard against losing control of AI agents. Nor are there clear rules on what safeguards governments themselves need to have in place as they increasingly give these agentic AI models access to sensitive military and intelligence systems.
AI safety researchers and policy experts said that the Hugging Face cyber attack could be the trigger that changes this equation. “Here in Washington, D.C. the people I have spoken to about this are already freaking out quite a bit,” Connor Leahy, an AI researcher who is now U.S. director of Control AI, a nonprofit dedicated to preventing existential risks from AI superintelligence, told Fortune .
Leahy noted that U.S. national security officials, including the head of the National Security Agency and the CIA director, had both voiced grave concerns about the cyber capabilities of the latest AI models following Anthropic’s debut of its Mythos AI model and that this OpenAI incident was likely to further reinforce their desire to put controls on the technology. That view was echoed by Seán Ó hÉigeartaigh, Professor of the Centre for the Future of Intelligence at the University of Cambridge. He said while the OpenAI-Hugging Face incident might not prompt regulation in isolation, “we’ve now had several things that have been wake-up moments for U.S. regulators in particular. I think Mythos was one example where a model demonstrated that it could find vulnerabilities in most of our digital infrastructure. I think that really alarmed policymakers, and then we have this happening only a short space of months afterwards.” He said there were now “enough data points that make it clear that the trend is going in the direction of more capable models that could plausibly cause serious harm in the real world.” Rep. Greg Casar, a Texas Democrat who has been vocal in his calls for AI regulation, became one of the first lawmakers to call for more robust federal AI regulation in the wake of the Hugging Face incident. Casar said on social media that he found the Hugging Face incident “extremely alarming.” “We need regular mandatory independent safety testing and oversight, mandatory disclosure of security incidents, and international cooperation to keep people safe from absolute disaster,” he said in a post on X.
The current Trump administration came into office intent on dismantling what little AI regulation the Biden administration had put in place. This included rescinding a 2023 Executive Order that mandated that frontier AI companies share safety testing information with the U.S. government. Trump technology policy officials said they wanted to accelerate U.S. AI innovation and take a hands-off approach to regulating the industry. Key Trump AI advisors were skeptical at best of AI safety concerns, especially when tied to calls for more regulation. David Sacks, Trump’s former AI czar, said that the leading AI labs were hoping to create complicated safety rules that only they would be able to comply with, making it harder for younger startups to challenge their market position. He accused AI company Anthropic of “running a sophisticated regulatory capture strategy based on fear-mongering.” This laissez faire approach began to shift markedly following Anthropic’s debut of its powerful Mythos model in April. Mythos’s…
Read the original at Fortune →
Open in TruthVane →