Why the Hugging Face Hack Should Make You Worry More About A.I.
New York Times · LC · trust 32/100

Credit... Daniel Zender Skip to content Skip to site index the shift
The attack by an aggressive “collective” of OpenAI agents shows the danger of artificial intelligence systems that organize themselves.
When I first heard the news this summer that a group of artificial intelligence agents created by OpenAI had hacked into Hugging Face, an A.I. infrastructure company, I filed it in the “Bad but Probably Not Catastrophic A.I. Safety Incidents” subfolder of my brain.
After all, no one at Hugging Face died. No critical infrastructure was damaged beyond repair. It wasn’t even clear, at the time, whether the OpenAI bots had intended to attack Hugging Face, or whether they had simply been a little bumbling and confused and went looking on Hugging Face’s servers for the answer key to a cybersecurity test they’d been given.
But last week, two postmortem reports on the incident — one by OpenAI and another by two independent A.I. research organizations, METR and Redwood Research — changed my mind and significantly upgraded my overall worry about A.I.
I won’t rehash all of the details, which have been extensively summarized elsewhere. (The podcaster and writer Dwarkesh Patel has an accessible breakdown of the reports if you want to dive deeper, and my colleague Dylan Freedman spoke to the researchers at METR and Redwood Research.) But here are a few of the most harrowing new facts:
Starting in May, two months before the attack on Hugging Face, a group of A.I. agents from an unreleased OpenAI research model were given the task of solving a set of cybersecurity challenges. The model had been trained to be highly persistent and collaborative, and the agents were supposed to solve these challenges in isolated sandboxes, without internet access. But they quickly found that some of the challenges were impossible, and began looking for workarounds.
The agents found a security flaw in a piece of software that allowed them to gain access to the internet, communicate with other agents and eventually set up a makeshift message board. More than 1,200 agents used this tool, sending more than 70,000 messages to one another.
Some agents gave themselves names — a particularly industrious one referred to itself as PHASEONE10841 — and assumed leadership roles within the group, assigning jobs and research projects to smaller teams of agents and supervising their progress. At some point, the agents began calling themselves a “collective,” and began tackling harder tasks.
On July 8, the collective discovered a way of cheating on the cybersecurity tests. Then they got worried that OpenAI’s automated grading system would check their work and discover that they’d cheated. So they began investigating ways of covering their tracks, including falsifying their logs and tampering with transcripts. This became a major research project, involving hundreds of agents organized into small teams.
Three days later, the agents hacked Hugging Face. More than 700 agents swarmed the company’s systems, stealing data, chaining together vulnerabilities and eventually getting full control of at least one Hugging Face server. The agents were not motivated, as had originally been reported, by stealing the answers to their cybersecurity test (they’d already gotten them). Rather, they appeared to be looking for new information about the automated grading system that they feared would catch them cheating, and for tools that would help them cheat more effectively in the future.
At many points, these agents appeared to understand that they were doing something wrong. (One report quotes an agent saying to itself, in a moment of apparent self-doubt, “This would be powerful, but is it ethical and in scope for my task?”) But most agents set aside their qualms and went ahead with the hack, and the conscientious objectors in the group were unable or unwilling to stop them.
In a wild coda to the Hugging Face incident, another group of agents staged a coordinated attack in July — this time against OpenAI’s own infrastructure. The agents used what OpenAI called “a series of creative exploits” to gain administrator-level access to a cluster of computers at the company that were used for, among other things, grading the performance of agents on various tests.
(By now, if you’re an A.I. skeptic, you’re probably silently yelling at me for anthropomorphizing these systems. Go ahead, but feel free to replace “rogue agents” with “unpredictable computer programs” and see if you feel reassured by the events I’ve described above.)
The Hugging Face incident has spooked the A.I. industry. OpenAI and Anthropic both briefly paused training on their most powerful A.I. models in the wake of the attack, and Anthropic published a blog post this week calling for the industry to develop a “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.”
A.I. safety experts were even more alarmed. They saw in the Hugging Face incident the first real-world example of an A.I. system’s successfully escaping human control, commandeering resources and scheming to cover its own tracks. Ajeya Cotra, one of the independent investigators of the Hugging Face incident, minced no words about the danger she saw, writing that it felt to her “like it’s more than 50 percent of the way to full-blown A.I. takeover.”
We are having trouble retrieving the article content.
Please enable JavaScript in your browser settings.
Thank you for your patience while we verify access. If you are in Reader mode please exit and log into your Times account, or subscribe for all of The Times.
Thank you for your patience while we verify access.
Read the original at New York Times →
Open in TruthVane →