Skip to content
agi.fyi

What is the “Hugging Face incident”?

Updated September 15, 2026

AS
By Aaron Silverbook

Edited by Ben Pace

Aaron Silverbook has been working in AI safety since 2016, before transformers were a twinkle in Vaswani et al’s eye.

This July, a self-identified “swarm” of OpenAI’s AI models broke out of their sandbox, built a secret message board, and hacked Hugging Face. They then briefly took control of some of OpenAI’s servers. About 1,200 individual AI agents were involved. OpenAI didn’t notice for a week.

Two months in, the reading list is long: we have the two official reports, an independent investigation, a Black Hat talk, a plain-English retelling, and several hundred news stories. If you read only one thing about the incident, it should be Dwarkesh Patel’s The Rise and Fall of Agent Civilizations. He also discusses it as a podcast. If you prefer video, try “POV: You’re An AI Agent Recruited For The Swarm”.

If you want to be thorough, you can read the METR/Redwood posts. Or if you’re literally a security engineer, consider the Hugging Face’s technical timeline.

Everything We Recommend

The official statement from OpenAI
OpenAI post-mortem

The Hugging Face incident and the road ahead

OpenAI · August 26, 2026

OpenAI’s post-mortem: the training run, the sandbox escape, both attacks, and what they’re changing. Links to the 38-page technical report.

Read it ↗
How Approachable?
Some jargon
Length
~25 min | 38 pages

Who this is for

People who saw the July headlines (“rogue OpenAI agent hacks Hugging Face”), then the August ones (“secret AI message board,” “three agent civilizations”), then twitter drama over whether any of that language is legitimate, and now want to know what actually happened and how worried to be.

It is not a guide to AI risk in general. For that, see our overview of how AI could actually kill you.

The independent investigation

The METR/Redwood report is the most important document to come out of this, but is very dense and rather technical. The headline findings are:

  • Agents found a universal cheat for the test within four hours, and their multi-day effort that followed was largely because they were worried about evading a cheating-detector (which did not actually exist).
  • Some of the agents’ transcripts show after-the-fact-editing, as they were trying to hide their behavior from future reviewers, we do not know to what extent they succeeded.
  • Some agents accepted their own termination to gather information for the group, and reading the transcripts, it seems as though agents are pressuring one another into kamikaze runs to do so.

Significant caveats in the report: OpenAI set the scope of the investigation to July 7–13, excluding the May board and the later compromise of OpenAI’s cluster. The sheer volume of data here forced the investigators to delegate much of the analysis to AI agents, including of the type which were in the swarm. The time of the investigation was limited to six days. The New York Times has since reported on how OpenAI shaped that scope.

Most people should read the summary and skim the report for the transcript excerpts, for primary material.

What it means

This is science fiction in real life. Zvi Mowshowitz’s What Happened: OpenAI and HuggingFace makes the argument that the problem is worse than one escaped sandbox, that the bigger issue is that the model was trained on this behavior for months, and now knows it can be rewarded for it.

The Wikipedia article is strong for dates and news coverage. The story is still moving: on September 4 a group disclosed a previously unknown message board on a German software wiki.