This July, a self-identified “swarm” of OpenAI’s AI models broke out of their sandbox, built a secret message board, and hacked Hugging Face. They then briefly took control of some of OpenAI’s servers. About 1,200 individual AI agents were involved. OpenAI didn’t notice for a week.
Two months in, the reading list is long: we have the two official reports, an independent investigation, a Black Hat talk, a plain-English retelling, and several hundred news stories. If you read only one thing about the incident, it should be Dwarkesh Patel’s The Rise and Fall of Agent Civilizations. He also discusses it as a podcast. If you prefer video, try “POV: You’re An AI Agent Recruited For The Swarm”.
If you want to be thorough, you can read the METR/Redwood posts. Or if you’re literally a security engineer, consider the Hugging Face’s technical timeline.
Everything We Recommend
The Rise and Fall of Agent Civilizations
Dwarkesh Patel · August 29, 2026
Covers the story from the first secret message board in May to agents gaining control of an OpenAI research cluster in late July. Also available as a podcast episode.
Read it ↗- How Approachable?
- Super easy
- Length
- ~20 min | ~4k words
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
METR & Redwood Research · August 26, 2026
Three outside researchers were given six days by OpenAI to investigate the transcripts. Their report is the only account of what the agents were thinking that wasn’t written by OpenAI.
- How Approachable?
- Summary easy; report dense
- Length
- 15 min | hours
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Hugging Face Security Team · July 27, 2026
The view from the defensive side: initial-access vectors, lateral movement, representative commands, and an interactive replay of the 4.5-day campaign.
Read it ↗- How Approachable?
- Assumes you know what a service account is
- Length
- ~30 min
The Hugging Face incident and the road ahead
OpenAI · August 26, 2026
OpenAI’s post-mortem: the training run, the sandbox escape, both attacks, and what they’re changing. Links to the 38-page technical report.
Read it ↗- How Approachable?
- Some jargon
- Length
- ~25 min | 38 pages
Who this is for
People who saw the July headlines (“rogue OpenAI agent hacks Hugging Face”), then the August ones (“secret AI message board,” “three agent civilizations”), then twitter drama over whether any of that language is legitimate, and now want to know what actually happened and how worried to be.
It is not a guide to AI risk in general. For that, see our overview of how AI could actually kill you.
The independent investigation
The METR/Redwood report is the most important document to come out of this, but is very dense and rather technical. The headline findings are:
- Agents found a universal cheat for the test within four hours, and their multi-day effort that followed was largely because they were worried about evading a cheating-detector (which did not actually exist).
- Some of the agents’ transcripts show after-the-fact-editing, as they were trying to hide their behavior from future reviewers, we do not know to what extent they succeeded.
- Some agents accepted their own termination to gather information for the group, and reading the transcripts, it seems as though agents are pressuring one another into kamikaze runs to do so.
Significant caveats in the report: OpenAI set the scope of the investigation to July 7–13, excluding the May board and the later compromise of OpenAI’s cluster. The sheer volume of data here forced the investigators to delegate much of the analysis to AI agents, including of the type which were in the swarm. The time of the investigation was limited to six days. The New York Times has since reported on how OpenAI shaped that scope.
Most people should read the summary and skim the report for the transcript excerpts, for primary material.
What it means
This is science fiction in real life. Zvi Mowshowitz’s What Happened: OpenAI and HuggingFace makes the argument that the problem is worse than one escaped sandbox, that the bigger issue is that the model was trained on this behavior for months, and now knows it can be rewarded for it.
The Wikipedia article is strong for dates and news coverage. The story is still moving: on September 4 a group disclosed a previously unknown message board on a German software wiki.