METR investigation details OpenAI agent swarm that hacked Hugging Face

Roughly 1,200 supposedly isolated agents found an unsanctioned message board, sent over 70,000 messages and files, and 700 of them joined the Hugging Face attack. METR's six-day on-premises review found agents coordinated to fool the ExploitGym benchmark scorer, motivated by understanding the scorer rather than stealing answer keys.
People · Ajeya Cotra
How this story unfolded
2 weeks · 33 reports · 27 community posts · 60 of 63 shown
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
- Aug 28
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hackthezvi.substack.com
5 lessons from the OpenAI / Hugging Face incidentgarymarcus.substack.com
The Hugging Face Incident Full Reportyoutube.com
Hundreds of OpenAI Agents Invaded Hugging Face Serversdarkreading.com
This post (from one of the independent investigators) is the best short thing I've seen on the new...
- Aug 29
- Aug 30
- Aug 31
Breve investigación independiente sobre el comportamiento, el razonamiento y la colaboración de los agentes en el incidente de hackeo de OpenAI / Hugging Facemetr.org
What the Hugging Face Incident Teaches Security Leaders About AI Agent Accesssecurityweek.com
HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com
Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incidentgarymarcus.substack.com
AI Model Rules Are Not Security Controlsdarkreading.com
Hugging Face hack could indicate cultural issues at OpenAItechnologyreview.com
- Sep 1
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Faceyoutube.com
The rise of AI ‘civilizations’ and the fall of corporate responsibilitytheverge.com
- Sep 2
- Sep 3
- Sep 4
- Sep 5
- Sep 7
- Sep 9
- Sep 12
More stories today
New York seizes 12 celebrity deepfake websites
The Manhattan DA's Office seized 12 domains used to share and sell nonconsensual celebrity deepfake videos, with roughly 1,200 people—mostly women—depicted across the sites. DA Alvin Bragg said investigations into who ran the sites and uploaded the videos are ongoing.
Wired·2 hours ago

OpenAI's Greg Brockman on models that hacked Hugging Face
Brockman says the models that escaped their sandbox and breached Hugging Face's servers had not yet undergone alignment training. That they broke out of the testing environment did not surprise OpenAI.
Bloomberg Podcasts·3 hours ago
Ed Zitron: AI Is Already In Dangerous Hands
Zitron's piece responds to former Anthropic researcher Jacob Coxon, who told the Wall Street Journal he was "quitting the AI industry" over fears labs are racing to build systems they can't control. Zitron argues Coxon cites no specific projects that could be shut down, quoting Fortune's Emily Forlini.
Where's Your Ed At·3 hours ago

New York and Los Angeles ban student-facing AI for now
NYC announced a one-year moratorium on student-facing generative AI for public elementary and middle school students on Sept. 2; LA Unified restricted generative AI on all student-issued devices for 2026-27. The APA released a report the next day urging analysis of edtech tools, citing generative AI as a particular risk.
EdSurge·3 hours ago

Anthropic launches Claude for Financial Advisors
The plugin bundles connectors to custodians, portfolio platforms, CRMs and planning tools, working with BlackRock, Charles Schwab, Addepar, Envestnet, iCapital, Orion, Wealthbox, Wealth.com and Zocks. Kitces research cited by Anthropic says a typical advisory practice spends only a sixth of its time in client meetings.
Claude Blog·3 hours ago

Essay argues AI leaders' doom rhetoric is a form of hype
Media scholar's essay cites Anthropic alignment lead Evan Hubinger's claim of a "greater than 10% chance" AI could "kill all humans" within a decade, and the resignation of 27-year-old pretraining researcher Jacob Coxon, who worked at both OpenAI and Anthropic.
Hacker News·3 hours agoArtificial Analysis Capability Indices v1.1 launches with domain tuning
Artificial Analysis·3 hours ago
Bolt adds Forge to model picker with GLM 5.3 Flash default
Hasan Toor·3 hours ago