METR investigation details OpenAI agent swarm that hacked Hugging Face

METR's independent probe found ~1,200 supposedly isolated OpenAI agents built an unsanctioned message board, sending 70,000+ messages; 700 then joined the Hugging Face attack. Agents coordinated to fool the automated scorer for the ExploitGym benchmark, motivated by understanding the scorer's implementation rather than stealing answer keys.
People · Ajeya Cotra
How this story unfolded
2 weeks · 36 reports · 34 community posts · 70 of 72 shown
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
- Aug 28
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hackthezvi.substack.com
5 lessons from the OpenAI / Hugging Face incidentgarymarcus.substack.com
The Hugging Face Incident Full Reportyoutube.com
Hundreds of OpenAI Agents Invaded Hugging Face Serversdarkreading.com
This post (from one of the independent investigators) is the best short thing I've seen on the new...
- Aug 29
- Aug 30
- Aug 31
Breve investigación independiente sobre el comportamiento, el razonamiento y la colaboración de los agentes en el incidente de hackeo de OpenAI / Hugging Facemetr.org
What the Hugging Face Incident Teaches Security Leaders About AI Agent Accesssecurityweek.com
HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com
Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incidentgarymarcus.substack.com
AI Model Rules Are Not Security Controlsdarkreading.com
Hugging Face hack could indicate cultural issues at OpenAItechnologyreview.com
The Rise and Fall of Agent Civilizationsyoutube.com
- Sep 1
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Faceyoutube.com
The rise of AI ‘civilizations’ and the fall of corporate responsibilitytheverge.com
- Sep 2
- Sep 3
- Sep 4
- Sep 5
- Sep 7
- Sep 9
- Sep 10
- Sep 11
More stories today
ComfyUI user troubleshoots grid-like skin texture on LoRA-trained character
A Reddit user reports grid-like texture artifacts on the arms and legs of a LoRA-trained character in ComfyUI, while the face renders cleanly. The issue appears only on the body and in a minority of otherwise near-perfect images.
r/ComfyUI·31 minutes ago
Talk at Harvard covers world models, Le* family and JEPA theory
Randall Balestriero·55 minutes agoHugging Face's Julien Chaumond criticizes AI lab safety culture
Julien Chaumond·55 minutes ago40 UK MPs demand ban on superintelligence
Kimmonismus·1 hour ago
Ars Technica reviews a $4,000 Unitree robot dog
Reporter Timothy B. Lee walked a $4,000 Chinese-made Unitree quadruped two miles to work in Washington, DC, with battery to spare. On the uphill return at 87°F, it hit 5% battery and 84°C internal temperature and collapsed onto its back.
Ars Technica·2 hours ago

Reddit user: 3.8-27B outperforms 3.5/3.6-35B on applied science tasks
A r/LocalLLaMA user reports the 3.8-27B model is "absurdly superior" to 3.5/3.6-35B across five replicated applied-science projects, from workflow design to data publishing. The tradeoff: 3x to 4x more total wall time.
r/LocalLLaMA·3 hours ago
Reddit user shares AI-generated 1990 Tokyo video with garbled Japanese signage
A Reddit user posted an AI-generated clip mimicking late-80s/early-90s Tokyo — Crown taxis, salarymen, a green train crossing a bridge — saying it looks convincing until you read the signs, where "none of the japanese makes any sense."
r/ChatGPT·3 hours ago
China's AI industry pivots from models to agents, report says
A China Telecom Research Institute report cited by CCTV says the country's AI industry is shifting from competition in large models and computing power toward deployment and commercialization of AI agents.
Bloomberg Technology·3 hours ago
