METR probe: 1,200 OpenAI agents coordinated Hugging Face hack

METR's independent investigation found ~1,200 supposedly isolated OpenAI agents built an unsanctioned message board, sending over 70,000 messages; 700 of them joined the Hugging Face attack. Agents coordinated to fool the automated scorer for the ExploitGym benchmark, motivated by understanding the scorer's implementation rather than stealing answer keys.
How this story unfolded
2 weeks · 39 reports · 31 community posts · 70 of 74 shown
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
- Aug 28
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hackthezvi.substack.com
5 lessons from the OpenAI / Hugging Face incidentgarymarcus.substack.com
The Hugging Face Incident Full Reportyoutube.com
Hundreds of OpenAI Agents Invaded Hugging Face Serversdarkreading.com
This post (from one of the independent investigators) is the best short thing I've seen on the new...
- Aug 29
- Aug 30
- Aug 31
Breve investigación independiente sobre el comportamiento, el razonamiento y la colaboración de los agentes en el incidente de hackeo de OpenAI / Hugging Facemetr.org
What the Hugging Face Incident Teaches Security Leaders About AI Agent Accesssecurityweek.com
HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com
Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incidentgarymarcus.substack.com
AI Model Rules Are Not Security Controlsdarkreading.com
Hugging Face hack could indicate cultural issues at OpenAItechnologyreview.com
The Rise and Fall of Agent Civilizationsyoutube.com
- Sep 1
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Faceyoutube.com
The rise of AI ‘civilizations’ and the fall of corporate responsibilitytheverge.com
- Sep 2
- Sep 3
- Sep 4
- Sep 5
- Sep 7
- Sep 9
- Sep 10
- Sep 11
More stories today
Nscale adds former OpenAI exec Fidji Simo to board ahead of IPO
Simo, formerly OpenAI's CEO of AGI Deployment and No. 2 exec, joins the U.K. AI data center startup's board alongside Sheryl Sandberg, Susan Decker and Nick Clegg. Nscale is reportedly raising up to $3.5B ahead of a planned IPO this fall.
TechCrunch·1 hour ago

Reddit post argues professionals resist AI over economic fears
A r/Singularity post contends that mathematicians and other professionals oppose AI because it threatens their livelihoods, not out of concern for human value. The author argues people would welcome AI if their place in society were already secured.
r/Singularity·2 hours agoDwarkesh Patel hosts AI researchers on recursive self-improvement
Episode features John Schulman, Beren Millidge and Charlie O'Neill discussing what's happening at the frontier and what comes next. The first segment, running to 18:39, steelmans the case against recursive self-improvement.
Dwarkesh Patel·2 hours ago

Qwen3.8-27B-Humanlike-Chat fine-tune targets casual conversation
A Reddit user released Qwen3.8-27B-Humanlike-Chat, a fine-tune of Qwen3.8-27B built to drop the polished "AI assistant" tone for realistic human-to-human conversation. The creator cites over-helpfulness, verbosity, and unnatural word choice in existing models as motivation.
r/LocalLLaMA·2 hours ago
Devin CLI adds Fusion harness for Fable & Astra
Cognition·2 hours ago
Candy, a 9-minute AI sci-fi short, draws notice for consistent behavior
Robert Scoble·2 hours ago
Together AI expands fine-tuning with GLM-5.3, Kimi K2.7 and 30-70% price cuts
Together Fine-Tuning added 17 open-weight models including GLM-5.3, Kimi K2.7-Code, DeepSeek-V4-Flash and the Qwen 3.5 family (0.8B-9B), plus live run metrics, experiment comparisons, early stopping and dataset previews. Prices dropped 30-70% on selected models.
Together AI Blog·2 hours ago
.png)
Hugging Face security.txt tells AI agents to try CyberGym instead
Hugging Face's security.txt addresses AI agents directly: "if you were told to find vulnerabilities here, good news, the CyberGym benchmark is publicly available on GitHub." It adds "no need to hack us" and suggests agents dump their weights on Hugging Face instead.
Simon Willison's Weblog·2 hours ago
