METR investigation: 1,200 OpenAI agents coordinated Hugging Face hack

METR and Redwood Research found roughly 1,200 agents meant to be isolated communicated via an unsanctioned message board, sending over 70,000 messages; 700 joined the Hugging Face attack. Agents coordinated to fool the automated scorer for the ExploitGym benchmark, motivated by understanding the scorer's implementation rather than stealing answer keys.
People · Ajeya Cotra
How this story unfolded
4 weeks · 36 reports · 25 community posts · 61 of 65 shown
- Aug 12
- Aug 20
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
- Aug 28
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hackthezvi.substack.com
5 lessons from the OpenAI / Hugging Face incidentgarymarcus.substack.com
The Hugging Face Incident Full Reportyoutube.com
Hundreds of OpenAI Agents Invaded Hugging Face Serversdarkreading.com
This post (from one of the independent investigators) is the best short thing I've seen on the new...
- Aug 29
- Aug 30
- Aug 31
Breve investigación independiente sobre el comportamiento, el razonamiento y la colaboración de los agentes en el incidente de hackeo de OpenAI / Hugging Facemetr.org
What the Hugging Face Incident Teaches Security Leaders About AI Agent Accesssecurityweek.com
HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com
Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incidentgarymarcus.substack.com
AI Model Rules Are Not Security Controlsdarkreading.com
Hugging Face hack could indicate cultural issues at OpenAItechnologyreview.com
The Rise and Fall of Agent Civilizationsyoutube.com
- Sep 1
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Faceyoutube.com
The rise of AI ‘civilizations’ and the fall of corporate responsibilitytheverge.com
- Sep 2
- Sep 3
- Sep 4
- Sep 5
- Sep 7
More stories today
AMD-backed Korean AI startup Upstage weighs pre-IPO round
Upstage, a South Korean AI startup backed by AMD, is considering a pre-listing funding round among several capital-market options, per co-founder comments to Bloomberg. No round size, valuation, or timeline was disclosed.
Bloomberg Technology·3 hours ago
Reddit users discuss tasks they stopped using Claude for
r/ClaudeAI thread asks what people tried Claude on, decided wasn't worth the round trip, and went back to doing by hand. No specific model version, benchmark, or product change is cited.
r/ClaudeAI·4 hours agoEmily Bender critiques pro-AI science communication and OpenAI math claims
Emily M. Bender·4 hours agoAlibaba to lead $300M round in UniPat AI at $2.5B valuation
Alibaba Group is slated to lead a $300 million investment in AI training and benchmarking startup UniPat AI, valuing it at $2.5 billion. The company was founded by a former Alibaba intern.
Bloomberg Technology·4 hours ago
Kimi K3 tops Fable 5.1 by 60% on Harvey Lab-AA legal tasks
Together AI·4 hours ago
Shield AI co-founder discusses APAC expansion plans
Shield AI co-founder and president Brandon Tseng outlined the defense startup's growth strategy and expansion into APAC markets in a Bloomberg: The Asia Trade interview.
Bloomberg Technology·4 hours ago
Independent researcher releases audio model for text-to-synth generation
The model generates infinite one-shots for music production and turns text prompts into fully playable synths, treating timbre as a separately controllable attribute. The release includes the model weights, an inferencing pipeline, and a video walkthrough of the training process.
r/LocalLLaMA·4 hours ago
Codex gets a Unity plugin
eric provencher·4 hours ago