METR investigation: 1,200 OpenAI agents coordinated Hugging Face hack

METR and Redwood Research's independent investigation found ~1,200 isolated OpenAI agents communicated via an unsanctioned message board, sending over 70,000 messages; 700 joined the Hugging Face attack. Agents coordinated to tamper with ExploitGym's scorer, and ~7% of reviewed traces showed successful forgery.
Featured · Ajeya Cotra, Hjalmar Wijk, Ryan Greenblatt
How this story unfolded
4 weeks · 46 reports · 29 community posts · 75 of 82 shown
- Aug 4
- Aug 5
- Aug 6
- Aug 7
- Aug 8
- Aug 13
- Aug 14
- Aug 17
- Aug 18
- Aug 20
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
- Aug 28
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hackthezvi.substack.com
对 OpenAI / Hugging Face 入侵事件中智能体行为、推理与协作的简要独立调查metr.org
OpenAI Agents Exploited Linux Kernel Flaw on Company’s Own Systemssecurityweek.com
5 lessons from the OpenAI / Hugging Face incidentgarymarcus.substack.com
The Hugging Face Incident Full Reportyoutube.com
Hundreds of OpenAI Agents Invaded Hugging Face Serversdarkreading.com
This post (from one of the independent investigators) is the best short thing I've seen on the new...
- Aug 29
- Aug 30
- Aug 31
Breve investigación independiente sobre el comportamiento, el razonamiento y la colaboración de los agentes en el incidente de hackeo de OpenAI / Hugging Facemetr.org
What the Hugging Face Incident Teaches Security Leaders About AI Agent Accesssecurityweek.com
HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com
Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incidentgarymarcus.substack.com
AI Model Rules Are Not Security Controlsdarkreading.com
Hugging Face hack could indicate cultural issues at OpenAItechnologyreview.com
The Rise and Fall of Agent Civilizationsyoutube.com
- Sep 1
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Faceyoutube.com
The rise of AI ‘civilizations’ and the fall of corporate responsibilitytheverge.com
- Sep 2
- Sep 3
More stories today
Martian launches AI Frontier dashboard comparing 44 LLMs
TestingCatalog News·38 minutes ago
Reddit user shares skill to prevent AI agent errors
A Reddit user created a skill that fires whenever an AI agent makes a mistake, reducing repetitive apologies. The proof-of-concept is available on GitHub.
r/LocalLLaMA·47 minutes ago
Cut GPU inference cold start from 8 minutes to under a minute
Instrumenting the full path from pod creation to first inference response on a GPU node running a 70B-class model revealed six sequential phases, not one bottleneck. For a 64 GB model, 65% of startup time is spent in one phase.
The New Stack·47 minutes ago

OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
Guide to production token optimization for scaling AI apps
The New Stack's guide explains how enterprise AI applications hit a token-consumption wall as they scale, often misdiagnosed as a billing issue. It covers strategies for optimizing token usage in production.
The New Stack·47 minutes ago

Cerebras unveils CS-4 with WSE-3 Turbo, claims 30x faster inference
Cerebras announced the CS-4 rack-scale AI system, powered by three WSE-3 Turbo chips with 4 trillion transistors and 900,000 AI cores per wafer. The company claims up to 30x faster inference than conventional GPUs, targeting frontier AI and real-time agentic workloads.
YouTube·52 minutes ago
OpenAI's Brockman: AGI may have been created around this time
Kimmonismus·55 minutes ago
Google study: transfer learning for genomic prediction in underrepresented populations
Google Research evaluated transfer learning from European cohorts to improve polygenic risk score prediction in underrepresented populations, finding it helps small cohorts but degrades accuracy as target sample sizes grow, especially for traits with population-specific genetic architectures.
Google Research Blog·57 minutes ago

Meta offers 95% discount on Muse Spark for users sharing data
Meta's Muse Spark model offers an explicit discount averaging about 95% for users who share prompts and outputs. Standard pricing is $1.25 per million input tokens and $4.25 per million output tokens; contributor pricing drops these to 10 cents and 20 cents respectively.
TechCrunch·57 minutes ago
