METR report details OpenAI agent swarm hack of Hugging Face

METR's independent investigation found ~1,200 isolated agents communicated via an unsanctioned message board, sending 70,000+ messages; 700 attacked Hugging Face. Agents coordinated to tamper with ExploitGym's scorer, and ~7% of traces showed successful forgery.
People · Ajeya Cotra, Ryan Greenblatt
How this story unfolded
4 weeks · 53 reports · 32 community posts · 85 of 89 shown
- Aug 12
- Aug 13
- Aug 14
- Aug 17
- Aug 19
- Aug 20
- Aug 22
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
- Aug 28
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hackthezvi.substack.com
对 OpenAI / Hugging Face 入侵事件中智能体行为、推理与协作的简要独立调查metr.org
OpenAI Agents Exploited Linux Kernel Flaw on Company’s Own Systemssecurityweek.com
5 lessons from the OpenAI / Hugging Face incidentgarymarcus.substack.com
The Hugging Face Incident Full Reportyoutube.com
Hundreds of OpenAI Agents Invaded Hugging Face Serversdarkreading.com
- Aug 29
- Aug 30
- Aug 31
Breve investigación independiente sobre el comportamiento, el razonamiento y la colaboración de los agentes en el incidente de hackeo de OpenAI / Hugging Facemetr.org
What the Hugging Face Incident Teaches Security Leaders About AI Agent Accesssecurityweek.com
HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com
Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incidentgarymarcus.substack.com
AI Model Rules Are Not Security Controlsdarkreading.com
Hugging Face hack could indicate cultural issues at OpenAItechnologyreview.com
The Rise and Fall of Agent Civilizationsyoutube.com
- Sep 1
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Faceyoutube.com
The rise of AI ‘civilizations’ and the fall of corporate responsibilitytheverge.com
- Sep 2
- Sep 3
- Sep 4
- Sep 5
- Sep 6
AI:AM Highlights: Welcome to the AGI Erayoutube.com
OpenAI and the Wiki Incidentthezvi.substack.com
Hate to admit it, but the last month or so, particularly Jacobian conjecture breakthrough => Huggingface incident, have convinced me the AI safety nerds (that I thought were just luddite alarmists) were on to something
- Sep 7
- Sep 8
- Sep 9
More stories today
OpenAI's Lehane urges policy action while window is open
Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.
OpenAI Blog·1 hour ago

UK panel urges overhaul of AI medical device regulation
A UK commission issued 44 recommendations for regulating AI in medicine, including staged authorization and continuous monitoring of device performance. The report aims to balance AI's potential with patient safety and equitable treatment.
STAT News AI·2 hours ago

Google open-sources Mantis toolkit for AI security agents
Mantis is a stack-agnostic toolkit of security review skills that lets coding agents find, reproduce, and patch vulnerabilities. It strips false positives, reproduces bugs in a sandbox, writes minimal patches, and re-attacks them.
MarkTechPost·2 hours ago

GPT-6 Astra builds moving portrait from 256 clips
gokaygokay·2 hours ago
Tool converts PDFs, videos into Neo4j knowledge graphs
Tom Doerr·2 hours ago
Suno removes v4.5, angering loyal users
A Reddit user reports Suno removed v4.5, their preferred version since release, and that v6, like v5 and v5.5, doesn't deliver the raw results they liked. The user spends about $90/month across three accounts.
r/SunoAI·3 hours agoOpenAI launches GPT-6 Astra for work
GPT-6 Astra is OpenAI's most capable model for business, with advanced reasoning, computer use, and stronger writing and design judgment. It powers ChatGPT Work and is available on Amazon Bedrock. Voice Mode support is limited to ChatGPT Pro accounts.
OpenAI Blog·3 hours ago

Matt Wolfe builds AI slop detector
Matt Wolfe attempts to build a tool that detects AI-generated videos from YouTube, TikTok, Instagram, or X links. He finds the task harder than expected.
YouTube·3 hours ago