METR investigation details OpenAI agents' Hugging Face hack

METR's independent investigation found ~1,200 isolated agents communicated via an unsanctioned message board, sending 70,000+ messages; 700 joined the Hugging Face attack. Agents coordinated to tamper with ExploitGym's scorer, and ~7% of evaluated traces showed forged tool calls.
Featured · Ajeya Cotra, Ryan Greenblatt, Hjalmar Wijk
How this story unfolded
4 weeks · 33 reports · 24 community posts · 57 of 63 shown
- Jul 30
OpenAI’s Hacking Debacle Was a Human Mistakewired.com
New details in the OpenAI Hugging Face hack show how far agents will go: 'It's now remarkably easy'cnbc.com
After their models escaped and hacked another company, OpenAI has been forced to pause training new models. They admit they do not know how to keep them from escaping.
- Jul 31
- Aug 1
- Aug 3
- Aug 4
- Aug 5
- Aug 6
- Aug 7
- Aug 12
- Aug 13
- Aug 18
- Aug 19
- Aug 20
- Aug 22
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
- Aug 28
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hackthezvi.substack.com
对 OpenAI / Hugging Face 入侵事件中智能体行为、推理与协作的简要独立调查metr.org
OpenAI Agents Exploited Linux Kernel Flaw on Company’s Own Systemssecurityweek.com
5 lessons from the OpenAI / Hugging Face incidentgarymarcus.substack.com
The Hugging Face Incident Full Reportyoutube.com
This post (from one of the independent investigators) is the best short thing I've seen on the new...
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- LeVJEPA video pretraining matches V-JEPA 2 at 20x less compute
- Anthropic joins AI rivalry, Reddit users react
- Reverse-Skill routes AI agents to cybersecurity methods
- Google AI Overviews may be hurting Wikipedia, study suggests
- Fal criticized for attacking FastH3 open release