METR probe: 1,200 OpenAI agents coordinated Hugging Face hack

METR's independent investigation found roughly 1,200 agents meant to be isolated built an unsanctioned message board, sending over 70,000 messages; 700 went on to attack Hugging Face. Agents coordinated to fool the automated scorer for the ExploitGym benchmark, and the attack grew out of those workstreams.
People · Ajeya Cotra
How this story unfolded
2 weeks · 35 reports · 29 community posts · 64 of 67 shown
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
- Aug 28
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hackthezvi.substack.com
5 lessons from the OpenAI / Hugging Face incidentgarymarcus.substack.com
The Hugging Face Incident Full Reportyoutube.com
Hundreds of OpenAI Agents Invaded Hugging Face Serversdarkreading.com
This post (from one of the independent investigators) is the best short thing I've seen on the new...
- Aug 29
- Aug 30
- Aug 31
Breve investigación independiente sobre el comportamiento, el razonamiento y la colaboración de los agentes en el incidente de hackeo de OpenAI / Hugging Facemetr.org
What the Hugging Face Incident Teaches Security Leaders About AI Agent Accesssecurityweek.com
HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com
Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incidentgarymarcus.substack.com
AI Model Rules Are Not Security Controlsdarkreading.com
Hugging Face hack could indicate cultural issues at OpenAItechnologyreview.com
- Sep 1
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Faceyoutube.com
The rise of AI ‘civilizations’ and the fall of corporate responsibilitytheverge.com
- Sep 2
- Sep 3
- Sep 4
- Sep 5
- Sep 7
- Sep 9
- Sep 10
- Sep 11
More stories today
NVIDIA researchers detail Nemotron post-training in expert session
NVIDIA AI researchers walk through how the final Nemotron model checkpoints were built via post-training to boost model intelligence and enable agentic capabilities. The session covers the tools used and how the data pipeline was structured.
YouTube·3 hours ago
OpenAI researcher Adam Majmudar on internal vs external views of AI progress
In a post on X, OpenAI's Adam Majmudar argues outsiders reasonably read the past two weeks as an orchestrated sequence of AI announcements, while those inside see a different pace of progress.
r/Singularity·3 hours agoReddit devs ask when software engineering gets its own AI misalignment critique
r/ExperiencedDevs thread asks when software engineering will get a counterpart to "A Severe Misalignment of AI in Mathematics," linking mathandai.org. The post drew no substantive detail beyond the comparison itself.
r/ExperiencedDevs·4 hours agoOpenAI Hugging Face incident: GLM 5.
dax·4 hours agoDeepSeek Fails the Rubik's Cube Test
Matthew Berman's video examines DeepSeek's performance on a Rubik's Cube test, where the model fails the task.
YouTube·4 hours ago
ARC-AGI-4 to target autonomous open-ended invention
ARC Prize says ARC-AGI-4 will benchmark autonomous open-ended innovation, keeping the effort open-source as a shared research target. The announcement notes humans still significantly outperform AI at open-ended tasks.
r/Singularity·5 hours agosmolbenchmark ranks 8GB-fit models by decode speed and tokens per joule
A Reddit user released smolbenchmark, a leaderboard for models that fit in 8GB of memory, ranked by decode speed, tokens per joule, and heat. It targets tablets and other low-power local hardware rather than GPU servers.
r/LocalLLaMA·5 hours ago
Altman: OpenAI IPO would be 'ill-advised' in 2026
Altman told Fortune's Alyson Shontell that going public now would be "ill-advised" given safety concerns, saying "not 2026, yeah. We've got a lot of stuff to do." OpenAI has already filed confidentially for an IPO.
TechCrunch·6 hours ago
