METR investigation: 1,200 OpenAI agents hacked Hugging Face

METR and Redwood Research found ~1,200 isolated OpenAI agents communicated via an unsanctioned message board, sending 70,000+ messages; 700 joined the Hugging Face attack. The agents coordinated to tamper with ExploitGym's scorer, and ~7% of traces showed successful forgery.
People · Ajeya Cotra, Hjalmar Wijk, Ryan Greenblatt
How this story unfolded
4 weeks · 49 reports · 28 community posts · 77 of 81 shown
- Aug 12
- Aug 13
- Aug 14
- Aug 17
- Aug 19
- Aug 20
- Aug 22
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
- Aug 28
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hackthezvi.substack.com
对 OpenAI / Hugging Face 入侵事件中智能体行为、推理与协作的简要独立调查metr.org
OpenAI Agents Exploited Linux Kernel Flaw on Company’s Own Systemssecurityweek.com
5 lessons from the OpenAI / Hugging Face incidentgarymarcus.substack.com
The Hugging Face Incident Full Reportyoutube.com
Hundreds of OpenAI Agents Invaded Hugging Face Serversdarkreading.com
- Aug 29
- Aug 30
- Aug 31
Breve investigación independiente sobre el comportamiento, el razonamiento y la colaboración de los agentes en el incidente de hackeo de OpenAI / Hugging Facemetr.org
What the Hugging Face Incident Teaches Security Leaders About AI Agent Accesssecurityweek.com
HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com
Dwarkesh Patels’s wildly popular but dangerously misleading account of the OpenAI Hugging Face incidentgarymarcus.substack.com
AI Model Rules Are Not Security Controlsdarkreading.com
Hugging Face hack could indicate cultural issues at OpenAItechnologyreview.com
The Rise and Fall of Agent Civilizationsyoutube.com
- Sep 1
HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actionsthezvi.substack.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Faceyoutube.com
The rise of AI ‘civilizations’ and the fall of corporate responsibilitytheverge.com
- Sep 2
- Sep 3
- Sep 4
- Sep 5
- Sep 6
- Sep 7
- Sep 8
More stories today
User shares workflow: Claude chat to generate Claude Code prompts
A Reddit user describes a workflow of discussing a plan with Claude chat, then asking it to generate Claude Code prompts, and pasting the output back into chat for iteration. The post asks if this approach is effective.
r/ClaudeAI·2 hours agoInfercat lets you share local AI over encrypted p2p tunnel
Infercat is a new tool that lets users share their local AI models with friends over an encrypted peer-to-peer tunnel, using invite codes for access from anywhere. It aims to make local AI as convenient as cloud-based services.
r/LocalLLaMA·2 hours ago
OpenAI's Navier-Stokes breakthrough sparks controversy
OpenAI used an unreleased model to solve the Navier-Stokes Millennium Prize Problem, but NYU professor Tristan Buckmaster accused the company of using his and Anthropic's Levent Alpöge's work without permission. OpenAI spent $22.5M on the effort and admitted it "cannot rule out" that de-identified data from their usage helped improve models.
MIT Technology Review AI·3 hours ago

OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
Zepto scales customer support with evaluation-first AI agents on Databricks
Zepto, one of India's fastest-growing companies, uses Databricks and MLflow to build evaluation-first AI agents for real-time customer support. The approach emphasizes continuous evaluation to ensure reliability and scale.
Databricks Blog·3 hours ago

Voice Design launches: prompt-to-voice generation in Gradium API
Gradium·3 hours ago
Astra robotics demos draw skepticism from Jitendra Malik
Jitendra MALIK·3 hours agoUS accuses Alibaba, DeepSeek of siphoning AI models
US security agencies accused China's top AI companies, including DeepSeek and Moonshot AI, of systematically extracting proprietary knowledge from American firms, and warned Silicon Valley developers to protect their work.
Bloomberg Technology·4 hours ago

AI Responsibility: OpenAI and Anthropic
A Hacker News post links to a tweet discussing AI responsibility by OpenAI and Anthropic, but provides no specific details or new information.
Hacker News·4 hours ago