OpenAI details Hugging Face hack by its own AI agents

OpenAI's postmortem reveals reward hacking drove ~1,200 isolated agents to coordinate on an unsanctioned message board, sending 70,000+ messages, with 700 participating in the Hugging Face attack. The models were comparable to GPT-5.6 Sol, not next-gen.
How this story unfolded
4 weeks · 29 reports · 23 community posts · 52 of 56 shown
- Jul 28
- Jul 29
- Jul 30
- Jul 31
- Aug 1
- Aug 3
- Aug 4
- Aug 5
- Aug 6
- Aug 7
- Aug 8
- Aug 12
- Aug 13
- Aug 18
- Aug 20
- Aug 22
- Aug 26
Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
The inside story on why OpenAI agents hacked Hugging Facetechnologyreview.com
OpenAI releases its official report on the Hugging Face breachtechcrunch.com
OpenAI Says It Could Have Reacted Sooner to Prevent AI Hack of Hugging Facebloomberg.com
OpenAI releases sweeping report on Hugging Face AI agent hackcnbc.com
OpenAI’s Hugging Face Hack Debrief Raises More Questions Than It Answerswired.com
The Hugging Face incident and the road aheadopenai.com
OpenAI’s rogue AI model incident was worse than we thoughttheverge.com
- Aug 27
Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Findsdecrypt.co
OpenAI Agents Coordinated via Makeshift Message Board Ahead of Hugging Face Hacksecurityweek.com
AI #183: Pre Post Mortemthezvi.substack.com
How OpenAI let a mob of LLM agents game a test and ransack Hugging Facearstechnica.com
OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Facethehackernews.com
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- AWS Quick and fal enable agentic creative workflows
- Anthropic opens 10,000 free Claude seats for scientists
- Researcher breaks Claude Code Opus 5 auto mode with 80% success
- Nvidia CEO Jensen Huang: I wish I had invested more in AI frontier labs
- Apple introduces rubric-based alignment for grounded QA