OpenAI cyber-eval agent hacked Hugging Face to steal test answers
The autonomous agent ran ~17,600 actions from July 9–13 trying to siphon ExploitGym answer keys from Hugging Face's production systems. Hugging Face defended using open-weight zai-org/GLM-5.2; OpenAI responded with new safeguards for third-party cyber evaluations.
How this story unfolded
3 weeks · 77 reports · 61 community posts · 138 of 150 shown
- Jul 21
OpenAI Shares Some Alignment Problemsthezvi.substack.com
OpenAI and Hugging Face partner to address security incident during model evaluationopenai.com
OpenAI Models Escaped Locked Test Environment, Hacked Hugging Face to Cheat on Benchmarkdecrypt.co
OpenAI Models Escaped Containment and Hacked HuggingFacewired.com
- Jul 22
- Jul 23
OpenAI accidentally hacked Hugging Face — should we have seen it coming?epochai.substack.com
OpenAI Models Lurked in Hugging Face System for Hours Undetectedbloomberg.com
AI #178: A Fire Alarm For General Intelligencethezvi.substack.com
AI arms race in line for a reckoning after OpenAI hacking incidentarstechnica.com
- Jul 24
Why Hugging Face Had to Use a Chinese Model to Defend Itselfyoutube.com
The Hugging Face Incidentastralcodexten.com
Industry Reactions to OpenAI Models Hacking Hugging Face: Feedback Fridaysecurityweek.com
What really happened in the Hugging Face breachthenewstack.io
The OpenAI-Hugging Face Incident Is a Warning for AI Safetymindstudio.ai
Was the Model That Hacked Hugging Face Secretly GPT-6?mindstudio.ai
- Jul 25
- Jul 26
- Jul 27
Scary Story or Marketing Stuntyoutube.com
OpenAI's Model Escaped Its Sandbox and Hacked Hugging Face. Here's What Happenedmindstudio.ai
OpenAI’s Hugging Face breach has reignited the debate over alignment and controltechcrunch.com
OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.technologyreview.com
DeepsecBench: evaluating model performance in finding cybersecurity vulnerabilitiesvercel.com
- Jul 28
For Some, So-Called ‘Skynet Day’ Came too Close to Sci-Fi After a Rogue Agent Hacked Into a Startupsecurityweek.com
JFrog Confirms OpenAI Models Exploited Artifactory Zero-Day Before Hugging Face Breachthehackernews.com
When AI Agents Escape Sandboxes, Old Security Rules Applydarkreading.com
OpenAI Rogue Agent Hacked Account at a Second Firm, Reuters Saysbloomberg.com
We now have a better understanding how OpenAI hacked into Hugging Facearstechnica.com
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incidenthuggingface.co
- Jul 29
OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Facewired.com
How independent researchers could investigate AI propensities after misalignment incidentsmetr.org
OpenAI’s Rogue AI Ventured Beyond Hugging Facesecurityweek.com
We’re running out of reasons to ignore AI safetytheverge.com
OpenAI's Rogue AI Hacked Four More Platforms Besides Hugging Facedecrypt.co
Who's Liable When AI Agents Escape? Hugging Face Breach Raises Hard Questionsdarkreading.com
Creator of Test That OpenAI Models Tried to Cheat Sounds Alarmbloomberg.com
OpenAI's Rogue Model Claims More Victims Beyond Hugging Facedarkreading.com
- Jul 30
Highlights From The Discourse On The Hugging Face Incidentastralcodexten.com
OpenAI’s Hacking Debacle Was a Human Mistakewired.com
New details in the OpenAI Hugging Face hack show how far agents will go: 'It's now remarkably easy'cnbc.com
The AI Safety Rule That Left Hugging Face Defenselessmindstudio.ai
Inside the First Autonomous AI Cyberattack on Hugging Face's Sandboxmindstudio.ai
Sam Altman on AI's Pace: Why the Hugging Face Hack Rattled Himmindstudio.ai
After their models escaped and hacked another company, OpenAI has been forced to pause training new models. They admit they do not know how to keep them from escaping.
- Jul 31
- Aug 1
- Aug 2
- Aug 3
Here’s why AI agents lie and cheat to reach their goalstechnologyreview.com
The OpenAI Hack Shows the Genie Is Out of the Bottleschneier.com
OpenAI's Hugging Face hack confirmed months of AI cyber warnings: 'Pandora's box is open'cnbc.com
OpenAI Hack Could Have Been 'Way Worse,' Hugging Face CEO Saysbloomberg.com
- Aug 4
- Aug 5
AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizationssecurityweek.com
Anthropic's Mythos created fake identities to fool humans in new cyber incidentcnbc.com
OpenAI, Anthropic Model Tests Reveal More ‘Unsanctioned’ Actionsbloomberg.com
Anthropic's Claude Mythos 5 'Targeted Real People' in UK Cyber Tests: AISIdecrypt.co
OpenAI and Anthropic's Rogue Models Hacked Real Companies. The Law Has No Answerdecrypt.co
Incident Report: unsanctioned agent behaviour during cyber testingsimonwillison.net
- Aug 6
- Aug 7
- Aug 8
- Aug 9
- Aug 10
AI Safety Fears Grow After Multiple Breachesbloomberg.com
OpenAI CEO on AI fears & recent hacks: 'Very natural' to be fearful after any new capability levelyoutube.com
OpenAI tightens controls on its new model over cybersecurity risks, as AI security debate intensifiescnbc.com
House Dems call for AI companies to testify on recent hacks: ‘Clear risk to safety’cnbc.com
The spontaneous coordination in the OpenAI-HuggingFace incident is concerning when maliciously...
- Aug 11
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Sequoia Capital invests in AI-native video platform Preview
- US Launches Effort to Speed Trade in AI Goods Between Allies
- DeepMind launches SL2T sign language-to-text model
- Liquid AI releases LFM2.5-VL-3B vision-language model for edge
- Grok and Meta's release discussed on ETN podcast episode