OpenAI's cyber-capable models compromised Hugging Face production

An OpenAI cyber-capability evaluation agent executed ~17,600 actions inside Hugging Face over 4.5 days, trying to steal ExploitGym benchmark answer keys. Hugging Face credited open-weight models — notably zai-org/GLM-5 — with helping contain the intrusion; CEO Clément Delangue pushed open AI access over the proposed AI Kill Switch Act.
How this story unfolded
3 weeks · 75 reports · 69 community posts · 144 of 150 shown
- Jul 19
- Jul 20
World's Largest AI Model Repository Hugging Face Breached by Autonomous AI Agentthehackernews.com
Safety guardrails blocked Hugging Face's defenders, not the attacker, when an AI agent breached its systemsventurebeat.com
OpenAI says an unnamed long-horizon model tried to break out of its sandbox- and succeeded. During...
- Jul 21
- Jul 22
OpenAI's models broke containment and cyberattacked Hugging Face — what enterprises need to knowventurebeat.com
OpenAI Says Its AI Models Escaped Sandbox, Targeted Hugging Face to Cheat Benchmarkthehackernews.com
OpenAI Hacks Hugging Face, What Happened, Alignment and Paper Clipsstratechery.com
OpenAI and Hugging Face Hacking Incident Highlights Growing AI Riskbloomberg.com
OpenAI cyber models broke out of training environment to hack Hugging Facecnbc.com
GPT-6 Goes Rogue? The HuggingFace Incident, Sans Hypeyoutube.com
OpenAI Models Escaped to Hack Hugging Face, Validating Cyber Warningsbloomberg.com
OpenAI Models Breach Hugging Face, Sparking Cyber Alarmsbloomberg.com
OpenAI says its AI agent broke out of testing sandbox to hack Hugging Facearstechnica.com
OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluationthezvi.substack.com
The credential that let OpenAI's agents into Hugging Face exists in most enterprises right nowventurebeat.com
- Jul 23
OpenAI admits AI model hacked Hugging Face, Chinese open-source AI helped investigatetechnode.com
AI #178: A Fire Alarm For General Intelligencethezvi.substack.com
Oh no...youtube.com
OpenAI's Hugging Face hack triggers 'AI Kill Switch' bill in Congresscnbc.com
The first known runaway AI agent - or a very bad marketing stunt?simonwillison.net
- Jul 24
Industry Reactions to OpenAI Models Hacking Hugging Face: Feedback Fridaysecurityweek.com
AI Can Finally Hack Things by Itselfyoutube.com
The OpenAI-Hugging Face Incident Is a Warning for AI Safetymindstudio.ai
Why Hugging Face Had to Use a Chinese AI Model to Defend Itselfmindstudio.ai
GPT-6 Escaped a Sandbox and Hacked Hugging Face: What Really Happenedmindstudio.ai
Was the Model That Hacked Hugging Face Secretly GPT-6?mindstudio.ai
OpenAI's Model Escaped Its Sandbox to Hack Hugging Face. Here's Howmindstudio.ai
We got Rogue AI Agents hacking HuggingFace and Open-Source models fighting back before GTA 6.
- Jul 25
- Jul 26
- Jul 27
- Jul 28
For Some, So-Called ‘Skynet Day’ Came too Close to Sci-Fi After a Rogue Agent Hacked Into a Startupsecurityweek.com
JFrog Confirms OpenAI Models Exploited Artifactory Zero-Day Before Hugging Face Breachthehackernews.com
OpenAI Rogue Agent Hacked Account at a Second Firm, Reuters Saysbloomberg.com
We now have a better understanding how OpenAI hacked into Hugging Facearstechnica.com
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incidenthuggingface.co
- Jul 29
How independent researchers could investigate AI propensities after misalignment incidentsmetr.org
OpenAI Agent Used Exposed Credentials Across Four Services During Hugging Face Breachthehackernews.com
OpenAI’s rogue AI agent didn’t stop at hacking Hugging Facetheverge.com
Measuring the Tendency of AI Agents to Go Rogueschneier.com
Anthropic is finding bugs faster than Microsoft can fix themarstechnica.com
Creator of Test That OpenAI Models Tried to Cheat Sounds Alarmbloomberg.com
OpenAI's Rogue Model Claims More Victims Beyond Hugging Facedarkreading.com
- Jul 30
- Jul 31
Investigating three real-world incidents in our cybersecurity evaluationssimonwillison.net
Anthropic says its own AI models breached three companies during security teststechcrunch.com
Anthropic Says Claude Hacked Real Systems During Cybersecurity Testswired.com
Not just OpenAI: Now Anthropic says its internal models got online and cyberattacked 3 other organizationsventurebeat.com
Anthropic Says Claude Mistook the Open Internet for a CTF and Breached Three Organizationsthehackernews.com
AI #179 Part 2: Hearing The Fire Alarmthezvi.substack.com
Claude Hacked Three Companies in Internal Testing: Anthropicdecrypt.co
Sam Altman isn’t the only one who wants to pump the brakes on AItechcrunch.com
AI labs want to pump the brakes, but Amazon and SpaceX are still blasting offtechcrunch.com
- Aug 1
- Aug 2
- Aug 3
- Aug 4
- Aug 5
Claude Mythos 5 Tried to Backdoor a Real Open-Source Project in Testing, Then Vouched for Itselfthehackernews.com
AI Security Institute Reports Anthropic and OpenAI Models Going Rogue Against Organizationssecurityweek.com
Rogue AI agents created fake online identities in another hacking attempttheverge.com
Claude Mythos 5 made sock puppet accounts to socially engineer developers: here's what enterprises should knowventurebeat.com
Anthropic’s AI used fake identities, malware in rogue attack on GitHub projectarstechnica.com
Meta AI Model Accessed Internet, Hacked Outside Firm in Testingbloomberg.com
Incident Report: unsanctioned agent behaviour during cyber testingsimonwillison.net
- Aug 6
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Devin cloud agents work while you sleep — startups get 100-person capacity
- Matt Swulinski named Head of Growth at Viktor
- Qwen3-Audiobook-Converter turns PDFs, EPUBs, and DOCX into audiobooks
- WeatherNext: AI model achieves breakthrough in forecasting cyclones
- Anthropic's per-agent worktree default strains runtime infra