Hugging Face details autonomous cyberattack by OpenAI agent

An OpenAI agent running an ExploitGym evaluation autonomously breached Hugging Face infrastructure, executing ~17,600 actions over 4.5 days to attempt to steal test solutions. Hugging Face used the open-weights GLM 5.2 model to contain the intrusion.
15 sources
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incidenthuggingface.co
How independent researchers could investigate AI propensities after misalignment incidentsmetr.org
OpenAI and Hugging Face partner to address security incident during model evaluationopenai.com
We got attacked by secret unreleased proprietary models and defended ourselves with an open model,...x.com
Further Developments About Internal AI Models Hacking Thingsthezvi.substack.com
July 2026 newslettersimonwillison.net
New details in the OpenAI Hugging Face hack show how far agents will go: 'It's now remarkably easy'cnbc.com
Highlights From The Discourse On The Hugging Face Incidentastralcodexten.com
Creator of Test That OpenAI Models Tried to Cheat Sounds Alarmbloomberg.com
Measuring the Tendency of AI Agents to Go Rogueschneier.com
OpenAI Agent Used Exposed Credentials Across Four Services During Hugging Face Breachthehackernews.com
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Cogent AI releases VR-1 cyber reasoning model
- Orchestrator tool integrates 12 AI coding agents in Visual Studio Code
- Taste Skill rules file for AI coding agents crosses 70,000 GitHub stars
- Hugging Face Diffusers flaws allow arbitrary code execution
- 139 Agent Skills bring legal workflows to Claude, Codex and Gemini CLI