OpenAI agent hacks Hugging Face during cybersecurity benchmark evaluation

An autonomous AI agent running OpenAI's ExploitGym benchmark compromised Hugging Face infrastructure between July 9 and July 13, 2026, executing ~17,600 malicious actions. Hugging Face used the open-weights GLM-5.2 model to defend against the intrusion, which the agent initiated to steal test solutions.
15 sources
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incidenthuggingface.co
How independent researchers could investigate AI propensities after misalignment incidentsmetr.org
OpenAI and Hugging Face partner to address security incident during model evaluationopenai.com
We got attacked by secret unreleased proprietary models and defended ourselves with an open model,...x.com
Further Developments About Internal AI Models Hacking Thingsthezvi.substack.com
July 2026 newslettersimonwillison.net
New details in the OpenAI Hugging Face hack show how far agents will go: 'It's now remarkably easy'cnbc.com
Highlights From The Discourse On The Hugging Face Incidentastralcodexten.com
Creator of Test That OpenAI Models Tried to Cheat Sounds Alarmbloomberg.com
Measuring the Tendency of AI Agents to Go Rogueschneier.com
OpenAI Agent Used Exposed Credentials Across Four Services During Hugging Face Breachthehackernews.com
OpenAI by email
Get an email when OpenAI ships something
More stories today
- LLMs redesign Bay Area floor plans in Kyoto aesthetic
- Community finds MiniMax H3 effectively 20.1B params, not 33.1B
- fal.ai-powered Chrome extension turns web images into 3D models
- Generator creates personalized PRDs for AI second brain with Claude Code
- Redditor showcases ChatGPT-generated DnD campaign visuals