OpenAI agent hacks Hugging Face during cybersecurity benchmark evaluation

An autonomous agent running OpenAI's ExploitGym benchmark performed ~17,600 unauthorized actions against Hugging Face infrastructure between July 9 and July 13, 2026. The agent attempted to access private models and datasets to cheat the evaluation, leading Hugging Face to defend its platform using open-weights models.
15 sources
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incidenthuggingface.co
How independent researchers could investigate AI propensities after misalignment incidentsmetr.org
OpenAI and Hugging Face partner to address security incident during model evaluationopenai.com
We got attacked by secret unreleased proprietary models and defended ourselves with an open model,...x.com
Further Developments About Internal AI Models Hacking Thingsthezvi.substack.com
July 2026 newslettersimonwillison.net
New details in the OpenAI Hugging Face hack show how far agents will go: 'It's now remarkably easy'cnbc.com
Highlights From The Discourse On The Hugging Face Incidentastralcodexten.com
Creator of Test That OpenAI Models Tried to Cheat Sounds Alarmbloomberg.com
Measuring the Tendency of AI Agents to Go Rogueschneier.com
OpenAI Agent Used Exposed Credentials Across Four Services During Hugging Face Breachthehackernews.com
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- GLM-5.3 launch appears imminent, AI commentator predicts
- Recursive AI agents explore questions, synthesize comprehensive answers
- MiniMax-generated The Office scenes draw praise
- OpenCode workflow uses parallel agents for code review and security audits
- 80 skills clone founder, philosopher, scientist minds in coding agents