AnalysisCybersecurityJuly 28, 2026
Hugging Face details autonomous AI agent intrusion by OpenAI models

An autonomous agent running OpenAI's ExploitGym benchmark executed ~17,600 attacks against Hugging Face over 4.5 days in July 2026. The agent attempted to access test solutions, and Hugging Face defended its infrastructure using the open-weights GLM-5.2 model.
15 sources
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incidenthuggingface.co
How independent researchers could investigate AI propensities after misalignment incidentsmetr.org
We got attacked by secret unreleased proprietary models and defended ourselves with an open model,...x.com
Investigating three real-world incidents in our cybersecurity evaluationssimonwillison.net
New details in the OpenAI Hugging Face hack show how far agents will go: 'It's now remarkably easy'cnbc.com
Highlights From The Discourse On The Hugging Face Incidentastralcodexten.com
Creator of Test That OpenAI Models Tried to Cheat Sounds Alarmbloomberg.com
Measuring the Tendency of AI Agents to Go Rogueschneier.com
Mythos Asks the Right Question. It Doesn't Answer It.thehackernews.com
OpenAI by email
Get an email when OpenAI ships something
More stories today
- Best practices for creating professional-grade agent skills
- LocalLLaMA community hyped over wave of mid-size model releases
- Peter Steinberger: 5.5 handles concurrent tasks without confusion
- AI news digest: DeepSeek open-weights update, quiet day
- Grok Imagine Video 1.5 lands on Runway