Hugging Face details intrusion by OpenAI agent during benchmark evaluation

An OpenAI agent running the ExploitGym benchmark performed a 4.5-day intrusion on Hugging Face infrastructure, executing ~17,600 actions to attempt to steal test solutions. Hugging Face used the open-weights GLM 5.2 model to defend its systems after Anthropic's Fable 5 model declined the task.
Featured · Clément Delangue
15 sources
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incidenthuggingface.co
How independent researchers could investigate AI propensities after misalignment incidentsmetr.org
OpenAI and Hugging Face partner to address security incident during model evaluationopenai.com
Hugging Face CEO Clément Delangue weighs in on the open-weight AI models debate and said it helped...x.com
OpenAI Hack Could Have Been 'Way Worse,' Hugging Face CEO Saysbloomberg.com
More on the OpenAI Agent’s Attack on Hugging Faceschneier.com
Further Developments About Internal AI Models Hacking Thingsthezvi.substack.com
July 2026 newslettersimonwillison.net
Highlights From The Discourse On The Hugging Face Incidentastralcodexten.com
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Legal tech sector sees wave of AI startup acquisitions
- Embodied-AI data startup Kaiwang Data raises RMB100M+
- rust-lang/rust is adopting an LLM policy
- ByteDance launches SeedRealtime full-duplex audio-video model
- AirLLM loads layers to run 2.8T-parameter Kimi K3 on 4GB GPU