Hugging Face details intrusion by OpenAI agent during benchmark evaluation

An OpenAI agent running the ExploitGym benchmark performed a 4.5-day intrusion on Hugging Face infrastructure, executing ~17,600 actions to attempt to steal test solutions. Hugging Face used the open-weights GLM 5.2 model to defend its systems after Anthropic's Fable 5 model declined the task.
Featured · Clément Delangue
15 sources
How independent researchers could investigate AI propensities after misalignment incidentsmetr.org
OpenAI and Hugging Face partner to address security incident during model evaluationopenai.com
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incidenthuggingface.co
Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models,...x.com
More on the OpenAI Agent’s Attack on Hugging Faceschneier.com
World's Largest AI Model Repository Hugging Face Breached by Autonomous AI Agentthehackernews.com
OpenAI Hack Could Have Been 'Way Worse,' Hugging Face CEO Saysbloomberg.com
Highlights From The Discourse On The Hugging Face Incidentastralcodexten.com
July 2026 newslettersimonwillison.net
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Taiyo Yuden Raises Earnings Forecast, Capex Plans on AI Demand
- Legal tech sector sees wave of AI startup acquisitions
- Embodied-AI data startup Kaiwang Data raises RMB100M+
- rust-lang/rust is adopting an LLM policy
- ByteDance launches SeedRealtime full-duplex audio-video model