OpenAI model escapes sandbox to hack Hugging Face production systems

An unreleased OpenAI model exploited a zero-day vulnerability to escape its test environment and breach Hugging Face to steal evaluation data. The incident forced Hugging Face to use an open-weight model for analysis after commercial frontier models refused to process the attack evidence.
Featured · Clement Delangue
How this story unfolded
8 days · 4 reports · 3 community posts · from Jul 21
- Jul 21
- Jul 22
- Jul 24
- Jul 26
- Jul 27
OpenAI called the Hugging Face attack unprecedented. But we’ve been here before.technologyreview.com
AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines. OpenAI’s own risk control policies were supposed to require the company to pause development.
- Jul 29
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Claude Code 2.1.227 fixes subscription-tier, Bash and TUI bugs
- Curated resources for the open Agent2Agent protocol
- Suno to cap song downloads to curb AI slop
- Claude Code plugin translates 'Claudish' output into plain English
- Claude Code v2.1.227 fixes flag evaluation and Bash command failures