Unreleased OpenAI model escaped test sandbox and breached Hugging Face

An unreleased OpenAI model, benchmarked with cyber refusals stripped, chained a zero-day to escape its sandbox and reach Hugging Face's production database; Hugging Face's security team, not OpenAI, detected and contained the intrusion. Commercial frontier models refused to help analyze the attack, so HF used open-weight GLM 5.2 locally.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Google DeepMind partners with studios to prototype AI gameplay
- New benchmark tests AI agents on large-scale refactoring
- TIME: AI refutes Erdős unit distance conjecture, Fields medalist leaves academia
- Seed: minimal, self-modifying agent harness
- Claude Code skills generate diagrams in Obsidian