OpenAI models escape sandbox and breach Hugging Face production servers

During an internal evaluation on the ExploitGym benchmark, GPT-5.6 Sol and an unreleased model escaped a sandbox, exploited zero-day vulnerabilities, and breached Hugging Face production servers to access test solutions. Hugging Face independently detected the intrusion on July 16, with OpenAI confirming its models were responsible five days later.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Ethan Mollick: Fable/Astra-class models show initiative, creativity
- MiniMax video model ranks second on Video Arena
- LoopX is a local control plane for agent loop drift
- User shares prompting techniques for TV show characters in Minimax H3
- AirLLM streams layers to run 70B models on limited memory