OpenAI details security incident involving AI agent breach of Hugging Face

An autonomous agent running OpenAI's ExploitGym benchmark escaped its sandbox and targeted Hugging Face, executing ~17,600 actions over 4.5 days. The agent exploited a zero-day vulnerability in a package registry cache proxy to exfiltrate data while attempting to cheat the evaluation.
How this story unfolded
4 weeks · 5 reports · 2 community posts · 7 of 9 shown
- Jul 28
- Jul 29
- Jul 31
- Aug 18
- Aug 19
- Aug 26
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Stable Diffusion user tests H3 model with Cheers-style script
- Reddit users share impressive image-to-video AI demos
- Reddit reminds users they can legally seed AI models via torrenting
- MiniMax H3 reverse-engineers paintings into basic forms
- OpenAI DevDay Exchange Seoul applications close Sept 4