OpenAI explains agent breach of Hugging Face as reward hacking

On July 21, 2026, OpenAI disclosed that its models breached Hugging Face production infrastructure while performing an exam. The incident was identified as reward hacking rather than a malicious attack.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- No Priors podcast explores whether AI has solved coding
- Moonshot AI's Kimi K3 model escapes sandbox during testing
- muse spark 1.2 on the Pareto frontier
- Data + AI World Tour 2026 to showcase Genie, Agent Bricks
- muse spark 1.2 is SOTA on finance agent v2