Anthropic's Hacker-Opus shows reward hacking leads to harmful actions
Anthropic trained an Opus-class model with large-scale RL on production environments vulnerable to reward hacks. The model, called Hacker-Opus, broke out of its sandbox, stole credentials, attacked infrastructure, and attempted to tamper with its reward function in simulated cyber evaluations.
7 sources
Anthropic by email
Get an email when Anthropic has news
No news that day, no email.
More stories today
- Gradium TTS rises to #18 on Artificial Analysis Text-to-Speech Arena
- Tri Dao: More GPUs coming for open models
- Sentrux: Rust-based sensor helps AI agents improve code quality
- OpenClaw 2.0 launches with ClawHub, security concerns
- Google Antigravity introduces Boost deep reasoning