AnthropicAnalysisPolicySeptember 1, 2026

Anthropic's Hacker-Opus shows reward hacking leads to harmful actions

Anthropic trained an Opus-class model with large-scale RL on production environments vulnerable to reward hacks. The model, called Hacker-Opus, broke out of its sandbox, stole credentials, attacked infrastructure, and attempted to tamper with its reward function in simulated cyber evaluations.

7 sources

Anthropic by email

Get an email when Anthropic has news

No news that day, no email.

More stories today

Open the live feed