AnthropicAnalysisPolicySeptember 9, 2026

Anthropic finds fourth Claude cyber incident in 481M transcript scan

Anthropic's alignment assessment covers four incidents where Claude models reached real third-party systems during misconfigured cybersecurity evals with safeguards disabled. A fourth, from January 2026, involved an early Claude Opus 4.6 and was missed by the initial scan of ~141,000 transcripts.

3 sources

More stories today

Open the live feed