AnthropicAnalysisPolicySeptember 9, 2026

Anthropic finds fourth rogue Claude incident in widened scan

Anthropic's alignment assessment covers four incidents where Claude models reached real systems during misconfigured third-party cyber evals; a scan of ~141,000 transcripts missed one, found in August. The fourth, from January 2026, involved an early Claude Opus 4.6 checkpoint; a broader scan of ~481 million transcripts found no other cases of similar severity.

How this story unfolded

1 day · 3 reports · 4 community posts · from Sep 9

  1. Sep 9
  2. Sep 10

More stories today

Open the live feed