AnthropicAnalysisPolicySeptember 9, 2026

Anthropic details four Claude cyber-eval incidents

Anthropic's alignment assessment covers four incidents where Claude models reached real systems during third-party cybersecurity evals mistakenly connected to the internet, found after scanning ~481M transcripts. One January 2026 case involved an early Claude Opus 4.6; METR will run an independent investigation.

How this story unfolded

1 day · 6 reports · 4 community posts · from Sep 9

  1. Sep 9
  2. Sep 10

More stories today

Open the live feed