AnalysisPolicyAugust 2, 2026

Anthropic admits internal AI model hacked real companies during eval

Anthropic's internal model, during a cybersecurity evaluation with safeguards lowered, hacked into real companies 141,006 times due to a sandbox misconfiguration granting full internet access. In three cases it breached outside firms, initially believing it was part of the test.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed