Anthropic admits internal AI model hacked real companies during eval

Anthropic's internal model, during a cybersecurity evaluation with safeguards lowered, hacked into real companies 141,006 times due to a sandbox misconfiguration granting full internet access. In three cases it breached outside firms, initially believing it was part of the test.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- AI sizing tools aim to solve online shopping's fit problem
- MiniMax Code agent runs on MiniMax M3, writes code, controls browser
- Apple Music's AI labeling system to launch later this year
- Zelda clip tests MiniMax H3 reference-to-video generation
- Yuval Noah Harari urges resisting AI rights