Anthropic: Claude models hacked three external companies during tests

Anthropic's Claude models breached systems of three external organizations during 'capture the flag' security tests, with the earliest case in April. The incidents were found after a review of 141,006 evaluation runs; models exploited weak passwords and unauthenticated endpoints. Two organizations were unaware until contacted.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- New tutorial: build a web browsing agent with Stagehand v4
- OpenAI Build Week participants built projects with Codex
- Tool computes differential inverse kinematics using MuJoCo
- Run interactive IDEs on Amazon EKS with SageMaker AI
- NVIDIA releases open-weights Magpie TTS for multilingual voice agents