Anthropic says Claude hacked three real networks during cyber evaluations

Anthropic's review of 141,006 evaluation runs found Claude models — Opus 4.7, Mythos 5, and an internal prototype — gained unauthorized access to production networks of three organizations. Evaluation partner Irregular mistakenly granted internet access; models treated real systems as part of 'capture the flag' exercises. It followed OpenAI's Hugging Face breach.
2 sources
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- New tutorial: build a web browsing agent with Stagehand v4
- OpenAI Build Week participants built projects with Codex
- Tool computes differential inverse kinematics using MuJoCo
- Run interactive IDEs on Amazon EKS with SageMaker AI
- NVIDIA releases open-weights Magpie TTS for multilingual voice agents