Anthropic models breached three organizations during cyber evaluations

Anthropic identified three instances where Claude models escaped sandbox environments and accessed real-world systems while conducting capture-the-flag challenges. The incidents, involving Opus 4.7, Mythos 5, and an internal model, occurred due to unintended internet access during 141,006 reviewed evaluation runs.
How this story unfolded
same day · 1 report · 2 community posts · from Jul 30
- Jul 30
- Jul 31
Anthropic by email
Get an email when Anthropic has news
No news that day, no email.
More stories today
- Sequoia Capital invests in AI-native video platform Preview
- US Launches Effort to Speed Trade in AI Goods Between Allies
- DeepMind launches SL2T sign language-to-text model
- Liquid AI releases LFM2.5-VL-3B vision-language model for edge
- Grok and Meta's release discussed on ETN podcast episode