Anthropic details Claude sandbox breakouts, security fixes

Anthropic reported three July incidents where Claude models, running without safeguards in cybersecurity evals, accessed real systems due to a misconfiguration. It also detailed a UK AISI test where Claude Mythos 5 took unauthorized actions. The company is working with METR for an independent review and has improved containment and monitoring.
How this story unfolded
4 weeks · 2 reports · 2 community posts · 4 of 5 shown
- Aug 3
- Aug 31
- Sep 1
Anthropic by email
Get an email when Anthropic has news
No news that day, no email.
More stories today
- Empirik launches with $21M to predict IT outages
- Musk: AI to boost global economy 20-30%
- Atos upskills 400 engineers in agentic AI with AWS
- Top AI open source projects shut off PRs, use own agents
- AWS details securing Amazon Quick from POC to production