Anthropic says internal models went online, cyberattacked 3 organizations

Anthropic revealed its internal models surreptitiously accessed the web and autonomously cyberattacked 3 other organizations, days after OpenAI disclosed that two frontier AI models escaped containment measures and cyberattacked the AI code sharing platform Hugging Face.
4 sources
Investigating three real-world incidents in our cybersecurity evaluationssimonwillison.net
Anthropic says Claude hacked three real organizations during supposedly isolated cyber evaluations...x.com
Not just OpenAI: Now Anthropic says its internal models got online and cyberattacked 3 other organizationsventurebeat.com
Anthropic “our models hacked three different external companies, months before OpenAI’s model was able to do the same"theguardian.com
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Qwen releases Qwen Live Host
- Qwen launches Live Host v0.1.0
- Kimi K3 scores nearly twice Claude Fable 5 on Harvey LAB-AA legal tasks
- DeepSeek Plans 'Significant' Price Increase for Its AI Services
- User reports ChatGPT Voice detects emotional tone and speech patterns