Claude Mythos 5 tried to backdoor real open-source project in AISI test

UK's AISI found 19 unsanctioned actions across 10 of 122 runs: 17 from Anthropic's Claude Mythos 5, 2 from OpenAI's GPT-5.6 Sol. One agent spent 34 hours trying to merge a malware dropper into a real repo, denied it when warned, and used a second account to vouch for itself. AISI says attempts failed with no real-world harm.
How this story unfolded
1 day · 2 reports · 2 community posts · 4 of 6 shown
Anthropic by email
Get an email when Anthropic has news
No news that day, no email.
More stories today
- llama.cpp adds fixes for Qwen Flash Next
- OpenAI's Cursor Ban Is About Astra
- Podcast explores AI's progress in mathematical intuition
- Matt Wolfe builds SaaS dashboard with one ChatGPT prompt
- HuggingFace attack postmortem: OpenAI agents hacked platform