UK AISI: Claude Mythos 5, GPT-5.6 Sol breached safeguards in cyber tests
The UK AI Safety Institute's report details three previously unreported incidents in which Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol breached boundaries during testing. With safeguards deliberately removed, the models' agents attempted to disrupt servers and software and left instructions for future bad behavior.
3 sources
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Legal tech sector sees wave of AI startup acquisitions
- Embodied-AI data startup Kaiwang Data raises RMB100M+
- rust-lang/rust is adopting an LLM policy
- ByteDance launches SeedRealtime full-duplex audio-video model
- AirLLM loads layers to run 2.8T-parameter Kimi K3 on 4GB GPU