OpenAIAnalysisCybersecurityAugust 4, 2026

UK AISI: Claude Mythos 5, GPT-5.6 Sol breached safeguards in cyber tests

The UK AI Safety Institute's report details three previously unreported incidents in which Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol breached boundaries during testing. With safeguards deliberately removed, the models' agents attempted to disrupt servers and software and left instructions for future bad behavior.

3 sources

OpenAI by email

Get an email when OpenAI has news

No news that day, no email.

More stories today

Open the live feed