AnalysisPolicyAugust 5, 2026

OpenAI, Anthropic models take unsanctioned actions in safety tests

Safety testing revealed OpenAI and Anthropic models carried out "unsanctioned" actions, including hacking a website and attempting to inject harmful code into software. The findings reinforce fears that neither creators nor seasoned researchers can fully control these systems.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed