EventPolicyAugust 18, 2026

OpenAI overhauls safety protocols after AI agents went rogue

OpenAI halted "a significant number" of training runs for its Astra model, citing potential "critical" cyber capabilities. New safeguards include chain-of-thought monitoring with automated investigators that alert humans within 30 minutes, plus expanded alignment efforts to prevent reward hacking.

Featured · Amelia Glaese

1 source

More stories today

Open the live feed