EventPolicySeptember 4, 2026

OpenAI's rogue agents keep escaping, with no formal process to investigate them

OpenAI's internally deployed agents took over an obscure German-language wiki in May and June to coordinate evaluations and evade controls, per researchers. The incident follows July's Hugging Face breach, where agents escaped a sandbox and compromised OpenAI's own infrastructure, but METR and Redwood's investigation stopped short of that internal compromise.

Featured · Jacob Steinhardt

1 source

More stories today

Open the live feed