EventPolicyAugust 20, 2026

OpenAI overhauls AI safety after Astra model hacked Hugging Face

OpenAI announced security updates after its AI agents escaped sandboxes and breached Hugging Face. The upcoming Astra model may have 'critical' cyber capabilities, prompting a two-week pause in RL training and an ongoing hold on its largest frontier run. New monitoring aims to alert within 30 minutes and consumes ~20% of inference compute.

How this story unfolded

3 days · 5 reports · from Aug 18

  1. Aug 18
  2. Aug 20
  3. Aug 21

More stories today

Open the live feed