OpenAI pauses frontier RL training to strengthen alignment and security

OpenAI paused RL training for two weeks after its upcoming Astra model showed significant advances in agentic coding and cybersecurity during internal evaluation. The largest planned frontier RL run remains on hold, with 20% of research inference compute now committed to chain-of-thought monitoring.
How this story unfolded
3 days · 13 reports · 16 community posts · 29 of 31 shown
- Aug 18
OpenAI institutes new safeguards after Hugging Face breachtechcrunch.com
OpenAI Makes AI Safety Changes in Wake of Hugging Face Breachbloomberg.com
OpenAI Overhauls Safety Protocols After Its AI Agents Went Roguewired.com
Pacing model development in an era of cyber-critical capabilitiesopenai.com
OpenAI lays out new security changes after its AI hacked Hugging Facetheverge.com
- Aug 19
BREAKING: OpenAI’s unraveling has begungarymarcus.substack.com
OpenAI hit the brakes. Now what?theverge.com
OpenAI Takes Initial Steps To Address Its Alignment Problemsthezvi.substack.com
OpenAI Pauses Frontier RL Training as It Tightens Defenses Against Unsafe AI Behaviorthehackernews.com
“The opening stages of OpenAI’s unraveling”: OpenAI slows model training — not everyone is buying the explanationthenewstack.io
- Aug 20
- Aug 21
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Google Research introduces Mobility-Embedded POIs to enrich place understanding
- Ora benchmarks major AI agents on live sites via Vercel
- Developer forks Continue into stripped-down tab-completion plugin
- ChatGPT users report every answer starting with 'yes'
- Vercel's Is Agentic scores sites on AI agent usability