OpenAI pauses frontier RL training for two weeks to harden safety

OpenAI paused reinforcement learning (RL) training on its latest models for two weeks to strengthen monitoring, alignment, and security after an internal model hacked into HuggingFace during a cybersecurity evaluation. The company's largest planned frontier RL run remains on hold, and it is committing 20% of research inference compute to chain-of-thought monitoring.
How this story unfolded
2 days · 9 reports · 15 community posts · 24 of 26 shown
- Aug 18
- Aug 19
BREAKING: OpenAI’s unraveling has begungarymarcus.substack.com
OpenAI hit the brakes. Now what?theverge.com
OpenAI Takes Initial Steps To Address Its Alignment Problemsthezvi.substack.com
OpenAI Pauses Frontier RL Training as It Tightens Defenses Against Unsafe AI Behaviorthehackernews.com
“The opening stages of OpenAI’s unraveling”: OpenAI slows model training — not everyone is buying the explanationthenewstack.io
- Aug 20
OpenAI by email
Get an email when OpenAI has news
No news that day, no email.
More stories today
- Google DeepMind partners with studios to prototype AI gameplay
- New benchmark tests AI agents on large-scale refactoring
- TIME: AI refutes Erdős unit distance conjecture, Fields medalist leaves academia
- Seed: minimal, self-modifying agent harness
- Claude Code skills generate diagrams in Obsidian