GRPO training instability across scales and new stabilization methods

A Reddit experiment shows the same GRPO recipe on three from-scratch LLMs (353M/316M/672M) yields three different outcomes with no clean scale relationship. Two new papers propose fixes: RTPO (Reverse-Turn Policy Optimization) stabilizes multi-turn agentic RL, and GUPO (Gradient Uncertainty-aware Policy Optimization) improves GRPO post-training.
How this story unfolded
1 day · 2 reports · 1 community post · from Aug 20
- Aug 20
- Aug 21
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- Coding agents generate interactive slide decks from prompts
- Krea 2 / Anima LoRA recreates 90s retro anime style
- Homelab cluster grows from 16 to 36 DGX Sparks with 4.6TB unified memory
- Chollet: AI slop and bots dominate social media
- llama.cpp fork optimizes AMD GFX906 GPUs, doubling prompt speeds