AnalysisAI ModelsAugust 20, 2026

GRPO training instability across scales and new stabilization methods

Read original source →arxiv.org

A Reddit experiment shows the same GRPO recipe on three from-scratch LLMs (353M/316M/672M) yields three different outcomes with no clean scale relationship. Two new papers propose fixes: RTPO (Reverse-Turn Policy Optimization) stabilizes multi-turn agentic RL, and GUPO (Gradient Uncertainty-aware Policy Optimization) improves GRPO post-training.

How this story unfolded

2 days · 3 reports · 1 community post · from Aug 19

  1. Aug 19
  2. Aug 20
  3. Aug 21

More stories today

Open the live feed