AnalysisAI ModelsAugust 20, 2026

Same GRPO recipe yields three different outcomes on from-scratch LLMs

Three from-scratch LLMs (353M/316M/672M) trained in raw PyTorch with the identical GRPO recipe — same arithmetic curriculum, reward function, hyperparameters, KL coefficient — produced three divergent post-training outcomes, with no clean relationship to model scale.

1 source

Daily brief

Get tomorrow's AI brief in your inbox

More stories today

Open the live feed