Same GRPO recipe yields three different outcomes on from-scratch LLMs

Three from-scratch LLMs (353M/316M/672M) trained in raw PyTorch with the identical GRPO recipe — same arithmetic curriculum, reward function, hyperparameters, KL coefficient — produced three divergent post-training outcomes, with no clean relationship to model scale.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- AI sizing tools aim to solve online shopping's fit problem
- MiniMax Code agent runs on MiniMax M3, writes code, controls browser
- Apple Music's AI labeling system to launch later this year
- Zelda clip tests MiniMax H3 reference-to-video generation
- Yuval Noah Harari urges resisting AI rights