AnalysisAI ModelsSeptember 12, 2026

Paper: RL for LLMs mostly improves easy tasks, NGU fixes it

A paper on RL post-training of LLMs finds gains are proportional to problem difficulty — on AIME 2025, easy-subset pass@1 rose from 22.7% while hard problems barely moved. The proposed NGU adaptive sampling reallocates compute to harder problems.

2 sources

More stories today

Open the live feed