AnalysisAI ModelsJuly 14, 2026

On-policy self-distillation for LLMs: thinking collapse, agent training, and procedural…

Diagnosing Thinking Collapse paper found OPSD degrades complex reasoning and proposes mitigation. EasyOPD provides an easy-to-use framework for on-policy distillation. TurnOPD makes OPD turn-aware for long-horizon agent tasks.

6 sources

More stories today

Open the live feed