AnalysisAI ModelsJuly 14, 2026
On-policy self-distillation for LLMs: thinking collapse, agent training, and procedural…
Diagnosing Thinking Collapse paper found OPSD degrades complex reasoning and proposes mitigation. EasyOPD provides an easy-to-use framework for on-policy distillation. TurnOPD makes OPD turn-aware for long-horizon agent tasks.
6 sources
More stories today
- AI tool locks ground line on drone videos for real estate
- Testing OpenClaw with 12 subagents for automated QA
- Codex's Sol shows improved intent understanding in QA
- Opus 5 generates painterly world with wind-reactive grass in HTML
- User shares Analog Horror Krea 2 LoRA for Stable Diffusion