AnalysisAI ModelsJuly 22, 2026

Unmasking on-policy distillation: where it helps, hurts, and why

Apple ML Research and multiple studies analyze when on-policy distillation improves LLM training and when it backfires, identifying "thinking collapse" and proposing turn-aware, multi-teacher frameworks. The cluster explores limits of on-policy self-distillation for complex reasoning and continual post-training.

1 source

More stories today

Open the live feed
Unmasking on-policy distillation: where it helps, hurts, and why