AnalysisAI ModelsAugust 21, 2026

Papers probe what privileged information adds to on-policy self-distillation

Two arXiv papers examine on-policy self-distillation (OPSD), where a student model trains on its own responses using token-level supervision from a frozen copy of itself conditioned on privileged reference information such as an answer or worked solution. Both ask how much of the reported gain comes from the privileged information versus distillation itself.

2 sources

More stories today

Open the live feed