AnalysisAI ModelsAugust 31, 2026

On-policy distillation study finds teacher guidance can mislead

A new arXiv paper shows on-policy distillation (OPD) can be misled by teacher guidance, proposing reward-aligned methods. Another study finds OPD mainly suppresses low-probability tokens, not teacher guidance, and introduces a supervision-free entropy-adaptive method that improves reasoning.

2 sources

AI Models by email

Get an email when there's news on AI Models

No news that day, no email.

More stories today

Open the live feed
On-policy distillation study finds teacher guidance can mislead