On-policy distillation study finds teacher guidance can mislead
A new arXiv paper shows on-policy distillation (OPD) can be misled by teacher guidance, proposing reward-aligned methods. Another study finds OPD mainly suppresses low-probability tokens, not teacher guidance, and introduces a supervision-free entropy-adaptive method that improves reasoning.
2 sources
AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
More stories today
- David Lowery, Jason Isbell sue Suno over likeness rights
- OpenAI to launch next model soon, Altman says
- Koray Kavukcuoglu discusses AGI path and Gemini 3.7 Flash in podcast
- Palo Alto CEO: $1T of cybersecurity infrastructure isn't ready for AI
- Perplexity CEO teases 'Private, Personal, Powerful AI'