AnalysisAI ModelsSeptember 29, 2026

Three papers probe on-policy distillation for LLM reasoning

Read original source →arxiv.org

Sparse crosscoders are used to inspect what on-policy distillation actually writes into a student's internal representations. Direct-OPD work argues only some tokens deserve supervision; LastOPD targets collapse in latent-state alignment.

How this story unfolded

3 days · 14 reports · from Sep 28

  1. Sep 28
  2. Sep 29
  3. Sep 30
  4. Oct 1

More stories today

Open the live feed