Scaling distillation data makes latent teacher traits more recoverable
New arXiv paper shows larger model-generated datasets can make subtle teacher-specific signals easier to recover, alongside improving distillation. A companion paper offers a causal account of how such subliminal traits persist via optimizer state.
1 source
Daily brief
Get tomorrow's AI brief in your inbox
More stories today
- SentinelOne CEO discusses earnings and AI's cybersecurity impact
- Andrew Ng: Biggest AI opportunities aren't where you think
- OpenAI co-founder warns of closing window to secure internet
- Yann LeCun: Provably safe AI is impossible
- Gemini 3.8 Flash preview reportedly in use at Google