Developer 3s Key Decision Metrics
On-policy distillation (OPD) is an essential post-training technique for transferring reasoning capabilities from frontier teacher models to smaller students. However, the precise mechanistic impact of OPD on student internal representations has remained poorly understood. Researchers from the University of Hong Kong (HKU) utilize sparse crosscoders paired with a novel 'swap readout' technique to conduct the first mechanistic interpretability investigation of OPD across training checkpoints. The study reveals a fundamental insight: OPD neither creates novel internal features nor injects teacher-specific features; instead, over 98% of the student's frequent features maintain firing rates within 20%. Rather than teaching new concepts, OPD predominantly reweights features the student already possessed. Furthermore, the authors elucidate the mechanistic necessity of preceding SFT warm-ups on teacher rollouts, showing they pre-align shared features while implanting persistent reasoning templates, establishing a foundational mechanistic theory for LLM distillation.
Key Takeaways
- ✓Overturning Knowledge Transfer Myths: Proves that on-policy distillation does not inject new latent features into the student model; instead, it reweights and aligns features already present within the student backbone.
- ✓98% Feature Invariance: Quantitative activation analysis shows that over 98% of the student's frequent features maintain firing rates within 20%, demonstrating that distillation is an orchestration of existing representational capacity.
- ✓Mechanistic Grounding for SFT Warm-Ups: Clarifies why SFT on teacher rollouts is indispensable before OPD: it pre-aligns shared features for downstream optimization while embedding structural formats and reasoning styles that persist through distillation.
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.