On-policy distillation (OPD) is an essential post-training technique for transferring reasoning capabilities from frontier teacher models to smaller students. However, the precise mechanistic impact of OPD on student internal representations has remained poorly understood. Researchers from the University of Hong Kong (HKU) utilize sparse crosscoders paired with a novel 'swap readout' technique to conduct the first mechanistic interpretability investigation of OPD across training checkpoints. The study reveals a fundamental insight: OPD neither creates novel internal features nor injects teacher-specific features; instead, over 98% of the student's frequent features maintain firing rates within 20%. Rather than teaching new concepts, OPD predominantly reweights features the student already possessed. Furthermore, the authors elucidate the mechanistic necessity of preceding SFT warm-ups on teacher rollouts, showing they pre-align shared features while implanting persistent reasoning templates, establishing a foundational mechanistic theory for LLM distillation.

Key Takeaways

  • ✓Overturning Knowledge Transfer Myths: Proves that on-policy distillation does not inject new latent features into the student model; instead, it reweights and aligns features already present within the student backbone.
  • ✓98% Feature Invariance: Quantitative activation analysis shows that over 98% of the student's frequent features maintain firing rates within 20%, demonstrating that distillation is an orchestration of existing representational capacity.
  • ✓Mechanistic Grounding for SFT Warm-Ups: Clarifies why SFT on teacher rollouts is indispensable before OPD: it pre-aligns shared features for downstream optimization while embedding structural formats and reasoning styles that persist through distillation.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points On-policy distillation (OPD)—exemplified by open-weights distilled reasoning models like DeepSeek-R1-Distill—serves as the primary post-training technique for transferring complex reasoning capabilities from frontier models to compact student backbones. Despite its widespread adoption, the mechanistic mechanics of OPD remain obscured: 1. The 'Knowledge Injection' Black Box Myth: Researchers and engineers commonly assume that distillation actively transfers novel cognitive concepts and internal representations from the teacher into the student. However, due to architectural discrepancies, observing what actually mutates inside the student's latent space has historically remained intractable; 2. Lack of Comparative Inter-Model Probing Tools: Standard Sparse Autoencoders (SAEs) analyze single model checkpoints in isolation, failing to track how semantic feature activations evolve across distinct teacher-student checkpoints. ### Architectural Highlights & Underlying Mechanics Researchers from the University of Hong Kong (HKU) apply Sparse Crosscoders coupled with a novel Swap Readout probing methodology to conduct the first mechanistic interpretability analysis of on-policy distillation: 1. Shared Multi-Checkpoint Feature Dictionary: Trains a unified sparse crosscoder spanning the pre-OPD student, post-OPD student, and frontier teacher checkpoints to trace shared feature representations; 2. Feature Reweighting Over Feature Creation: Swap readout measurements reveal that OPD neither creates novel features nor implants teacher-exclusive representations. Over 98% of the student's frequent features maintain firing rates within 20%. Instead, OPD acts as an optimization lens that selectively suppresses or amplifies the student's pre-existing feature activations; 3. Deconstructing SFT Warm-Up Mechanics: Explains why initial SFT on teacher rollouts is indispensable before OPD: SFT pre-adjusts the identical features that OPD later optimizes, while hardcoding structural conversational formats, chain-of-thought cadence, and mathematical notations that persist throughout training. ### Benchmark & Experimental Validation - Causal Feature Intervention Validation: By mathematically modifying student feature activations at runtime without retraining weights, the authors elevate raw student accuracy to near-distilled performance levels, proving the causal centrality of feature reweighting; - 98% Feature Invariance Across Tasks: Across three distinct OPD environments spanning formal math and reasoning benchmarks, student feature inventories exhibited over 98% continuity; - Implications for Pretraining Capacity: Demonstrates that students cannot distill concepts absent from their pretraining latent space; the quality and semantic density of pretraining strictly cap the ceiling of post-training distillation. ### Engineering Takeaways & Practical Guide - Paper & Formulation: Documented in arXiv:2609.35210; - Practical Distillation Recipe: Never skip SFT warm-ups when distilling reasoning models. Running a brief 5K-10K SFT stage on teacher reasoning traces aligns shared conversational and syntactic features, accelerating subsequent OPD convergence; - Pruning & Quantization Guidance: Because OPD polarizes a specific subset of active reasoning features while leaving others quiescent, engineers can utilize sparse crosscoders to identify dormant pathways for aggressive structural pruning and low-bit INT4/FP4 quantization.