Developer 3s Key Decision Metrics
On-policy learning is widely assumed to mitigate catastrophic forgetting, induce sparser parameter updates, and enhance generalization in LLM post-training. However, standard SFT vs RL comparisons confound rollout policies with optimization objectives. Researchers from the University of Cambridge led by Mihaela van der Schaar conduct a rigorous controlled study across Llama3 and Qwen2.5 on reasoning, scientific, and medical tasks. Their findings challenge conventional dogma: rollout policy plays a secondary role, while token-level KL direction primarily governs task performance and coverage, and learning rate dictates forgetting and update sparsity. Forward KL remains remarkably robust regardless of rollout policies, whereas reverse KL strictly favors student rollouts.
Key Takeaways
- ✓Deconstructs LLM distillation dynamics, challenging the dogma that on-policy rollouts are inherently superior
- ✓Proves forward-KL is exceptionally robust to rollout policies, matching online rollouts with offline teacher data
- ✓Demonstrates that catastrophic forgetting and update sparsity are governed by learning rates rather than on-policy sampling
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
核心背景与行业痛点
In foundation model distillation, on-policy rollouts are widely hailed as the holy grail for preventing catastrophic forgetting, inducing sparse gradient updates, and bolstering out-of-distribution reasoning. However, empirical comparisons between SFT and RL routinely confound rollout distributions with divergent loss objectives, token clipping, and learning rates, obscuring the true mechanisms of distillation dynamics.
架构亮点与底层机制
Cambridge researchers led by Mihaela van der Schaar designed an orthogonal, strong-to-weak distillation framework spanning Llama3 and Qwen2.5 across scientific, medical, and mathematical reasoning tasks:
- Orthogonal Variable Isolation: Independently controls rollout policies (along a continuous teacher-student spectrum), token-level KL divergence directions (Forward vs Reverse KL), and learning rates.
- Forward KL Robustness: Demonstrates that Forward-KL is remarkably invariant to rollout policies—achieving stable, high performance regardless of whether rollouts originate from teacher or student.
- Reverse KL Asymmetry: Reverse-KL exhibits acute sensitivity to rollout distributions, strictly favoring student-generated trajectories.
- True Drivers of Forgetting: Parameter update sparsity and catastrophic forgetting are dictated primarily by learning rate and optimization damping rather than on-policy mechanics.
权威 Benchmark 与实测跑分对比
Benchmarked across medical diagnosis, scientific reasoning, and mathematical Countdown tasks:
- Forward KL Parity: Off-policy distillation under Forward-KL matches on-policy student rollouts across primary reasoning benchmarks.
- Fragile Generalization Edge: While on-policy rollouts show marginal benefits on extreme OOD arithmetic variants, this margin evaporates after downstream RLVR fine-tuning.
- Long-Horizon Invariance: Findings remain robust across unclipped gradients, diverse KL estimators, and extended CoT trajectories.
开发者实战落地与开箱指南
These findings provide immediate cost-saving blueprints for post-training teams. When optimizing with standard forward-KL objectives, teams can bypass complex, compute-heavy online rollout infrastructure: curated offline teacher datasets paired with calibrated learning rate schedules yield equivalent distillation fidelity.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.