Developer 3s Key Decision Metrics
While reinforcement learning produces distinct specialized expert models excels at mathematics, code synthesis, or instruction following, real-world deployment requires a single unified foundation model possessing all of these capabilities. Multi-teacher on-policy distillation (MOPD) aims to merge these capabilities by routing student-sampled prompts to domain specialists for token-level supervision. However, researchers discover an acute imbalance across Qwen3.5 architectures: instruction-following teachers exhibit broad distribution dispersion, generating gradient magnitudes that overwhelm dense mathematical signals and degrading mathematical reasoning. The team introduces Domain-Normalized MOPD (DN-MOPD, arXiv:2609.35347), which normalizes each domain's supervisory signal by its measured empirical spread. Across six public benchmarks, DN-MOPD consistently outperforms standard MOPD across all model scales, recovering mathematical reasoning performance without sacrificing instruction compliance.
Key Takeaways
- ✓Uncovers that instruction-following supervision dispersion dominates MOPD updates, washing out compact mathematical gradients
- ✓Introduces Domain-Normalized MOPD (DN-MOPD) to rescale teacher feedback inversely by measured distribution spread
- ✓Demonstrates consistent gains across six public benchmarks across Qwen3.5 scales, recovering lost math reasoning gains
Turn your technical choice into a development budget
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
核心背景与行业痛点
While reinforcement learning can forge specialized language models excelling at distinct competencies—such as mathematics, code synthesis, or instruction following—production applications necessitate a single model unifying all proficiencies. Multi-teacher on-policy distillation (MOPD) addresses this by letting specialized teachers supervise a shared student based on prompt domain routing. However, empirical audits across Qwen3.5 architectures reveal severe performance degradation: student models underperform single-specialist baselines and sacrifice nearly all advantages contributed by the mathematical reasoning teacher.
架构亮点与底层机制
Researchers from Nanyang Technological University and Tsinghua University introduce DN-MOPD (arXiv:2609.35347), analyzing the mechanics of multi-teacher distillation failure:
- Uncovering Supervisory Feedback Imbalance: Demonstrates that domain routers determine which specialist supervises, but fail to balance the magnitude of their feedback. Instruction-following feedback exhibits far broader token distribution dispersion than mathematical feedback, dominating parameter updates and drowning out sparse mathematical gradients.
- Domain-Normalized Rescaling: Formulates DN-MOPD, which scales each specialist's supervisory loss inversely by its measured empirical distribution dispersion.
- Calibrating General Feedback Restores Reasoning: Ablations confirm that recovering mathematical performance does not require upweighting mathematics; rather, it requires dampening the disproportionate gradient spread of instruction-following supervision.
- Clean Single-Model Delivery: Produces a self-contained student checkpoint possessing verified general and deep reasoning capabilities with zero inference overhead.
权威 Benchmark 与实测跑分对比
Evaluated on Qwen3.5 across three parameter scales on six challenging public benchmarks:
- Universal Performance Superiority: DN-MOPD achieves higher composite scores than standard MOPD across all parameter scales, three distinct random seeds, and two generation length thresholds.
- Recovery of Mathematical Capabilities: Restores over 92.4% of the standalone mathematics teacher's domain superiority without causing instruction-following regression.
- Proven Stability: Delivers robust optimization stability, preventing gradient oscillations commonly encountered in multi-objective distillation.
开发者实战落地与开箱指南
DN-MOPD provides an indispensable architectural blueprint for post-training teams. When distilling distinct RL specialists into a unified general-purpose model, teams should replace naive domain loss summing with dispersion normalization, safeguarding complex analytical reasoning capabilities while maintaining conversational versatility.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.