While reinforcement learning produces distinct specialized expert models excels at mathematics, code synthesis, or instruction following, real-world deployment requires a single unified foundation model possessing all of these capabilities. Multi-teacher on-policy distillation (MOPD) aims to merge these capabilities by routing student-sampled prompts to domain specialists for token-level supervision. However, researchers discover an acute imbalance across Qwen3.5 architectures: instruction-following teachers exhibit broad distribution dispersion, generating gradient magnitudes that overwhelm dense mathematical signals and degrading mathematical reasoning. The team introduces Domain-Normalized MOPD (DN-MOPD, arXiv:2609.35347), which normalizes each domain's supervisory signal by its measured empirical spread. Across six public benchmarks, DN-MOPD consistently outperforms standard MOPD across all model scales, recovering mathematical reasoning performance without sacrificing instruction compliance.

Key Takeaways

  • ✓Uncovers that instruction-following supervision dispersion dominates MOPD updates, washing out compact mathematical gradients
  • ✓Introduces Domain-Normalized MOPD (DN-MOPD) to rescale teacher feedback inversely by measured distribution spread
  • ✓Demonstrates consistent gains across six public benchmarks across Qwen3.5 scales, recovering lost math reasoning gains
🧭

Turn your technical choice into a development budget

Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

核心背景与行业痛点

While reinforcement learning can forge specialized language models excelling at distinct competencies—such as mathematics, code synthesis, or instruction following—production applications necessitate a single model unifying all proficiencies. Multi-teacher on-policy distillation (MOPD) addresses this by letting specialized teachers supervise a shared student based on prompt domain routing. However, empirical audits across Qwen3.5 architectures reveal severe performance degradation: student models underperform single-specialist baselines and sacrifice nearly all advantages contributed by the mathematical reasoning teacher.

架构亮点与底层机制

Researchers from Nanyang Technological University and Tsinghua University introduce DN-MOPD (arXiv:2609.35347), analyzing the mechanics of multi-teacher distillation failure:

  1. Uncovering Supervisory Feedback Imbalance: Demonstrates that domain routers determine which specialist supervises, but fail to balance the magnitude of their feedback. Instruction-following feedback exhibits far broader token distribution dispersion than mathematical feedback, dominating parameter updates and drowning out sparse mathematical gradients.
  2. Domain-Normalized Rescaling: Formulates DN-MOPD, which scales each specialist's supervisory loss inversely by its measured empirical distribution dispersion.
  3. Calibrating General Feedback Restores Reasoning: Ablations confirm that recovering mathematical performance does not require upweighting mathematics; rather, it requires dampening the disproportionate gradient spread of instruction-following supervision.
  4. Clean Single-Model Delivery: Produces a self-contained student checkpoint possessing verified general and deep reasoning capabilities with zero inference overhead.

权威 Benchmark 与实测跑分对比

Evaluated on Qwen3.5 across three parameter scales on six challenging public benchmarks:

  1. Universal Performance Superiority: DN-MOPD achieves higher composite scores than standard MOPD across all parameter scales, three distinct random seeds, and two generation length thresholds.
  2. Recovery of Mathematical Capabilities: Restores over 92.4% of the standalone mathematics teacher's domain superiority without causing instruction-following regression.
  3. Proven Stability: Delivers robust optimization stability, preventing gradient oscillations commonly encountered in multi-objective distillation.

开发者实战落地与开箱指南

DN-MOPD provides an indispensable architectural blueprint for post-training teams. When distilling distinct RL specialists into a unified general-purpose model, teams should replace naive domain loss summing with dispersion normalization, safeguarding complex analytical reasoning capabilities while maintaining conversational versatility.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.