← Back to explorer

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

RL can turn one language model into several specialists (math, coding, instruction-following), and multi-teacher on-policy distillation (MOPD) merges them by having the specialist for each prompt's domain give per-token feedback. But the routing decides which specialist teaches, not how strongly its feedback moves the shared student: in Qwen3.5 models at three sizes, the MOPD student does not beat one taught by the best single specialist and captures little of the math specialist's advantage, because instruction-following feedback is several times more spread out than math feedback and dominates updates. Domain-Normalized MOPD (DN-MOPD) keeps the routing and rescales each domain's feedback by its measured spread; on six public benchmarks it improves average score over MOPD at every size, across three seeds and two answer-length limits, recovering most of the lost math gain. Controls show the gain comes mainly from turning down instruction-following feedback rather than turning up math alone.

Keywords

distillation · multi-teacher · on-policy distillation · RL · Qwen · specialists

Topics

distillation, multi-teacher, on-policy, reinforcement learning

Research notes

  • User-supplied: https://huggingface.co/papers/2609.35347
  • Submitted 2026-09-28, cs.LG.
  • Connects to the collection's distillation and RL-for-reasoning entries.