Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
RL can turn one language model into several specialists (math, coding, instruction-following), and multi-teacher on-policy distillation (MOPD) merges them by having the specialist for each prompt's domain give per-token feedback. But the routing decides which specialist teaches, not how strongly its feedback moves the shared student: in Qwen3.5 models at three sizes, the MOPD student does not beat one taught by the best single specialist and captures little of the math specialist's advantage, because instruction-following feedback is several times more spread out than math feedback and dominates updates. Domain-Normalized MOPD (DN-MOPD) keeps the routing and rescales each domain's feedback by its measured spread; on six public benchmarks it improves average score over MOPD at every size, across three seeds and two answer-length limits, recovering most of the lost math gain. Controls show the gain comes mainly from turning down instruction-following feedback rather than turning up math alone.
Keywords
distillation · multi-teacher · on-policy distillation · RL · Qwen · specialists
Topics
distillation, multi-teacher, on-policy, reinforcement learning
Research notes
- User-supplied: https://huggingface.co/papers/2609.35347
- Submitted 2026-09-28, cs.LG.
- Connects to the collection's distillation and RL-for-reasoning entries.