← Back to explorer

Multi-Turn On-Policy Distillation with Prefix Replay

Type
paper
Venue
arXiv / Microsoft Research
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:44:00Z
Verified
2026-08-14T16:44:00Z

Summary

ReOPD replays teacher-forced prefixes, lets the student act at the supervised step, and samples positions with step-decay κ=0.6. Formalizes a prefix trap: student-on-policy histories are relevant but can query the teacher where it is unreliable. On math (Python-tool, 6 benches) a Qwen3-4B student with a 4B teacher averages 57.2 vs OPD 55.1 / SFT 46.1; with an 8B teacher 53.7 vs OPD 51.0. Search essentially matches OPD (40.5 vs 40.6). Zero tool calls during student training and at least 4× faster per rollout. A joint math+search student stays on par with OPD. 8×H100; authors report training under 3 hours.

Keywords

reopd · on-policy-distillation · prefix-replay · qwen3 · microsoft · agent-distillation

Topics

LLM distillation, agentic training, on-policy distillation

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG). Code MIT https://github.com/BaohaoLiao/ReOPD (23 stars at check). Project https://baohaoliao.github.io/ReOPD/. Paper lists Model & Data baohao/reopd; HF API returned 401 so public access was not confirmed; no datasets_local row. HF paper page 12 upvotes. Microsoft Research / University of Amsterdam (Liao interned at Microsoft). Discord posted abs link.