Rethinking On-Policy Distillation of Large Language Models II: One Training Example
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Argues on-policy distillation is data-overfed but algorithm-starved: state coverage grows quickly with few queries while teacher alignment remains slow. One query reaches 71.5% of full-data OPD state coverage; 16 diverse queries reach 98.9% and match full-data training.
Keywords
distillation · on-policy · data efficiency
Topics
distillation, on-policy, data efficiency
Research notes
- Discovery: Shared in #random-papers as an arXiv link.
- Method: On-policy distillation with state-coverage analysis; measures how much of the full-data state distribution a tiny query set covers.
- Key findings: One training query covers 71.5% of full-data state coverage; 16 diverse queries cover 98.9% and match full-data OPD training performance.
- GitHub code is Apache-2.0. Repo states no dataset ships directly; training data is rebuilt from public sources.