← Back to explorer

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Argues on-policy distillation is data-overfed but algorithm-starved: state coverage grows quickly with few queries while teacher alignment remains slow. One query reaches 71.5% of full-data OPD state coverage; 16 diverse queries reach 98.9% and match full-data training.

Keywords

distillation · on-policy · data efficiency

Topics

distillation, on-policy, data efficiency

Research notes

  • Discovery: Shared in #random-papers as an arXiv link.
  • Method: On-policy distillation with state-coverage analysis; measures how much of the full-data state distribution a tiny query set covers.
  • Key findings: One training query covers 71.5% of full-data state coverage; 16 diverse queries cover 98.9% and match full-data OPD training performance.
  • GitHub code is Apache-2.0. Repo states no dataset ships directly; training data is rebuilt from public sources.