← Back to explorer

DAPD: Dual-Anchored Policy Distillation

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:24:56Z
Verified
2026-08-14T16:24:56Z

Summary

Diagnoses OPSD failure as information asymmetry: a privileged teacher (reference/tool) supervises a student that lacks that information at inference, so the student learns unreproducible privilege-dependent behavior. DAPD adds Dual-Path Anchoring (self-conditioned bridge aligning reference and rollout with and without privilege) and Dual-Source Anchoring (reference-to-rollout and rollout-to-reference). On Qwen3-4B, +2.00 avg over OPSD across reasoning/coding/instruct; gains hold at scale (+2.69 at 4B, +2.78 at 32B vs OPSD on reasoning Avg@12) while OPSD's gains vanish. Code https://github.com/uanu2002/DAPD.

Keywords

opsd · distillation · privilege-illusion · post-training · reasoning · qwen3 · dapd

Topics

LLM post-training, on-policy distillation, reasoning

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.AI). Code https://github.com/uanu2002/DAPD (HF lists 111 stars). HF paper page 147 upvotes; org tag Shanghai AI Laboratory. Discord posted AlphaXiv abs.