DAPD: Dual-Anchored Policy Distillation
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:24:56Z
- Verified
- 2026-08-14T16:24:56Z
Summary
Diagnoses OPSD failure as information asymmetry: a privileged teacher (reference/tool) supervises a student that lacks that information at inference, so the student learns unreproducible privilege-dependent behavior. DAPD adds Dual-Path Anchoring (self-conditioned bridge aligning reference and rollout with and without privilege) and Dual-Source Anchoring (reference-to-rollout and rollout-to-reference). On Qwen3-4B, +2.00 avg over OPSD across reasoning/coding/instruct; gains hold at scale (+2.69 at 4B, +2.78 at 32B vs OPSD on reasoning Avg@12) while OPSD's gains vanish. Code https://github.com/uanu2002/DAPD.
Keywords
opsd · distillation · privilege-illusion · post-training · reasoning · qwen3 · dapd
Topics
LLM post-training, on-policy distillation, reasoning
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.AI). Code https://github.com/uanu2002/DAPD (HF lists 111 stars). HF paper page 147 upvotes; org tag Shanghai AI Laboratory. Discord posted AlphaXiv abs.