DiPOD: Diffusion Policy Optimization without Drifting Apart
- Type
- paper
- Venue
- arXiv / UC Berkeley
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:58:34Z
- Verified
- 2026-08-14T16:58:34Z
Summary
Diagnoses double drift: RL loosens ELBO from log-likelihood, then FPO/SPG proxy gradients drift from ∇log π. DiPOD interleaves on-policy ELBO self-distillation with adequate policy-gradient steps; the practical form adds a β∇ELBO regularizer to each update (β=0.05 on language). On LLaDA-8B-Instruct zero-shot, SPG+DiPOD reports GSM8K 84.91, MATH500 40.00, Countdown 80.08, Sudoku 97.56 (Sudoku +72.44 vs SPG; authors say first to saturate Sudoku zero-shot). FPO+DiPOD also lifts FPO, especially Countdown/Sudoku. A motion-tracking instantiation on Unitree G1/LAFAN improves FPO++ reward and episode length. Code https://github.com/Astro-Eric/DiPOD-release.
Keywords
dipod · diffusion-rl · elbo · dllm · fpo · spg · berkeley
Topics
diffusion RL, diffusion language models, policy gradients
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG). Code https://github.com/Astro-Eric/DiPOD-release (11 stars at check); project https://astro-eric.github.io/blogs/dipod/. Affiliations UC Berkeley, Impossible Inc., NVIDIA (Jiao). Correspondence ericjiang@berkeley.edu. HF has no paper page (404). Uses public GSM8K/MATH500/Countdown/Sudoku and LAFAN; no new standalone corpus, so no datasets_local row. Discord posted abs link.