← Back to explorer

DiPOD: Diffusion Policy Optimization without Drifting Apart

Type
paper
Venue
arXiv / UC Berkeley
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:58:34Z
Verified
2026-08-14T16:58:34Z

Summary

Diagnoses double drift: RL loosens ELBO from log-likelihood, then FPO/SPG proxy gradients drift from ∇log π. DiPOD interleaves on-policy ELBO self-distillation with adequate policy-gradient steps; the practical form adds a β∇ELBO regularizer to each update (β=0.05 on language). On LLaDA-8B-Instruct zero-shot, SPG+DiPOD reports GSM8K 84.91, MATH500 40.00, Countdown 80.08, Sudoku 97.56 (Sudoku +72.44 vs SPG; authors say first to saturate Sudoku zero-shot). FPO+DiPOD also lifts FPO, especially Countdown/Sudoku. A motion-tracking instantiation on Unitree G1/LAFAN improves FPO++ reward and episode length. Code https://github.com/Astro-Eric/DiPOD-release.

Keywords

dipod · diffusion-rl · elbo · dllm · fpo · spg · berkeley

Topics

diffusion RL, diffusion language models, policy gradients

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG). Code https://github.com/Astro-Eric/DiPOD-release (11 stars at check); project https://astro-eric.github.io/blogs/dipod/. Affiliations UC Berkeley, Impossible Inc., NVIDIA (Jiao). Correspondence ericjiang@berkeley.edu. HF has no paper page (404). Uses public GSM8K/MATH500/Countdown/Sudoku and LAFAN; no new standalone corpus, so no datasets_local row. Discord posted abs link.