Diffusion Reward Models
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Recasts reward modeling as conditional density estimation over p(r|x,y): a lightweight Diffusion Transformer conditioned on a frozen LLM encoder denoises Gaussian noise into a reward vector, making no parametric assumption on the output distribution. A single architecture handles multi-attribute regression and pairwise preference data; N samples at inference form an empirical reward distribution aggregated as scalar, variance, or quantiles.
Keywords
reward models · diffusion · RLHF · uncertainty
Topics
reward models, diffusion, RLHF, uncertainty
Research notes
- Discovery: Shared in #random-papers as an X post (Sep 29, 8:03 AM); resolved to arXiv 2609.33803.
- Method: Diffusion Reward Model: conditional density estimation with a lightweight Diffusion Transformer conditioned on a frozen LLM encoder; uncertainty-aware rejection and lower-confidence-bound (LCB) aggregation.
- Key findings: Matches or surpasses baselines across five benchmarks under matched data and backbone; stays competitive with much larger discriminative, distributional, and generative RMs at modest training scale; recovers multimodal reward structure where conventional heads collapse to a point; downstream RLHF with DRM as the training-time reward improves policy performance.
- arXiv ID resolved via arXiv API query (2609.33803). Author-name boundaries taken from the API rendering; verify 'Jiaze Wang' / 'Ziqing Qiao' split before publishing.