Vector Policy Optimization: Training for Diversity Improves Test-Time Search
- Type
- paper
- Venue
- arXiv / MIT Improbable AI Lab
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T20:26:00Z
- Verified
- 2026-08-14T20:26:00Z
Summary
MIT Improbable AI / Sakana (cite Bahlous-Boldi et al.; arXiv 2605.22817). Argues scalar GRPO collapses entropy so extra samples become near-duplicates; when test-time search (best@k, AlphaEvolve) handles exploitation, training should preserve reward diversity. VPO: emit m answers in one chain; score the set as E_w[max_y w·r] with w~Dirichlet(1); drop-in GRPO advantage on the set reward. Beats GRPO/Multi-RLVR/Max-at-k/MaxRL/random-w/goal-conditioned GRPO on Maze, MuSiQue, EUREQA, ToolRL as k grows. LiveCodeBench Qwen2.5-Coder-7B: GRPO wins pass@1 but VPO wins best@k and OpenEvolve on 32 hard problems GRPO never solves. Helps when reward components are non-collinear; UltraFeedback ArmoRM dims are near-collinear and VPO does not beat scalar GRPO.
Keywords
vpo · grpo · diversity · test-time-search · pass-at-k · alphaevolve · mit
Topics
RL post-training, diversity, test-time search
Research notes
- Primary: arxiv abs 2605.22817 (cs.LG; 24 pages). No official GitHub on abs/HF. Discord posted abs only. License left blank (no CC on abs). Algorithm paper, not a hosted corpus.