← Back to explorer

Vector Policy Optimization: Training for Diversity Improves Test-Time Search

Type
paper
Venue
arXiv / MIT Improbable AI Lab
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T20:26:00Z
Verified
2026-08-14T20:26:00Z

Summary

MIT Improbable AI / Sakana (cite Bahlous-Boldi et al.; arXiv 2605.22817). Argues scalar GRPO collapses entropy so extra samples become near-duplicates; when test-time search (best@k, AlphaEvolve) handles exploitation, training should preserve reward diversity. VPO: emit m answers in one chain; score the set as E_w[max_y w·r] with w~Dirichlet(1); drop-in GRPO advantage on the set reward. Beats GRPO/Multi-RLVR/Max-at-k/MaxRL/random-w/goal-conditioned GRPO on Maze, MuSiQue, EUREQA, ToolRL as k grows. LiveCodeBench Qwen2.5-Coder-7B: GRPO wins pass@1 but VPO wins best@k and OpenEvolve on 32 hard problems GRPO never solves. Helps when reward components are non-collinear; UltraFeedback ArmoRM dims are near-collinear and VPO does not beat scalar GRPO.

Keywords

vpo · grpo · diversity · test-time-search · pass-at-k · alphaevolve · mit

Topics

RL post-training, diversity, test-time search

Research notes

  • Primary: arxiv abs 2605.22817 (cs.LG; 24 pages). No official GitHub on abs/HF. Discord posted abs only. License left blank (no CC on abs). Algorithm paper, not a hosted corpus.