Vector Policy Optimization: Training for Diversity Improves Test-Time Search
- Type
- paper
- Venue
- arXiv preprint
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Proposes Vector Policy Optimization (VPO), which replaces the GRPO advantage estimator and optimizes policies against vector-valued rewards so that different outputs specialize to different reward tradeoffs, improving test-time search.
Research notes
- Key findings: Matches or beats scalar RL baselines on pass@k/best@k over four tasks; Advantage grows with larger search budgets; Enables evolutionary search solutions that GRPO models could not solve
- Submitted 2026-05-21; arXiv subjects include cs.LG.