← Back to explorer

Vector Policy Optimization: Training for Diversity Improves Test-Time Search

Type
paper
Venue
arXiv preprint
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Proposes Vector Policy Optimization (VPO), which replaces the GRPO advantage estimator and optimizes policies against vector-valued rewards so that different outputs specialize to different reward tradeoffs, improving test-time search.

Research notes

  • Key findings: Matches or beats scalar RL baselines on pass@k/best@k over four tasks; Advantage grows with larger search budgets; Enables evolutionary search solutions that GRPO models could not solve
  • Submitted 2026-05-21; arXiv subjects include cs.LG.