Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
- Type
- paper
- Venue
- arXiv:2609.36178 (cs.CL), submitted 28 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
GRPO gives every token in a trajectory the same advantage, so the training signal cannot distinguish decisive steps from the rest. ProVer targets potentially pivotal decisions for fine-grained credit assignment: given a rollout group, an agentic judge contrasts successful and failed trajectories to propose the segment responsible for their divergent outcomes; rather than trusting the judge directly, ProVer verifies the segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment, folding positive estimates into the GRPO advantages of policy tokens within the segment. Model judgment is used only to select where to verify, grounding local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B; informed segment selection works with modest additional generation overhead even without a frontier-scale judge.
Keywords
ProVer · credit assignment · agent RL · GRPO · LLM judge · pivotal decisions · ALFWorld · WebShop · SearchQA
Topics
agent RL, credit assignment, GRPO, LLM judge, pivotal decisions, agentic reasoning
Research notes
- Discovery: @omarsar0 (dair_ai) X post 2026-10-02 (https://x.com/omarsar0/status/2105930871534690714)
- Chat-with-paper page: https://academy.dair.ai/papers/targeting-pivotal-decisions-for-credit-assignment-in-agentic-reinforcement-learn-2609.36178
- Method: judge compares successful vs failed rollouts to name the segment causing the difference, then samples continuations from just before and just after that segment; the change in success rate becomes the segment's advantage
- Works even when the judge is a smaller model
- Affiliations on paper: UC Davis, Microsoft, University of Washington, Purdue
- License: CC BY 4.0