← Back to explorer

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Type
paper
Venue
arXiv:2609.36178 (cs.CL), submitted 28 Sep 2026
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

GRPO gives every token in a trajectory the same advantage, so the training signal cannot distinguish decisive steps from the rest. ProVer targets potentially pivotal decisions for fine-grained credit assignment: given a rollout group, an agentic judge contrasts successful and failed trajectories to propose the segment responsible for their divergent outcomes; rather than trusting the judge directly, ProVer verifies the segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment, folding positive estimates into the GRPO advantages of policy tokens within the segment. Model judgment is used only to select where to verify, grounding local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B; informed segment selection works with modest additional generation overhead even without a frontier-scale judge.

Keywords

ProVer · credit assignment · agent RL · GRPO · LLM judge · pivotal decisions · ALFWorld · WebShop · SearchQA

Topics

agent RL, credit assignment, GRPO, LLM judge, pivotal decisions, agentic reasoning

Research notes

  • Discovery: @omarsar0 (dair_ai) X post 2026-10-02 (https://x.com/omarsar0/status/2105930871534690714)
  • Chat-with-paper page: https://academy.dair.ai/papers/targeting-pivotal-decisions-for-credit-assignment-in-agentic-reinforcement-learn-2609.36178
  • Method: judge compares successful vs failed rollouts to name the segment causing the difference, then samples continuations from just before and just after that segment; the change in success rate becomes the segment's advantage
  • Works even when the judge is a smaller model
  • Affiliations on paper: UC Davis, Microsoft, University of Washington, Purdue
  • License: CC BY 4.0