cs.CLSep 28, 2026

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

Authors: Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang, Xiaomin Li, Yuexing Hao, Yu Hu, Muhao Chen, Varun Chandrasekaran, +2 more

Organizations: University of California, Davis · Microsoft · University of Washington

Abstract

Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

    Jun 30, 2026Yuanda Xu, Zhengze Zhou, Hejian Sang +6Credit AssignmentAgentic Reinforcement Learning

  2. SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning

    Sep 30, 2026Xinchen Du, Zhengze Zhou, Wenhui Zhu +4Agentic Reinforcement LearningCredit Assignment

  3. APPO: Agentic Procedural Policy Optimization

    Jun 10, 2026Xucong Wang, Ziyu Ma, Yong Wang +5Agentic Reinforcement LearningFrictive Policy Optimization