Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
Figures & tables
Method
ALFWorld
WebShop
SearchQA
Avg.
Qwen3.5-2B
GRPO
84.08 ± 3.02
42.53 ± 0.31
35.58 ± 0.38
54.07 ± 0.78
Budget-Matched GRPO
84.05 ± 4.13
39.00 ± 2.62
37.75 ± 0.25
53.60 ± 1.32
GiGPO
85.82 ± 2.24
44.60 ± 0.35
36.17 ± 1.66
55.53 ± 1.25
SPO-tree
56.96 ± 1.14
38.47 ± 0.50
31.67 ± 0.72
42.36 ± 0.27
SPO-chain
75.87 ± 1.72
47.80 ± 2.62
27.67 ± 0.76
50.45 ± 1.20
Table 1: Performance of the baselines and ProVer with Qwen3.5-2B and Qwen3.5-4B on three benchmarks. We report the mean and sample standard deviation over three evaluation runs.
Method
Generated Tokens / Step (K)
Judge Cost / Step (¢)
Time / Step (min)
ALF
WS
SQA
ALF
WS
SQA
ALF
WS
SQA
GRPO
237.0
97.1
37.6
–
–
–
3.39
0.82
1.39
SPO-tree
153.2
71.0
63.4
–
–
–
7.18
4.85
4.65
SPO-chain
421.2
172.3
72.3
–
–
–
6.73
5.07
4.77
CriticSearch
227.1
136.3
45.8
17.31
8.54
7.40
4.64
3.15
3.25
ProVer
242.6
108.4
43.9
4.00
10.81
5.69
3.73
2.00
2.25
Table 2: Per-step training cost for Qwen3.5-4B on ALFWorld (ALF), WebShop (WS), and SearchQA (SQA). Generated tokens count policy-rollout outputs. Judge cost is calculated from recorded gpt-5.4-mini API usage using OpenAI’s published pricing as of August 2026. Time is the mean wall-clock time per optimizer step in minutes.
Agentic Judge
Policy Performance (%)
SearchQA Proposal Diagnostics
ALFWorld
WebShop
SearchQA
Accept (%)
Mean Len.
Random
94.0
42.4
42.8
19.5
2.38
Qwen3.5-9B
98.5
44.2
45.8
71.2
1.45
gpt-5.4-nano
94.8
46.2
43.5
72.4
1.25
gpt-5.4-mini
98.3
47.9
47.0
62.5
1.18
gpt-5.4
96.3
47.2
43.8
58.5
1.03
Table 3: Effect of segment selection and judge model choice on Qwen3.5-4B policy training. Random replaces judge-guided selection with random segment sampling while retaining the same outcome-verification and credit-assignment procedure. We additionally report proposal diagnostics on SearchQA: acceptance rate ( Δseg>0 ) and mean segment length in agent turns.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Model
Benchmark
Phase 1 steps
Rollouts/group
Phase 2 steps
Rollouts/group
Qwen3.5-2B
ALFWorld
1–84
11
85–100
12
Qwen3.5-2B
WebShop
1–94
10
95–100
11
Qwen3.5-2B
SearchQA
1–70
10
71–100
11
Qwen3.5-4B
ALFWorld
1–26
8
27–100
9
Qwen3.5-4B
WebShop
1–35
10
36–100
11
Qwen3.5-4B
SearchQA
1–36
9
37–100
10
Appendix
Table 4: Number of trajectories sampled per group in each phase of Budget-Matched GRPO. Every update contains 16 groups.
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advantage over all action tokens. This outcome signal is useful but structurally incomplete: it punishes useful exploration in failed rollouts and reinforces redundant or regressive actions in successful rollouts. We propose TRIAGE, a role-typed credit assignment framework that adds a semantic role axis to outcome credit. A structured judge classifies each segment as decisive progress, useful exploration, no-progress infrastructure, or regression, and a fixed role-conditioned rule maps these labels to bounded segment-level process rewards. This keeps verifier outcomes as the source of optimization direction while correcting the two main blind spots of outcome-only credit. We further show that role-conditioned credit is the optimal segment-level correction expressible from role labels alone -- a projection of the per-segment advantage residual onto the role variable -- so that the fixed role constants reduce advantage estimation error whenever the judge is reliable, and we connect this to lower-variance policy gradients. Across ALFWorld, Search-QA, and WebShop, TRIAGE improves success rates over GRPO for two policy models and outperforms both a scalar judge-derived process reward and an outcome-supervised shared-backbone value baseline. Ablations show that the gain comes from role typing rather than merely adding dense rewards: reliable detection of regression inside successful trajectories is the dominant contributor, while exploration credit provides a consistent secondary gain; on completed ALFWorld and WebShop rollouts, TRIAGE also reduces environment-facing turns by an additional 10.4% and 14.8% relative to GRPO.
Yuanda Xu, Zhengze Zhou, Hejian Sang +6
1LinkedIn Corporation · 2Harvard University · 3Johns Hopkins University +1
Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.
Xinchen Du, Zhengze Zhou, Wenhui Zhu +4
LinkedIn Corporation · Georgia Institute of Technology
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit over coarse heuristic units, such as tool-call boundaries or fixed workflows, making it difficult to identify which intermediate decisions influence downstream outcomes. In this work, we study agentic RL from two perspectives: \textit{where to branch and how to assign credit after branching}. Our pilot analysis shows that influential decision points are broadly distributed throughout the generated sequence rather than concentrated at tool calls, while token entropy alone does not reliably reflect their impact on final outcomes. Motivated by these observations, we propose \textbf{Agentic Procedural Policy Optimization (APPO)}, which shifts branching and credit assignment from coarse interaction units to fine-grained decision points in the sequence. APPO selects branching locations using a Branching Score that combines token uncertainty with policy-induced likelihood gains of subsequent continuations, enabling more targeted exploration while filtering out spurious high-entropy positions. It further introduces procedure-level advantage scaling to better distribute credit across branched rollouts. Experiments on 13 benchmarks show that APPO consistently improves strong agentic RL baselines by nearly 4 points, while keeping efficient tool-calls and maintaining behavior interpretability.
Xucong Wang, Ziyu Ma, Yong Wang +5
University of Science and Technology of China · AMAP, Alibaba Group · Southern University of Science and Technology