Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
Figures & tables
Method
ALFWorld
WebShop
SearchQA
Avg.
Qwen3.5-2B
GRPO
84.08 ± 3.02
42.53 ± 0.31
35.58 ± 0.38
54.07 ± 0.78
Budget-Matched GRPO
84.05 ± 4.13
39.00 ± 2.62
37.75 ± 0.25
53.60 ± 1.32
GiGPO
85.82 ± 2.24
44.60 ± 0.35
36.17 ± 1.66
55.53 ± 1.25
SPO-tree
56.96 ± 1.14
38.47 ± 0.50
31.67 ± 0.72
42.36 ± 0.27
SPO-chain
75.87 ± 1.72
47.80 ± 2.62
27.67 ± 0.76
50.45 ± 1.20
Table 1: Performance of the baselines and ProVer with Qwen3.5-2B and Qwen3.5-4B on three benchmarks. We report the mean and sample standard deviation over three evaluation runs.
Method
Generated Tokens / Step (K)
Judge Cost / Step (¢)
Time / Step (min)
ALF
WS
SQA
ALF
WS
SQA
ALF
WS
SQA
GRPO
237.0
97.1
37.6
–
–
–
3.39
0.82
1.39
SPO-tree
153.2
71.0
63.4
–
–
–
7.18
4.85
4.65
SPO-chain
421.2
172.3
72.3
–
–
–
6.73
5.07
4.77
CriticSearch
227.1
136.3
45.8
17.31
8.54
7.40
4.64
3.15
3.25
ProVer
242.6
108.4
43.9
4.00
10.81
5.69
3.73
2.00
2.25
Table 2: Per-step training cost for Qwen3.5-4B on ALFWorld (ALF), WebShop (WS), and SearchQA (SQA). Generated tokens count policy-rollout outputs. Judge cost is calculated from recorded gpt-5.4-mini API usage using OpenAI’s published pricing as of August 2026. Time is the mean wall-clock time per optimizer step in minutes.
Agentic Judge
Policy Performance (%)
SearchQA Proposal Diagnostics
ALFWorld
WebShop
SearchQA
Accept (%)
Mean Len.
Random
94.0
42.4
42.8
19.5
2.38
Qwen3.5-9B
98.5
44.2
45.8
71.2
1.45
gpt-5.4-nano
94.8
46.2
43.5
72.4
1.25
gpt-5.4-mini
98.3
47.9
47.0
62.5
1.18
gpt-5.4
96.3
47.2
43.8
58.5
1.03
Table 3: Effect of segment selection and judge model choice on Qwen3.5-4B policy training. Random replaces judge-guided selection with random segment sampling while retaining the same outcome-verification and credit-assignment procedure. We additionally report proposal diagnostics on SearchQA: acceptance rate ( Δseg>0 ) and mean segment length in agent turns.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Model
Benchmark
Phase 1 steps
Rollouts/group
Phase 2 steps
Rollouts/group
Qwen3.5-2B
ALFWorld
1–84
11
85–100
12
Qwen3.5-2B
WebShop
1–94
10
95–100
11
Qwen3.5-2B
SearchQA
1–70
10
71–100
11
Qwen3.5-4B
ALFWorld
1–26
8
27–100
9
Qwen3.5-4B
WebShop
1–35
10
36–100
11
Qwen3.5-4B
SearchQA
1–36
9
37–100
10
Appendix
Table 4: Number of trajectories sampled per group in each phase of Budget-Matched GRPO. Every update contains 16 groups.