Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.
Figures & tables
Figure 1: Final validation success (%) with Qwen2.5-7B-Instruct. The dashed line marks GRPO. SHARPO improves on GRPO at both mixing weights on both benchmarks, reaching 84.90% on ALFWorld ( +14.32 points) and 75.26% on WebShop ( +9.11 points). Bars are means over three independent runs at training step 200; the vertical axis starts at 50% . Table 1 reports all baselines with standard deviations.
Figure 2: SHARPO Framework. A successful peer supplies context for a policy-copy teacher. A length-normalized likelihood comparison of each selected segment produces a bounded weight. Length balancing and interpolation with the base advantage then determine the segment’s effective advantage for the policy update.
ALFWorld
WebShop
Method
Success rate (%)
Gain
Success rate (%)
Gain
GRPO
70.57±7.83
–
66.15±1.19
–
SDAR
67.19±9.50
−3.39
61.72±6.10
−4.43
RLSD
73.96±5.76
+3.39
69.01±6.31
+2.86
GRPO+OPSD
72.92±4.51
+2.34
69.01±3.16
+2.86
StepOPSD ( λ=0.2 )
77.60±2.96
+7.03
69.27±1.63
+3.12
Table 1: Final validation success rate (%) at step 200 with Qwen2.5-7B-Instruct, mean ± sample standard deviation over three independent runs. Gain is the change in mean success rate relative to GRPO, in percentage points. Red and blue indicate the best and second-best mean results in each column, respectively.
Method
Pick
Look
Clean
Heat
Cool
Pick2
All
GRPO
89.22
73.33
81.39
61.90
58.21
60.06
70.57
SDAR
80.94
82.22
76.55
71.43
54.68
48.00
67.19
RLSD
89.65
74.44
88.49
55.24
55.04
70.71
73.96
GRPO+OPSD
90.95
81.72
77.92
39.05
62.29
66.04
72.92
StepOPSD ( λ=0.2 )
90.96
83.33
83.50
52.86
68.46
75.38
77.60
StepOPSD ( λ=0.05 )
90.32
85.86
88.60
69.52
65.70
64.67
78.12
Table 2: ALFWorld success rate (%) by task family at step 200, mean over three independent runs. Family rates are averaged separately; All reports the mean overall validation success rate from Table 1 . Red and blue indicate the best and second-best mean results in each column, respectively.
Method
Task score
GRPO
0.7781
StepOPSD ( λ=0.2 )
0.8231
StepOPSD ( λ=0.05 )
0.8268
SHARPO ( λ=0.2 )
0.8822
SHARPO ( λ=0.05 )
0.8480
Table 3: WebShop task score. Higher is better. Red and blue mark the best and second-best results among the listed methods.
Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially responsible for their divergent outcomes. Rather than directly trusting the judge's assessment, ProVer verifies the proposed segment by estimating its advantage from the difference in terminal success rates between current-policy continuations sampled before and after the segment. Positive estimates are then incorporated into the GRPO advantages of policy tokens within the proposed segment. By using model judgment only to select where to verify, ProVer grounds local credit in observed outcomes without exhaustively evaluating every intermediate state. Across ALFWorld, WebShop, and SearchQA, ProVer achieves the strongest average performance at both model scales, with relative improvements over GRPO of 9.91% and 7.12% for Qwen3.5-2B and Qwen3.5-4B, respectively. Further analyses demonstrate that informed segment selection improves policy training with modest additional generation overhead, even without a frontier-scale judge model, highlighting the effectiveness and efficiency of selectively targeting pivotal decisions for fine-grained credit assignment in agentic reinforcement learning.
Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang +7
University of California, Davis · Microsoft · University of Washington
Group-based reinforcement learning (RL) methods have achieved remarkable success in improving the performance of large language models (LLMs) and have been rapidly extended to agentic tasks. However, their credit assignment relies heavily on coarse-grained trajectory-level attribution according to final outcomes, making it difficult to capture the contribution of individual steps, such as valuable steps obscured within failed trajectories. To uncover latent information and enable more faithful step-level credit assignment, we propose Graph-based Group Policy Optimization (GraphGPO), which first aggregates all rollout trajectories into a unified state-transition graph and then estimates the distance from each state to the task goal using the global information encoded in the graph. Finally, GraphGPO assigns credit to each edge by estimating a graph-based advantage, based on how much the transition reduces the distance to the task goal. In this way, GraphGPO significantly improves training efficiency and achieves state-of-the-art performance across a range of challenging benchmarks.
Xin Cheng, Shuo He, Lang Feng +4
Nanyang Technological University, Singapore · Tongyi Lab, Alibaba Group · Southeast University, China
Group-based reinforcement learning effectively post-trains LLM agents for long-horizon, sparse-reward tasks by deriving step-level credit from trajectory outcomes. However, this ties a step's credit to its rollout's final outcome: semantically near-identical intermediate steps receive opposite credit depending on whether their trajectory eventually succeeded or failed. Such semantic credit inconsistency sends conflicting gradients to similar actions and wastes the partially-correct progress inside failed rollouts. Motivated by this, we propose Semantic Consistency Policy Optimization (SCPO), a value-free reward-shaping method that mitigates this inconsistency by recovering step-level credit from successful siblings in the same rollout group. Concretely, SCPO scores each failed step against a successful sibling and adds positive step-level credit for new progress along that sibling. On ALFWorld and WebShop, SCPO matches or exceeds strong group-based baselines, reaching 93.7+/-4.1 percent success on ALFWorld and 74.8+/-2.0 percent on WebShop at 1.5B parameters, with gains concentrated on the hardest multi-step tasks.
Peng Xu, Sijia Chen, Junzhuo Li +1
†The Hong Kong University of Science and Technology (Guangzhou) · ‡The Hong Kong University of Science and Technology