Organizations: Beihang University · Zhongguancun Academy · Communication University of China · Peking University · Hangzhou Innovation Institute of Beihang University
Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should incorporate evidence beyond the realized suffixes observed at an anchor while aggregating alternative continuations according to their empirical frequencies. Visit-local averaging pools realized suffix returns at shared anchors and respects observed frequencies, but does not recursively propagate evidence across rollouts, whereas shortest-path estimators have global reach but allow a rarely observed route to dominate an anchor's value. We introduce Cross-Rollout Bellman Closure (CRBC), which merges each rollout group into a finite empirical process with absorbing success and failure boundaries and evaluates its behavior-policy Bellman fixed point with one linear solve. This fixed point uses the same empirical action and transition frequencies to propagate evidence through shared anchors and aggregate alternative continuations. Backing up the resulting state values through observed transitions yields action values, whose gain over the corresponding state value provides step-level credit. A corresponding finite-depth family recovers visit-local return averaging at zero depth and converges to the exact closure as depth increases. The normalized closure credit is combined with the trajectory-level group advantage for policy optimization, without additional environment rollouts. Across ALFWorld, WebShop, and Sokoban benchmarks with multiple model scales, CRBC consistently improves final performance and learning efficiency. For example, CRBC outperforms the strongest evaluated baseline by 5.59 percentage points on ALFWorld with Qwen2.5-1.5B-Instruct.
Figures & tables
Figure 1: Reach and aggregation for step-level credit. For the rollout trajectory τ , node s denotes an anchor : an environment-relevant state representation that may be shared by multiple visits, and the edge represents an observed transition after executing semantic action a . (a) Shared anchors merge rollouts into a finite empirical process. (b) Shortest-path estimation aggregates by an extremum, ignoring transition frequencies; a rarely executed continuation can dominate anchor value. (c) The finite-depth family on this process, illustrated at s2 , uses behaviour-consistent expectations: K=0 pools visit-local suffix returns; K=1 adds one cross-rollout backup; and K→∞ evaluates all supported continuations at the behaviour-policy Bellman fixed point (CRBC), giving global reach.
Figure 2: Overview of CRBC. (a) Shared anchors merge rollouts into an empirical process with absorbing success and failure boundaries. (b) A Bellman fixed-point solve yields the state value V , followed by action backups to obtain the action value Q and normalized Q−V credit, which gives relative action credit. (c) Step-level credit and trajectory-level group-relative advantages are combined for policy optimization.
ALFWorld
WebShop
Type
Method
Pick
Clean
Cool
Look
Heat
Pick2
All
Score
Succ.
Closed-Source Models
Prompting
GPT-4o
75.30
60.80
31.20
56.70
21.60
49.80
48.00
31.80
23.70
Prompting
Gemini-2.5-Pro
92.80
63.30
62.10
69.00
26.60
58.70
60.30
42.50
35.90
Qwen2.5-1.5B-Instruct
Prompting
Qwen2.5
5.90
5.50
3.30
9.70
4.20
0.00
4.10
23.10
5.20
Table 1: Performance comparison on ALFWorld and WebShop. For ALFWorld, we report the average success rate (%) for each subtask and the overall success rate. For WebShop, we report the average task score and average success rate (%). Most results are averaged over three random seeds. The best results are highlighted in bold .
Depth K
Step-credit
Score
Success (%)
0
A(0)
7.00 (0.27)
92.97 (0.00)
1
A(1)
7.05 (0.53)
92.97 (1.56)
2
A(2)
7.16 (1.20)
93.75 (3.12)
3
A(3)
7.21 (0.13)
93.75 (0.78)
∞
A(∞)
8.01 (0.27)
96.09 (0.64)
Table 2: Closure-depth ablation on ALFWorld with Qwen2.5-1.5B-Instruct. Only K varies; all other settings are fixed. Best values are in bold .
wG
wS
Credit assignment
Pick
Clean
Cool
Look
Heat
Pick2
All
0
5
Step-level only
98.89 (1.57)
98.25 (2.48)
93.54 (1.74)
94.44 (3.93)
88.89 (2.24)
96.67 (2.36)
95.31 (1.69)
1
0
Trajectory-level only
84.62 (1.46)
64.91 (2.48)
86.92 (6.79)
58.33 (0.00)
77.78 (5.94)
73.33 (4.71)
76.56 (2.30)
1
1
Step + trajectory
98.89 (1.57)
87.72 (2.48)
94.82 (1.78)
94.44 (3.93)
88.89 (2.24)
93.33 (2.36)
93.75 (1.69)
1
2.5
Step + trajectory
98.89 (1.57)
92.98 (4.96)
97.38 (1.85)
86.11 (3.93)
92.06 (2.24)
93.33 (2.36)
94.53 (1.91)
[6pt][6pt] 1
5
Step + trajectory (Ours)
100.00 (0.00)
92.98 (6.56)
92.15 (3.33)
91.67 (0.00)
96.83 (4.49)
100.00 (0.00)
96.09 (0.64)
1
10
Step + trajectory
98.89 (1.57)
92.98 (2.48)
94.87 (3.63)
91.67 (0.00)
88.89 (2.24)
98.33 (2.36)
94.79 (1.95)
Table 4: Credit and weight ablations on ALFWorld with Qwen2.5-1.5B-Instruct. We report final validation success rates (%) for six subtasks and the overall success rate (All) at step 150, averaged over three random seeds. Best values are in bold .
Figure 3: Training reward curves for CRBC (Ours, red), GraphGPO (blue), GiGPO (green), and GRPO (gray) on ALFWorld, WebShop, and Sokoban. Light and dark curves denote per-update and EMA-smoothed rewards ( α=0.95 ), respectively. Qwen2.5-1.5B-Instruct is used for ALFWorld/WebShop and Qwen2.5-VL-3B-Instruct for Sokoban; panel scales are independent.
Figure 4: Per-update runtime breakdown of CRBC on ALFWorld with Qwen2.5-1.5B-Instruct. Blue bars denote shared stages, while red bars denote trajectory-level and step-level credit construction.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
ALFWorld
WebShop
Sokoban
Maximum environment steps
50
15
15
Maximum prompt tokens
2048
5120
1024
Maximum response tokens
512
512
512
Training groups per update
16
16
32
Rollouts per group
8
8
8
Validation trajectories
128
128
128
Appendix
Table 5: Environment-level training settings used in our reproduced experiments. Settings are shared across methods within each model scale and environment.
Figure 5: Cross-rollout evidence coverage measured on rollouts from an early-training checkpoint on ALFWorld. Left: distribution of the number of reachable rollouts at each closure depth K . Right: paired coverage at K=0 and full closure, with color indicating the percentage of all eligible visits.
Bellman discount β
Pick
Clean
Cool
Look
Heat
Pick2
All
0.95
98.89 (1.57)
98.25 (2.48)
94.82 (1.78)
91.67 (0.00)
88.89 (2.24)
98.33 (2.36)
95.57 (0.37)
0.98
100.00 (0.00)
92.98 (6.56)
92.15 (3.33)
91.67 (0.00)
96.83 (4.49)
100.00 (0.00)
96.09 (0.64)
0.99
100.00 (0.00)
86.84 (2.63)
94.23 (5.77)
95.83 (4.17)
92.86 (2.38)
100.00 (0.00)
95.31 (1.56)
1.00
100.00 (0.00)
89.47 (10.53)
92.23 (3.77)
83.33 (8.33)
95.24 (4.76)
92.50 (2.50)
93.36 (0.39)
Appendix
Table 6: Bellman-discount ablation on ALFWorld with Qwen2.5-1.5B-Instruct. Only β varies; all other CRBC settings are fixed. Results are averaged over three random seeds; best values are in bold .
Group size G
Method
Pick
Clean
Cool
Look
Heat
Pick2
All
Score
4
GRPO
60.00
36.84
46.15
50.00
38.10
50.00
47.66
2.27
4
GiGPO
90.00
63.16
88.46
66.67
80.95
75.00
79.69
5.34
4
GraphGPO
86.67
68.42
84.62
58.33
76.19
75.00
77.34
4.38
4
CRBC (Ours)
100.00
94.74
88.46
91.67
100.00
80.00
92.97
7.12
16
GRPO
93.33
73.68
80.77
58.33
85.71
80.00
81.25
5.08
16
GiGPO
100.00
89.47
88.46
83.33
100.00
90.00
92.97
7.01
Appendix
Table 7: Rollout group size comparison on ALFWorld with Qwen2.5-1.5B-Instruct at step 150. We report validation success rates (%) for six subtasks and overall (All), together with task score. Best reported values within each group size are in bold .
Figure 6: Validation success-rate curves for CRBC (Ours, red), GraphGPO (purple), GiGPO (green), and GRPO (gray) on ALFWorld, WebShop, and Sokoban. Light curves show validation measurements recorded every five training updates, while dark curves show EMA-smoothed trends ( α=0.95 per update). Qwen2.5-1.5B-Instruct is used for ALFWorld and WebShop, and Qwen2.5-VL-3B-Instruct for Sokoban.
Figure 7: Training dynamics for the closure-depth ablation on ALFWorld with Qwen2.5-1.5B-Instruct. Only the closure depth K varies; all other settings are fixed. From left to right, we show episode reward, validation success rate, and validation score. Light and dark curves denote raw and EMA-smoothed trajectories ( α=0.95 ); deeper closure gives stronger late-training performance.
Figure 8: Prompt templates for ALFWorld, WebShop, and Sokoban. Coloured placeholders denote runtime inputs. Text environments retain up to two observation–action records; Sokoban uses the current RGB observation without textual interaction history. Panel (c) includes an example RGB input.
Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL frameworks have shifted from trajectory-level to step-level training. However, long-horizon agentic RL suffers from severe reward sparsity and delay, as feedback is often deferred for dozens of interaction steps. While existing step-level frameworks refine training granularity, their credit assignment remains coarse-grained and still treats agent exploration as isolated, linear trajectories. This oversimplified perspective ignores the inherent graph structure of state transitions, leading to high-variance state-value estimation and myopic, localized credit assignment. To overcome these critical bottlenecks, we propose Group-Graph Policy Optimization (G2PO), a novel group-based RL algorithm tailored for multi-turn agentic tasks. G2PO explicitly transforms linear interaction trajectories into a global state-transition graph. By aggregating identical observations across different trajectories, we introduce group-aggregation state-value estimation that reduces sampling variance and trajectory-dependent bias. Furthermore, we redefine agent actions as transitions between state nodes and propose an edge-centric advantage estimation strategy. By globally standardizing Temporal Difference (TD) errors across the entire graph, G2PO explicitly identifies and prioritizes critical transitions that drive absolute task progress. Extensive experiments on representative long-horizon benchmarks-WebShop, ALFWorld, and AppWorld-demonstrate that G2PO substantially outperforms state-of-the-art prompt-based and RL baselines, achieving remarkable success rate improvements of up to 22.2% over GRPO.
Group-relative RL training (GRPO) samples a small group of parallel rollouts for every training prompt and uses their within-group reward spread to compute per-trajectory advantages. In agentic environments each rollout is a long multi-turn dialogue with one LLM call per step, so this multi-sample multiplier dominates the total training cost. When every rollout of a prompt ends with the same reward, the group has zero reward variance and contributes no gradient, so the extra rollouts add no information; such groups are common in practice (typically around 40% of all groups), so the wasted-compute fraction is substantial rather than marginal. Existing methods filter such groups at the prompt level, either after their rollouts are paid for or before any rollout begins, but both decide without using information that becomes available during the rollout itself. We instead ask whether the in-group divergence between the partial trajectories at an intermediate step can already predict that the group will be zero-variance: when the parallel rollouts have already converged on the same action prefix, the group is on track to produce a single reward, and we can stop early. We propose a one-parameter gate that stops a group when the mean pairwise prefix edit distance between its partial action sequences falls below a threshold. On a 60-iteration on-policy GRPO run on ALFWorld with Qwen2.5-7B, averaged over four random seeds, the gated arm finishes 10.7% faster in wall-clock (bootstrap 95% CI excludes 0) and shifts held-out success rate on 50 unseen tasks by +2.5 pp, with the held-out gain tracing to a measurable reduction in zero-advantage gradient-batch dilution. Code is available at https://github.com/zhiyuanZhai20/selective-rollout.
Training large language model agents in long-horizon environments requires assigning credit from sparse terminal outcomes to individual actions. Existing critic-free methods propagate trajectory-level rewards uniformly across steps, while recent approaches construct step-level groups by matching repeated states and compare actions within each group. The former cannot distinguish useful actions in failed trajectories from ineffective actions in successful ones. The latter rely on step credit derived directly from individual trajectory outcomes and fixed-weight fusion with episode-level credit. We propose Gated-BEPO, which derives step-level credit from empirical rollout graphs. For each rollout group, Gated-BEPO constructs an empirical graph and estimates node values through a mean-backup Bellman fixed point that reflects the empirical action distribution of the current policy. We then accumulate these temporal-difference residuals along each sampled trajectory using generalized advantage estimation, yielding step-level Bellman advantages that capture both immediate and downstream effects. To adaptively fuse episode- and step-level credit, a confidence gate incorporates Bellman credit only at states with multiple observed successors and otherwise uses episode-level credit. Experiments on WebShop, ALFWorld, and visual Sokoban show consistent improvements across language and vision-language models, while diagnostic ablations support the effectiveness of Bellman fixed-point value estimation and show that step-level credit should be incorporated selectively rather than uniformly into the final advantage.
Hongxi Yan, Ziyue Huang, Shichao Fan +1
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing, China · Zhongguancun Laboratory, Beijing, China · Qualcomm, Beijing, China