Long-horizon LLM agents require reinforcement learning methods that can assign credit to intermediate decisions under sparse and delayed rewards. Existing group-based methods such as GRPO and GiGPO alleviate this issue by comparing rollout returns or repeated anchor states, but they still fail when the compared returns have no variation. We identify this failure mode as zero-credit failure: during early training, many failed rollouts contain useful prefixes, yet existing methods assign them no task-discriminative advantage. To address this issue, we propose Milestone Viability Potential Policy Optimization (MVPO), a potential-routed policy optimization algorithm that learns from viable failure prefixes. MVPO estimates prefix potential over Union-Find viability regions, repairs zero-credit groups with potential-difference advantages, and attenuates the potential branch according to relative performance progress. Experiments with Qwen2.5-1.5B-Instruct show that MVPO outperforms eight strong baselines, including GRPO and GiGPO. Under the same training length, MVPO improves over the GiGPO baseline by +4.4 success points on ALFWorld and +5.3 on WebShop, while adding only 0.16%-0.20% advantage-construction overhead.
Figures & tables
Figure 1: Motivation and overview of Mvpo . Top : an ALFWorld case of zero-credit failure : all rollouts fail, making GRPO/GiGPO assign no task-discriminative advantage, while Mvpo identifies viable prefixes using prefix potential. Bottom (a) : zero-advantage steps affect over 90% of GRPO steps and about 65% of GiGPO steps in early WebShop training. Bottom (b) : higher prefix potential predicts higher future success on ALFWorld. Bottom (c) : Mvpo reaches 40% success earlier during training. Bottom (d) : Mvpo improves final success over GiGPO on both benchmarks.
Figure 2: Overview of Mvpo . Mvpo first estimates prefix potential over Union-Find viability regions using future success, loop statistics, and adaptive milestone progress (§ 4.1 ). It then routes step-level credit according to anchor-group availability: informative anchor groups use anchor-based comparative credit, while zero-credit groups are repaired with prefix-potential advantages (§ 4.2 ). Finally, the routed step-level credit is combined with the episode-level group advantage for critic-free policy optimization (§ 4.3 ).
Method
ALFWorld
WebShop
Pick
Clean
Cool
Look
Heat
Pick2
All
Score
Success
Closed-source LLMs
GPT 5.4
88.3 ±2.4
52.3 ±8.6
56.9 ±2.0
66.7 ±8.6
70.8 ±20.6
56.3 ±8.4
64.1 ±3.4
9.3 ±1.1
7.0 ±0.6
Claude Opus 4.7
92.8 ±1.9
81.7 ±3.2
73.6 ±2.0
72.7 ±0.0
72.9 ±10.6
66.2 ±0.6
77.1 ±1.0
23.6 ±6.1
19.8 ±4.8
Qwen2.5-1.5B-Instruct
Prompting
5.9
3.3
4.2
5.5
9.7
0.0
4.1
23.1
5.2
Table 1: Main results on ALFWorld and WebShop. For ALFWorld, we report success rates for six subtasks and the overall success rate. For WebShop, we report the average task score and success rate. Results are averaged over 3 random seeds. The best result of each group is in bold and the second-best result is underlined . GiGPO w/std uses Fnorm=std , while GiGPO w/ostd uses Fnorm=1 . Mvpo uses the w/std setting.
Method
Anchor
Potential
κ Decay
Success
Score
GiGPO w/std
✓
–
–
65.0 ±3.2
83.1 ±1.6
w/o GiGPO anchor
–
✓
✓
63.5 ±2.1
79.8 ±0.7
w/o κ decay
✓
✓
–
60.4 ±4.1
83.4 ±2.4
Mvpo
✓
✓
✓
70.3 ±1.9
84.6 ±1.0
Table 2: Ablation study on WebShop with Qwen2.5-1.5B-Instruct. We report success rate and task score, averaged over 3 random seeds. “Anchor” denotes GiGPO anchor credit, “Potential” denotes MVPO prefix-potential credit, and “ κ Decay” denotes performance-progress attenuation. Best results are bolded .
Env.
Method
Gen.
Logp+Ref
Adv.
Actor
Total
ALFWorld
GRPO
183.3
19.5
0.5
35.5
283.4
GiGPO
207.9
16.6
0.7
30.7
303.9
Mvpo
227.5
20.4
1.2
35.7
340.0
WebShop
GRPO
60.0
10.7
0.1
19.6
108.1
GiGPO
58.8
8.0
0.1
14.7
100.0
Mvpo
61.2
9.4
0.3
17.3
107.9
Table 3: Per-step training time. Gen. denotes rollout generation, Logp+Ref denotes old-policy and reference-policy log-probability computation, Adv. denotes advantage construction, and Actor denotes policy update.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Training dynamics of Mvpo . (a) The performance-progress attenuation mechanism works as intended. At the beginning of training, the EMA success rate is low and κt stays close to 1, allowing prefix-potential credit to repair zero-credit groups with full strength. As the policy improves, the EMA success rate increases and κt decreases accordingly, causing the potential branch to fade and leaving more gradient space to anchor-based comparative credit. (b) The learned prefix potential remains directionally meaningful throughout training. Successful trajectories maintain higher mean potential than failed trajectories, while the Union-Find region statistics remain stable. This supports the central assumption of Mvpo : potential can distinguish viable failure prefixes from unproductive prefixes before final rewards become frequent.
Method
Single-Hop QA
Multi-Hop QA
Avg.
NQ †
TriviaQA ∗
PopQA ∗
HotpotQA †
2Wiki ∗
MuSiQue ∗
Bamboogle ∗
R1-Instruct
27.0
53.7
19.9
23.7
29.2
7.2
29.3
27.1
Search-R1
34.1
54.5
37.8
32.4
31.9
10.3
26.4
32.5
ZeroSearch
41.4
57.4
44.8
27.4
30.0
9.8
11.1
31.7
StepSearch
–
–
–
34.5
32.0
17.4
34.4
–
GiGPO
42.0
59.5
42.4
36.9
37.0
12.6
64.1
42.1
Appendix
Table 4: Performance on search-augmented QA tasks. Models are trained on NQ and HotpotQA with Qwen2.5-3B-Instruct. † and ∗ indicate in-domain and out-of-domain datasets, respectively. Best results are in bold .
Step
GRPO
GiGPO
Mvpo
ΔGi
20
20.3
23.4
28.1
+4.7
30
22.7
19.5
32.0
+12.5
50
31.2
32.8
47.7
+14.9
60
21.1
37.5
57.0
+19.5
70
40.6
60.2
60.9
+0.7
90
46.9
76.6
78.1
+1.5
Appendix
Table 5: Empirical verification of the theoretical prediction on ALFWorld 1.5B. We report validation success rates at representative training steps. ΔGi denotes the improvement of Mvpo over GiGPO.
Long-horizon LLM agents require reinforcement learning methods that can assign credit to intermediate decisions under sparse and delayed rewards. Recent group-based methods such as GiGPO improve over GRPO by constructing step-level advantages at repeated anchor states. However, we show that such dense credit can be statistically unreliable: under limited rollouts, rare but lucky actions may receive overly large advantages, producing divergent anchor bias and late-stage training oscillation. We propose Evidence-Calibrated Policy Optimization (ECPO), a critic-free policy optimization algorithm that calibrates step-level credit before policy updates. ECPO combines Evidence-Calibrated Action Advantage, which groups rollouts by canonical actions and shrinks low-count estimates, with Variance-Gated Credit Weighting, which suppresses anchor states dominated by within-action noise. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B/7B show that ECPO consistently outperforms strong baselines, improving GiGPO by +5.2/+7.3 success points on ALFWorld/WebShop with Qwen2.5-1.5B while adding only 0.1% additional advantage-computation overhead.
Yuanfan Li, Qi Zhou, Wenjing Duan +1
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China · Faculty of Electronic and Information Engineering, Xi’an Jiaotong University
Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by comparing trajectories or steps within a group. However, on difficult long-horizon tasks, this comparison can suffer from a sampling imbalance: repeated or low-effect actions dominate the high-probability region of the policy while useful state-changing actions remain under-sampled. This imbalance produces many all-failed rollout groups, where outcome rewards provide no direction for correcting the policy. Together, these effects can form a self-reinforcing credit trap: failure-dominated sampling yields no outcome-based correction, allowing repeated low-effect actions to persist. To break this loop, we propose Progress-conditioned Group Policy Optimization (ProGPO), which uses first-visit observation coverage only when all samples in a group receive zero outcome reward. Specifically, within such groups, ProGPO assigns higher relative advantages to trajectories or steps that visit more new states since reaching new observations is a prerequisite for task success. Experiments on two challenging agentic benchmarks, ALFWorld and WebShop with Qwen2.5-1.5/7B-Instruct, show that ProGPO consistently improves over group-based baselines, with particularly large gains on hard tasks.
Kaibing Yang, Guangfeng Cai, Shengtian Yang +6
Southeast University · Kuaishou Technology · Work done during an internship at Kuaishou Technology, supervised by Jun Xu. +1
While long-horizon agentic tasks require language agents to perform dozens of sequential decisions, training such agents with reinforcement learning remains challenging. We identify two root causes: credit misattribution, where correct early actions are penalized due to terminal failures, and sample inefficiency, where scarce successful trajectories result in near-total loss of learning signal. We introduce a milestone-guided policy learning framework, BEACON, that leverages the compositional structure of long-horizon tasks to ensure precise credit assignment. BEACON partitions trajectories at milestone boundaries, applies temporal reward shaping within segments to credit partial progress, and estimates advantages at dual scales to prevent distant failures from corrupting the evaluation of local actions. On ALFWorld, WebShop, and ScienceWorld, BEACON consistently outperforms GRPO and GiGPO. Notably, on long-horizon ALFWorld tasks, BEACON achieves 92.9% success rate, nearly doubling GRPO's 53.5%, while improving effective sample utilization from 23.7% to 82.0%. These results establish milestone-anchored credit assignment as an effective paradigm for training long-horizon language agents. Code is available at https://github.com/ZJU-REAL/BEACON.