Organizations: State Key Laboratory of Novel Software Technology, Nanjing University · School of Intelligence Science and Technology, Nanjing University · ByteDance
Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that uses past interaction trajectories to provide step-level feedback for policy learning. T2SPO derives remaining-distance targets from successful trajectories and pairs them with representations of the states visited along the way. Conditioned on these examples, a pretrained TabPFN regressor estimates the remaining distance to success at each state of a new rollout. Changes in this distance estimate across consecutive states yield auxiliary credit for agent steps alongside task-level supervision. As training proceeds, newly completed trajectories refresh the estimator's context, incorporating new experience without updating its parameters. Experiments with 1.5B and 7B language models on ALFWorld and WebShop show that T2SPO consistently improves overall task success over GRPO.
Figures & tables
Figure 1: T2SPO converts historical trajectories into step-level credit. (a) Context construction. Successful historical trajectories pair each visited state with its own observed number of turns to completion (Section 4.1 ). State representations and these targets form a labeled context. (b) Step credit for policy optimization. A frozen TabPFN predicts the remaining distance at adjacent states of a current rollout using the same historical context. Their difference is normalized, gated, and scaled into auxiliary step credit, which augments the outcome advantage in GRPO. Current trajectories refresh the context only after completion and scoring.
Type
Method
ALFWorld
WebShop
Pick
Look
Clean
Heat
Cool
Pick2
All
Score
Succ.
Qwen2.5-1.5B-Instruct
Prompting
Qwen2.5
5.9
5.5
3.3
9.7
4.2
0.0
4.1
23.1
5.2
Prompting
ReAct
17.4
20.5
15.7
6.2
7.7
2.0
12.8
40.1
11.3
RL Training
PPO (with critic)
64.8±3.5
40.5±6.9
57.1±4.9
60.6±6.6
46.4±4.0
47.4±1.9
54.4±3.1
73.8±3.0
51.5±2.9
RL Training
RLOO
88.3±3.0
52.8±8.6
71.0±5.9
62.8±8.7
66.4±5.5
56.9±4.7
69.7±2.5
73.9±5.6
52.1±6.7
Table 1: Main results on ALFWorld and WebShop. ALFWorld reports success rates (%); WebShop reports score (0–100) and success rate (%). RL results are means over three seeds; subscripts denote standard deviations. Shaded rows denote T2SPO-GRPO using only successful context examples; bold numbers indicate the highest reported mean within each model size. Δ reports the difference between T2SPO-GRPO and the GRPO reference within each model size: percentage points for success rates and points for WebShop score.
ALFWorld
WebShop
Model
Context examples
Succ. (%)
Score
Succ. (%)
Qwen2.5-1.5B-Instruct
Success only
74.7±10.0
88.5±0.3
73.7±3.3
With 25% failure examples
74.0±5.9
85.0±0.6
68.2±3.0
Qwen2.5-7B-Instruct
Success only
83.6±6.4
85.4±4.1
75.8±3.6
With 25% failure examples
89.3±3.0
86.3±0.6
74.7±6.1
Table 2: Failure-augmented context on ALFWorld and WebShop. Means over three seeds; subscripts report standard deviations. Bold numbers indicate the higher mean within each model-size pair.
Figure 2: Offline distance prediction on WebShop. (a) Remaining-distance MAE across nine context/SVD settings; parameter axes use logarithmic spacing, and facets connect evaluated points. The square marks the main RL configuration; the star marks the lowest observed MAE. (b) TabPFN and the lowest kNN MAE across five tested neighbor counts per context budget, with matched context examples, SVD32 features, and queries.
Figure 3: Online computation cost on WebShop. (a) Mutually exclusive training stages across three seeds and 360 ordinary optimizer steps. (b) Substages within advantage computation; times are means in seconds. Both panels report shares of summed full-step time. Other advantage work includes GRPO advantage computation, shaping, and context commit.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
ALFWorld
WebShop
Task instances per update
16
16
Rollouts per task instance
8
8
Actor learning rate
10−6
10−6
KL loss coefficient
0.01
0.01
Invalid-action penalty
0.10
0.10
Policy history window
2 turns
2 turns
Appendix
Table 3: Policy training and evaluation settings shared by both model sizes.
Shared settings
Value
Estimator history window
Full prefix
TabPFN package version
8.3.0
Shared context budget
256
Successful-state pool capacity
256
New successful states per round
≤64
Recent trajectory window
256
Appendix
Table 4: Estimator settings for the main T2SPO-GRPO runs at both model sizes. History refers to the estimator input; context and successful-state pool budgets count state rows.
Type
Method
ALFWorld
Pick
Look
Clean
Heat
Cool
Pick2
All
Qwen2.5-1.5B-Instruct
RL Training
T2SPO-GRPO
87.4±7.4
56.8±15.9
90.3±10.1
65.9±18.2
86.5±3.0
43.7±23.9
74.7±10.0
RL Training
T2SPO-GRPO + Fail
81.6±4.6
65.2±17.0
82.9±13.1
70.2±11.4
77.9±2.6
54.1±21.2
74.0±5.9
Qwen2.5-7B-Instruct
RL Training
T2SPO-GRPO
91.3±5.7
73.8±2.1
97.3±2.4
84.5±16.8
78.4±7.7
63.5±10.5
83.6±6.4
Appendix
Table 5: Failure-augmented context across ALFWorld task categories. Success rates (%), averaged over three seeds; subscripts denote sample standard deviations. T2SPO-GRPO uses success-only context, while T2SPO-GRPO + Fail admits up to 25% failure examples. Bold marks the higher mean at each size.
Context
SVD
MAE ↓
RMSE ↓
ρ↑
128
32
0.8526
1.4784
0.7293
128
64
0.8370
1.4714
0.7500
128
128
1.0666
1.6504
0.7240
256
32
0.8215
1.4698
0.7486
256
64
0.8345
1.4681
0.7838
256
128
0.7987
1.4477
0.7815
Appendix
Table 6: Offline remaining-distance prediction on WebShop. All rows use TabPFN and evaluate the same 1,951 successful query states on held-out tasks. MAE and RMSE are measured in turns; ρ denotes Spearman correlation. Bold numbers mark the best value for each metric.
Estimator
M=128
M=256
M=512
kNN ( k=8 )
1.1042
1.0804
1.0022
kNN ( k=16 )
1.0429
1.0948
1.0401
kNN ( k=M/4 )
0.9856
0.9978
1.0093
kNN ( k=M/2 )
1.0874
1.0882
1.1083
kNN ( k=M )
1.4719
1.4678
1.4669
TabPFN
0.8526
0.8215
0.6500
Appendix
Table 7: TabPFN and kNN on WebShop. All results use SVD32; columns give context budgets M . MAE is in turns, with the lowest value in bold. At k=M , kNN returns the context mean.
Component
Mean ± SD
Median
Share (%)
Rollout generation
171.729±34.274
165.729
45.167
Reward extraction
0.560±0.215
0.515
0.147
Old-policy log-probability
21.657±6.640
19.245
5.696
Reference-policy log-probability
20.616±6.416
18.093
5.422
Advantage computation
80.682±13.747
79.454
21.221
Actor update
84.830±28.138
73.509
22.312
Appendix
Table 8: Online training-step timing on WebShop. Means and standard deviations are over 360 optimizer steps from three seeds; times are in seconds. Shares use summed full-step time. The lower block expands the advantage row in the upper block.
Model
Turn-PPO
T2SPO-Turn-PPO
Δ
Qwen2.5-1.5B-Instruct
91.4±6.6
91.4±3.3
0.0
Qwen2.5-7B-Instruct
90.2±9.4
91.4±1.1
+1.2
Appendix
Table 9: Turn-PPO with and without T2SPO on ALFWorld. Success rates (%) are means with sample standard deviations over the same two training seeds, 0 and 1 ( n=2 ). Each final checkpoint is evaluated on 128 episodes after 150 updates; Δ is the change in mean success in percentage points.
Training large language models (LLMs) as autonomous agents via reinforcement learning (RL) has enabled frontier models to achieve superhuman performance in long-horizon tasks. However, existing RL algorithms operate at the trajectory level, performing policy optimization only after collecting complete episode rollouts. This coarse-grained approach faces fundamental challenges in multi-turn agent settings where rewards are sparse, delayed, and credit assignment across individual steps is critical. In this work, we propose \textbf{State-Score-Supervised Policy Optimization (3SPO)}, a novel RL algorithm that performs post-step policy optimization with dynamic state score supervision. At each step, 3SPO computes the state score based on historical success rates, supervising step-wise credit assignment, adaptive rollout and post-step policy optimization without requiring value function estimation or additional auxiliary models. Theoretically, under a per-state bandit abstraction, we show that the proposed score-supervised allocation mechanism achieves logarithmic allocation regret and provide sample-complexity guarantees for action identification, score distinguishability, and filtering stability. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B/7B-Instruct show that 3SPO consistently outperforms GRPO by +22.6% on ALFWorld and +15.6 points on WebShop, while using comparable resources to achieve 2.4× more state exploration and 1.8× faster convergence. Code is available at https://github.com/genalyu/3SPO.
Yu Han, Kailing Li, Yang Jiao +4
School of Computer Science and Technology, East China Normal University · Fudan University · KAUST +1
Recently, Reinforcement Learning (RL) has emerged as a crucial paradigm for the post-training of Large Language Model (LLM) agents. However, existing methods predominantly rely on sparse task rewards for policy optimization, failing to fully exploit another class of inherently dense supervisory signals naturally present during online interaction: environmental feedback following action execution. Recent theoretical studies suggest that generalization in multi-step, goal-oriented tasks hinges on predictive knowledge of environmental consequences. Inspired by this, we propose TAPO: Transition-Aware Policy Optimization for LLM Agents, a unified training framework that alternates between policy optimization and transition supervision. Beyond standard RL updates, TAPO repurposes rollout data to apply action-conditioned next-observation prediction supervision on a shared backbone model. This approach enhances the model's sensitivity to environmental transition dynamics and action consequences while concurrently optimizing the policy. It serves as a computationally lightweight, plug-and-play enhancement module for existing agent RL algorithms, requiring no additional expert data, extra sampling costs, or inference-time overhead. We conduct systematic experiments on WebShop and ALFWorld, integrating foundation models of various scales with different policy optimization algorithms. Empirical results demonstrate that TAPO consistently improves task performance over pure policy optimization baselines.
Cong Li, Peixi Peng, Yisen Zhao +4
School of Electronic and Computer Engineering, Peking University · Pengcheng Laboratory
Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to trajectory neglect, in which agents lose focus on the task goal and interaction history at intermediate steps. Prior work has explored step-level supervision using Shannon-entropy-based uncertainty signals, which conflate inherent state complexity with agent confidence and therefore provide unreliable estimates of decision reliability. To address this issue, we propose normalized entropy, which measures confidence deviations relative to an agent's average behavior under a given state, thereby strengthening the association between low-quality actions and trajectory neglect. Building on this insight, we introduce Selective Trajectory-Aware Policy Optimization (STAPO), a hierarchical group-based RL framework. STAPO leverages normalized entropy to locate outlier steps associated with trajectory neglect and optimizes them via a joint mechanism of trajectory-aware reward and trajectory-independent penalty, enhancing trajectory awareness while preserving training stability. Extensive experiments on ALFWorld, WebShop, and Search-Augmented QA demonstrate that STAPO achieves state-of-the-art performance while substantially alleviating trajectory neglect, validating its effectiveness and robustness for agentic tasks.
Qiuyi Qi, Tian Liang, Mutian Bao +8
Zhejiang University · Ant Group · City University of Hong Kong