Organizations: State Key Laboratory of Novel Software Technology, Nanjing University · School of Intelligence Science and Technology, Nanjing University · ByteDance
Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that uses past interaction trajectories to provide step-level feedback for policy learning. T2SPO derives remaining-distance targets from successful trajectories and pairs them with representations of the states visited along the way. Conditioned on these examples, a pretrained TabPFN regressor estimates the remaining distance to success at each state of a new rollout. Changes in this distance estimate across consecutive states yield auxiliary credit for agent steps alongside task-level supervision. As training proceeds, newly completed trajectories refresh the estimator's context, incorporating new experience without updating its parameters. Experiments with 1.5B and 7B language models on ALFWorld and WebShop show that T2SPO consistently improves overall task success over GRPO.
Figures & tables
Figure 1: T2SPO converts historical trajectories into step-level credit. (a) Context construction. Successful historical trajectories pair each visited state with its own observed number of turns to completion (Section 4.1 ). State representations and these targets form a labeled context. (b) Step credit for policy optimization. A frozen TabPFN predicts the remaining distance at adjacent states of a current rollout using the same historical context. Their difference is normalized, gated, and scaled into auxiliary step credit, which augments the outcome advantage in GRPO. Current trajectories refresh the context only after completion and scoring.
Type
Method
ALFWorld
WebShop
Pick
Look
Clean
Heat
Cool
Pick2
All
Score
Succ.
Qwen2.5-1.5B-Instruct
Prompting
Qwen2.5
5.9
5.5
3.3
9.7
4.2
0.0
4.1
23.1
5.2
Prompting
ReAct
17.4
20.5
15.7
6.2
7.7
2.0
12.8
40.1
11.3
RL Training
PPO (with critic)
64.8±3.5
40.5±6.9
57.1±4.9
60.6±6.6
46.4±4.0
47.4±1.9
54.4±3.1
73.8±3.0
51.5±2.9
RL Training
RLOO
88.3±3.0
52.8±8.6
71.0±5.9
62.8±8.7
66.4±5.5
56.9±4.7
69.7±2.5
73.9±5.6
52.1±6.7
Table 1: Main results on ALFWorld and WebShop. ALFWorld reports success rates (%); WebShop reports score (0–100) and success rate (%). RL results are means over three seeds; subscripts denote standard deviations. Shaded rows denote T2SPO-GRPO using only successful context examples; bold numbers indicate the highest reported mean within each model size. Δ reports the difference between T2SPO-GRPO and the GRPO reference within each model size: percentage points for success rates and points for WebShop score.
ALFWorld
WebShop
Model
Context examples
Succ. (%)
Score
Succ. (%)
Qwen2.5-1.5B-Instruct
Success only
74.7±10.0
88.5±0.3
73.7±3.3
With 25% failure examples
74.0±5.9
85.0±0.6
68.2±3.0
Qwen2.5-7B-Instruct
Success only
83.6±6.4
85.4±4.1
75.8±3.6
With 25% failure examples
89.3±3.0
86.3±0.6
74.7±6.1
Table 2: Failure-augmented context on ALFWorld and WebShop. Means over three seeds; subscripts report standard deviations. Bold numbers indicate the higher mean within each model-size pair.
Figure 2: Offline distance prediction on WebShop. (a) Remaining-distance MAE across nine context/SVD settings; parameter axes use logarithmic spacing, and facets connect evaluated points. The square marks the main RL configuration; the star marks the lowest observed MAE. (b) TabPFN and the lowest kNN MAE across five tested neighbor counts per context budget, with matched context examples, SVD32 features, and queries.
Figure 3: Online computation cost on WebShop. (a) Mutually exclusive training stages across three seeds and 360 ordinary optimizer steps. (b) Substages within advantage computation; times are means in seconds. Both panels report shares of summed full-step time. Other advantage work includes GRPO advantage computation, shaping, and context commit.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
ALFWorld
WebShop
Task instances per update
16
16
Rollouts per task instance
8
8
Actor learning rate
10−6
10−6
KL loss coefficient
0.01
0.01
Invalid-action penalty
0.10
0.10
Policy history window
2 turns
2 turns
Appendix
Table 3: Policy training and evaluation settings shared by both model sizes.
Shared settings
Value
Estimator history window
Full prefix
TabPFN package version
8.3.0
Shared context budget
256
Successful-state pool capacity
256
New successful states per round
≤64
Recent trajectory window
256
Appendix
Table 4: Estimator settings for the main T2SPO-GRPO runs at both model sizes. History refers to the estimator input; context and successful-state pool budgets count state rows.
Type
Method
ALFWorld
Pick
Look
Clean
Heat
Cool
Pick2
All
Qwen2.5-1.5B-Instruct
RL Training
T2SPO-GRPO
87.4±7.4
56.8±15.9
90.3±10.1
65.9±18.2
86.5±3.0
43.7±23.9
74.7±10.0
RL Training
T2SPO-GRPO + Fail
81.6±4.6
65.2±17.0
82.9±13.1
70.2±11.4
77.9±2.6
54.1±21.2
74.0±5.9
Qwen2.5-7B-Instruct
RL Training
T2SPO-GRPO
91.3±5.7
73.8±2.1
97.3±2.4
84.5±16.8
78.4±7.7
63.5±10.5
83.6±6.4
Appendix
Table 5: Failure-augmented context across ALFWorld task categories. Success rates (%), averaged over three seeds; subscripts denote sample standard deviations. T2SPO-GRPO uses success-only context, while T2SPO-GRPO + Fail admits up to 25% failure examples. Bold marks the higher mean at each size.
Context
SVD
MAE ↓
RMSE ↓
ρ↑
128
32
0.8526
1.4784
0.7293
128
64
0.8370
1.4714
0.7500
128
128
1.0666
1.6504
0.7240
256
32
0.8215
1.4698
0.7486
256
64
0.8345
1.4681
0.7838
256
128
0.7987
1.4477
0.7815
Appendix
Table 6: Offline remaining-distance prediction on WebShop. All rows use TabPFN and evaluate the same 1,951 successful query states on held-out tasks. MAE and RMSE are measured in turns; ρ denotes Spearman correlation. Bold numbers mark the best value for each metric.
Estimator
M=128
M=256
M=512
kNN ( k=8 )
1.1042
1.0804
1.0022
kNN ( k=16 )
1.0429
1.0948
1.0401
kNN ( k=M/4 )
0.9856
0.9978
1.0093
kNN ( k=M/2 )
1.0874
1.0882
1.1083
kNN ( k=M )
1.4719
1.4678
1.4669
TabPFN
0.8526
0.8215
0.6500
Appendix
Table 7: TabPFN and kNN on WebShop. All results use SVD32; columns give context budgets M . MAE is in turns, with the lowest value in bold. At k=M , kNN returns the context mean.
Component
Mean ± SD
Median
Share (%)
Rollout generation
171.729±34.274
165.729
45.167
Reward extraction
0.560±0.215
0.515
0.147
Old-policy log-probability
21.657±6.640
19.245
5.696
Reference-policy log-probability
20.616±6.416
18.093
5.422
Advantage computation
80.682±13.747
79.454
21.221
Actor update
84.830±28.138
73.509
22.312
Appendix
Table 8: Online training-step timing on WebShop. Means and standard deviations are over 360 optimizer steps from three seeds; times are in seconds. Shares use summed full-step time. The lower block expands the advantage row in the upper block.
Model
Turn-PPO
T2SPO-Turn-PPO
Δ
Qwen2.5-1.5B-Instruct
91.4±6.6
91.4±3.3
0.0
Qwen2.5-7B-Instruct
90.2±9.4
91.4±1.1
+1.2
Appendix
Table 9: Turn-PPO with and without T2SPO on ALFWorld. Success rates (%) are means with sample standard deviations over the same two training seeds, 0 and 1 ( n=2 ). Each final checkpoint is evaluated on 128 episodes after 150 updates; Δ is the change in mean success in percentage points.