cs.LGSep 30, 2026

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning

Authors: Bo-Wen Zhang, Junwei He, Maoqi Liu, Feiran Li, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, +1 more

Organizations: State Key Laboratory of Novel Software Technology, Nanjing University · School of Intelligence Science and Technology, Nanjing University · ByteDance

Abstract

Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that uses past interaction trajectories to provide step-level feedback for policy learning. T2SPO derives remaining-distance targets from successful trajectories and pairs them with representations of the states visited along the way. Conditioned on these examples, a pretrained TabPFN regressor estimates the remaining distance to success at each state of a new rollout. Changes in this distance estimate across consecutive states yield auxiliary credit for agent steps alongside task-level supervision. As training proceeds, newly completed trajectories refresh the estimator's context, incorporating new experience without updating its parameters. Experiments with 1.5B and 7B language models on ALFWorld and WebShop show that T2SPO consistently improves overall task success over GRPO.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. 3SPO: State-Score-Supervised Policy Optimization for LLM Agents

    Jun 8, 2026Yu Han, Kailing Li, Yang Jiao +4Large Language Model Reinforcement LearningOffline Reinforcement Learning

  2. TAPO: Transition-Aware Policy Optimization for LLM Agents

    Jul 30, 2026Cong Li, Peixi Peng, Yisen Zhao +4Frictive Policy OptimizationLarge Language Model Agents

  3. STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training

    Jul 6, 2026Qiuyi Qi, Tian Liang, Mutian Bao +8Agentic Reinforcement LearningTrajectory-Aware Hidden-State Analyses