cs.AIFeb 12, 2026

TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents

Authors: Aladin Djuhera, Swanand Kadhe, Farhan Ahmed, Syed Zawad, Heiko Ludwig, Holger Boche

Organizations: Technical University Munich · IBM Research

Abstract

Advances in large language models (LLMs) are driving a shift toward using reinforcement learning (RL) to train agents from iterative, multi-turn interactions across tasks. However, multi-turn RL remains challenging as rewards are often sparse or delayed, and environments can be stochastic. In this regime, naive trajectory sampling can hinder exploitation and induce mode collapse. We propose TSR (Trajectory-Search Rollouts), a training-time approach that repurposes test-time scaling ideas for improved per-turn rollout generation. TSR performs lightweight tree-style search to construct higher-quality trajectories by selecting promising actions and trajectory prefixes during rollout generation. This improves rollout quality while preserving stable policy optimization and remains compatible with standard policy-gradient optimizers by design. Across Sokoban, FrozenLake, and WebShop, TSR achieves success-rate gains of up to 15 percentage points and converges in fewer optimization steps, while trading additional training-time rollout compute for stronger policies that require no search at inference time. By moving search from test time to the rollout stage of training, TSR provides a modular mechanism for stronger multi-turn agent learning, complementary to existing frameworks and rejection-sampling-style selection methods.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Process Reward Informed Tree Rollout for Effective Multi-Turn RL

    Jul 17, 2026Xintong Li, Sha Li, Yuwei Zhang +8Autoregressive RolloutRollout

  2. RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training

    Aug 19, 2026Yugu Li, Zehong Cao, Jianglin Qiao +1Agentic Reinforcement LearningFrictive Policy Optimization

  3. T2^2PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning

    May 4, 2026Haixin Wang, Hejie Cui, Chenwei Zhang +7Multi-Turn Reinforcement LearningFrictive Policy Optimization