cs.AIAug 20, 2026

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

Authors: Dayang Liang, Lang Feng, Bo An, Yunlong Liu

Organizations: Xiamen University, China

Abstract

Agentic reinforcement learning (RL) has emerged as an important post-training approach for enhancing the capabilities of Large Language Models (LLMs). However, existing methods face a trade-off between policy performance and resource efficiency. Conventional Proximal Policy Optimization (PPO) implementations incur substantial memory overhead from a separate critic, whereas critic-free group-relative methods require multiple rollouts and face potential learning bottlenecks on long-horizon tasks. In this work, we propose Single-rollout Autoregressive Policy Optimization (SAPO), an efficient PPO-style framework that unifies policy optimization and value learning within a single causal language model. SAPO exploits the autoregressive structure of LLMs to sequentially generate action and value estimation at distinct causal boundaries with shared parameters, and then jointly optimizes the PPO objectives and an auxiliary on-policy SARSA objective with turn-level generalized advantage estimation, where the latter is designed to facilitate value learning. Extensive experiments on ALFWorld and WebShop with Qwen2.5-1.5B/7B and Qwen3-14B demonstrate that SAPO reduces peak GPU memory usage by 23.1% and per-iteration runtime by 24.8% over strong PPO baseline, while matching or slightly improving task success rate. Our experiments also show that SAPO outperforms Group Relative Policy Optimization (GRPO) and recent cutting-edge variants in both task success and training stability.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

    Jul 8, 2026Zhenyu Hou, Yujiang Li, Jie Tang +1Agentic LearningAutoregressive Rollout

  2. Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning

    Jul 15, 2026Bowei He, Yankai Chen, Xiaokun Zhang +1Frictive Policy OptimizationOffline Reinforcement Learning

  3. 3SPO: State-Score-Supervised Policy Optimization for LLM Agents

    Jun 8, 2026Yu Han, Kailing Li, Yang Jiao +4Large Language Model Reinforcement LearningOffline Reinforcement Learning