cs.LGSep 27, 2026

QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

Authors: Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang, Yuyan Luo, Zhiyu Yin, Chencan Wu, +4 more

Organizations: Alibaba Token Hub, Alibaba Group · University of Science and Technology of China · Tsinghua University

Abstract

Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% →\to 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to 1.85×1.85\times and 1.78×1.78\times speedups over Colocate and Async, respectively.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

    Sep 12, 2026Junyao Yang, Yucheng Shi, Zhongzhi Li +4Critic-Free Reinforcement LearningLong-Horizon Task Planning