QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
Organizations: Alibaba Token Hub, Alibaba Group · University of Science and Technology of China · Tsinghua University
Abstract
Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to and speedups over Colocate and Async, respectively.
Figures & tables
| ID | Evolution | Stream | Switch | Allocation | Total (h) | Time / E0 | ||
| A0 | Original Async | Disabled | – | 16 T / 16 R | 25.08 | 1.556 | ||
| A1 | A0 with adjusted partition | Disabled | – | 8 T / 24 R | 27.01 | 1.676 | ||
| A2 | A0 + multi-step burst | Disabled | – | 16 T / 16 R | 32.58 | 2.021 | ||
| A3 | A0 + streaming | Enabled | – | 16 T / 16 R | 22.18 | 1.376 | ||
| C0 | Original Colocate | Disabled | Global | 32 C / 0 S | 23.23 | 1.442 | ||
| C1 | C0 + standalone rollout nodes | Disabled | Coarse | 28 C / 4 S | 22.56 | 1.400 |
| Dataset | vs. Async | vs. Colocate | ||
|---|---|---|---|---|
| DeepSWE | 1 | 1.569 | 1.610 | |
| DeepSWE | 1.5 | 1.572 | 1.824 | |
| TerminalBench | 1 | 1.383 | 1.809 | |
| TerminalBench | 1.5 | 1.430 | 1.849 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| , | Groups per update; updates per burst. |
| Extra boundary dispatch in units of groups. | |
| Boundary depth after replenishment: all unconsumed groups. | |
| Continuous depth: unfinished groups only. | |
| , | Published version at admission; trainer version immediately before consumption. |
| Staleness ; denotes its boundary mean in Section A.5 . |