Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% → 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to 1.85× and 1.78× speedups over Colocate and Async, respectively.
Figures & tables
ID
Evolution
Stream
Switch
(φ,μ)
Allocation
Total (h)
Time / E0
A0
Original Async
Disabled
–
(1.5,1)
16 T / 16 R
25.08
1.556
A1
A0 with adjusted partition
Disabled
–
(1.5,1)
8 T / 24 R
27.01
1.676
A2
A0 + multi-step burst
Disabled
–
(0.5,3)
16 T / 16 R
32.58
2.021
A3
A0 + streaming
Enabled
–
(1.5,1)
16 T / 16 R
22.18
1.376
C0
Original Colocate
Disabled
Global
(0,4)
32 C / 0 S
23.23
1.442
C1
C0 + standalone rollout nodes
Disabled
Coarse
(0.5,3)
28 C / 4 S
22.56
1.400
Table 1: Scheduling ablations over the first 12 NL2RepoBench training steps with Qwen 3.6 122B at E[d]=1.5 . A0, C0, and E0 reuse the Async, Colocate, and QwenGyre base configurations from Section 6.2 . Arrows show the evolution paths; E1 and E2 branch independently from E0. Switch is – for fixed partitions, Global for switching all nodes together, Coarse for switching the Colocate pool together, and Fine for switching cells independently. Standalone nodes remain in rollout. Allocation lists training/rollout (T/R) or Colocate/standalone (C/S) node counts; cell entries give cells × nodes per cell. Total is the time for all 12 steps in hours; Time / E0 normalizes it to E0. Lower is better.
Dataset
E[d]
(φ,μ)
vs. Async
vs. Colocate
DeepSWE
1
(0.5,2)
1.569
1.610
DeepSWE
1.5
(0.5,3)
1.572
1.824
TerminalBench
1
(0.5,2)
1.383
1.809
TerminalBench
1.5
(0.5,3)
1.430
1.849
Table 2: 24-step speedup of QwenGyre over each baseline with Qwen 3.6 122B. The (φ,μ) column gives QwenGyre ’s extra dispatch in batch units and training steps per burst. Async uses (φ,μ)=(E[d],1) ; Colocate uses (φ,μ)=(0,2E[d]+1) , following Equation 2 .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
B , μ
Groups per update; updates per burst.
φ
Extra boundary dispatch in units of B groups.
Qb
Boundary depth after replenishment: all (φ+μ)B unconsumed groups.
Qc
Continuous depth: unfinished groups only.
vd , vt
Published version at admission; trainer version immediately before consumption.
d
Staleness vt−vd ; denotes its boundary mean in Section A.5 .
Appendix
Table 3: Notation for dispatch and step-time analysis.
A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment. The gap is especially important for AI agents, whose observations, tool outputs, documents, and prior decisions accumulate over long trajectories. LongStraw is an architecture-aware execution stack for million-token RL post-training under a fixed GPU budget, instantiated with Group Relative Policy Optimization (GRPO). It evaluates the shared prompt without autograd, retains only model-specific state needed by later tokens, and replays short response branches one at a time, reducing the live training graph at the cost of additional replay time. We implement it for the hybrid recurrent and full-attention Qwen3.6-27B and the compressed-attention mixture-of-experts GLM-5.2. On eight H20 GPUs, LongStraw completes grouped Qwen scoring and response backward at 2.1M positions for groups of 2 and 8; increasing the group size adds only 0.21 GB of peak allocated memory, while a separate stress test reaches 4.46M positions. On 32 H20 GPUs, we validate the end-to-end LongStraw execution path for a 2.1M-token prompt across all 78 layers of GLM-5.2. These experiments establish execution capacity rather than complete training correctness because the captured prompt state is detached and some distributed forward and gradient composition paths remain incomplete.
Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use also depends on the prefill--decode (PD) configuration. Both the choice between colocation and disaggregation and the optimal PD ratio vary with the workload, making resource scaling and PD configuration interdependent. Exploiting this opportunity requires selecting effective configurations and realizing their benefits within transient resource-availability windows despite reconfiguration costs. We present PEARL, an asynchronous agentic RL system that coordinates external resource elasticity, temporary reuse of idle training GPUs, and adaptive PD execution. PEARL maintains a unified GPU--worker--role state and uses runtime profiles to predict rollout batch completion time, accounting for environment-induced reductions in decode concurrency. It selects the PD mode and ratio under the current GPU budget and translates each decision into an incremental transition plan that minimizes worker and role changes. Cost-aware switching and borrowing policies suppress transitions with insufficient expected benefit while ensuring timely return of training GPUs. Our evaluation show that PEARL achieves 2.17--2.79× the throughput of fixed-resource ROLL across different LLMs. Compared with RLBoost+, throughput improves by up to approximately 26.9% for Qwen3-8B and 36.3% for Qwen3-30B-A3B.
Agent usage is shifting toward long-horizon tasks such as coding and scientific discovery, among which terminal tasks are especially important. We introduce T1, a Mixture-of-Experts model of 122B total trained with reinforcement learning, operating a real shell in a cloud sandbox for up to 300+ tool-call turns per task, rewarded by executing each task's own verifier. We provide a comprehensive recipe: First, an aggressively warm-started to stabilize actor-critic training, with a dense process reward scoring trajectories by the absolute number of passing verifiers. Second, stable optimization through TITO construction, training on the exact sampled token identifiers with drift repair at turn boundaries, and rollout routing replay, recording the sampler's per-token expert choices at every MoE layer and replaying them during training. Third, fully out-of-distribution training corpus: isolated seeds and synthesized tasks disjoint from Terminal-Bench 2.1 ensures gains reflect genuine capability transfer over benchmark overfitting. Together, TITO and R3 cut the training-to-inference log-probability difference from 0.021 to 0.013, with exactly aligned zero token drift in the loss region. On Terminal-Bench 2.1, our post-train pipeline raises initial base model from 43.8% to T1 with 64.0% resolved. On Long-Horizon Terminal Bench, T1 reaches 27.9% and surpasses GPT-5.4 and GLM-5.1.
Junyao Yang, Yucheng Shi, Zhongzhi Li +4
Tencent Hy Foundation Model · National University of Singapore · University of Georgia +2