Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we identify two fundamental sources of instability in this setting: (1) a granularity mismatch between token-level policy optimization and turn- structured interactions, and (2) high-variance and unreliable gradient updates induced by off- policy importance sampling and inaccurate advantage estimation. To address these challenges, we propose SORL, Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Train- ing. SORL introduces mechanisms that align policy optimization with the structure of multi- turn interactions and adaptively suppress unreliable off-policy updates, yielding more conserva- tive and robust learning dynamics. Within this framework, we instantiate two stabilized algo- rithms: SO-PPO and SO-GRPO. Both algorithms are designed to mitigate gradient variance and prevent optimization collapse without requiring careful early stopping or heuristic tuning. We evaluate SO-PPO and SO-GRPO on benchmarks spanning open-domain QA, multi-hop QA, and medical multiple-choice QA, and further assess their transfer to asynchronous RL for mathe- matical reasoning by training on DAPO-Math-17k and validating on AIME-2024. These results demonstrate that SORL provides a practical, scalable, and general framework for stabilizing re- inforcement learning in multi-turn LLM agent training and asynchronous RL for mathematical reasoning.
Figures & tables
Figure 1 : Illustration of the proposed SORL framework. The left panel shows the two key components of SORL: turn-level importance sampling (which induces turn-level credit assignment) and clipping-triggered normalization. The right panel reports policy gradient norms during training, demonstrating that the proposed framework effectively suppresses extreme gradient spikes and leads to more stable optimization dynamics.
Figure 2 : Observations from a failed training run using the Qwen2.5-7B model with token-level PPO. From left to right, we report the estimated advantage (batch mean and variance), the ratio of valid actions (i.e., successful tool invocations), the ℓ2 norm of the policy gradient, and the search-task success rate. All metrics are logged per training batch.
Figure 3 : Comparison and diagnostics on Qwen2.5-7B base model for the search task. Left: Token- vs. turn-level PPO, showing improved stability in (a) success rate and (b) policy gradient norm. Right: Clipping-bias diagnostics, including (c) heavy-tailed importance ratios (log scale) and (d) the growth of the clipping-bias norm alongside the gradient norm, indicating increasing off-policy instability. All metrics are logged per training batch.
Figure 4 : Experimental results on Qwen2.5-7B base model. (a,b) Comparison of PPO and SO-PPO on the NQ and HotpotQA datasets, where SO-PPO achieves more stable training and avoids performance collapse. (c,d) Comparison of GRPO variants on the NQ dataset. (c) Success rate shows that GRPO and GSPO suffer from instability, while SO-GRPO remains stable. (d) KL divergence exhibits large spikes for GRPO and GSPO, whereas SO-GRPO stays near zero. Results are averaged over three runs.
Retrieval-Augmented and RL-Enhanced Methods (Base: Llama-3.1-8B-Instruct)
Table 1 : Performance of 8B medical LLMs across in-domain and out-of-domain benchmarks (accuracy %). We record model performance over the course of training. For Search-R1, the best checkpoint corresponds to the highest-performing model observed within the first epoch. Bold numbers indicate the best performance, while underlined numbers indicate the second-best performance.
Figure 5 : Results under the AReaL asynchronous RL framework, averaged over three seeds. SO-PPO improves consistently and remains stable, whereas PPO with conventional optimizer gradient clipping undergoes substantial late-stage performance degradation.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Ablations of the main design choices in SO-PPO. Left: Turn-level importance sampling improves faster, reaches a higher final success rate, and exhibits more stable late-stage optimization than sequence-level importance weighting. Right: The settings δ∈{1,5,10} produce similar optimization trajectories and comparable final performance.
Figure 7 : Training dynamics in a more off-policy setting with a training batch size of 512 and a mini-batch size of 128. Top: Success rate, averaged over three trials. Bottom left: Clipping ratio. Bottom right: KL divergence. Compared with vanilla PPO, SO-PPO maintains stable performance and exhibits lower clipping ratios and KL divergence.
Figure 8 : Effect of optimizer gradient clipping on NQ. Gradient clipping delays but does not prevent the late-stage degradation of standard PPO. In contrast, SO-PPO remains stable with or without gradient clipping, demonstrating that its stabilization benefit does not rely on conventional gradient-norm clipping.
Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use external tools, and conduct iterative search beyond single-turn settings. Yet multi-turn RL training remains highly unstable, often causing severe performance degradation as the number of turns increases. Through theoretical analysis, we identify three tightly coupled sources of instability: rollout-training context mismatch, weak turn-level credit assignment under sparse terminal rewards, and asynchronous policy drift when short and long trajectories are optimized under different policy versions. We show that these issues share a common structural origin in flattened trajectory optimization and address them through a unified reverse-turn formulation. We propose Reverse-Turn Policy Optimization (RTPO), which organizes multi-turn rollouts as sparse reverse trees and performs turn-level policy updates in temporal reverse order, aligning each decision with its downstream continuation. RTPO enables causally consistent turn-level credit assignment and on-policy continuation to control asynchronous drift. We provide theoretical guarantees showing that RTPO eliminates context mismatch and asynchronous drift under the proposed turn-level formulation, reduces credit bias, and converges to recursive optimality. Experiments on multi-turn agentic RL benchmarks show that RTPO improves upon trajectory- and turn-level baselines by 21.50% and 10.76%, respectively, highlighting its potential to support more stable training for tool-using agents.
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged as a more efficient alternative by updating the model as rollouts arrive. However, existing asynchronous RL systems often emphasize throughput, while leaving training stability and task effectiveness largely underexplored. For example, a key challenge is that group-wise sampling in the widely-used GRPO framework does not naturally fit asynchronous agentic training. In this paper, we present Single-rollout Asynchronous Optimization (SAO) to address the stability and off-policy challenges in asynchronous RL. To reduce off-policy effects and improve generalization, we replace group-wise sampling with single-rollout sampling, that is, using one rollout per prompt. We further improve this single-rollout strategy with practical value-model training designs. To improve optimization stability, we introduce a strict double-side token-level clipping strategy. SAO is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks, such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench. We also demonstrate that single-rollout RL is particularly effective in a simulated online learning setting, where the model must adapt to changing evolving environments. To this end, SAO is successfully deployed in the agentic RL pipeline for training the open GLM-5.2 model (750B-A40B).
Zhenyu Hou, Yujiang Li, Jie Tang +1
Tsinghua University · Work done while ZH and YL interned at Z.AI.
Multi-agent LLM systems enable advanced reasoning and tool use via role specialization, yet reliable reinforcement learning (RL) post-training for such systems remains difficult. In this work, we theoretically pinpoint a key reason for training instability when extending group-based RL to multi-agent LLM systems. We show that under GRPO-style optimization, a global normalization baseline may deviate from diverse agents' reward distributions, which ultimately leads to gradient-norm instability. Based on this finding, we propose Dr. MAS, a simple and stable RL training recipe for multi-agent LLM systems. Dr. MAS uses an agent-wise remedy: normalizing advantages per agent using each agent's own reward statistics, which calibrates gradient scales and dramatically stabilizes training, both theoretically and empirically. Beyond the algorithm, Dr. MAS provides an end-to-end RL training framework for multi-agent LLM systems, supporting scalable orchestration, flexible per-agent LLM serving and optimization configs, and shared resource scheduling of LLM actor backends. We evaluate Dr. MAS on multi-agent math reasoning and multi-turn search benchmarks using Qwen2.5 and Qwen3 series models. Dr. MAS achieves clear gains over vanilla GRPO (e.g., +5.6% avg@16 and +4.6% pass@16 on math, and +15.2% avg@16 and +13.1% pass@16 on search) while largely eliminating gradient spikes. Moreover, it remains highly effective under heterogeneous agent-model assignments while improving efficiency.