Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
Organizations: Texas A&M University · GE HealthCare · Independent Researcher · University of Minnesota
Abstract
Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we identify two fundamental sources of instability in this setting: (1) a granularity mismatch between token-level policy optimization and turn- structured interactions, and (2) high-variance and unreliable gradient updates induced by off- policy importance sampling and inaccurate advantage estimation. To address these challenges, we propose SORL, Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Train- ing. SORL introduces mechanisms that align policy optimization with the structure of multi- turn interactions and adaptively suppress unreliable off-policy updates, yielding more conserva- tive and robust learning dynamics. Within this framework, we instantiate two stabilized algo- rithms: SO-PPO and SO-GRPO. Both algorithms are designed to mitigate gradient variance and prevent optimization collapse without requiring careful early stopping or heuristic tuning. We evaluate SO-PPO and SO-GRPO on benchmarks spanning open-domain QA, multi-hop QA, and medical multiple-choice QA, and further assess their transfer to asynchronous RL for mathe- matical reasoning by training on DAPO-Math-17k and validating on AIME-2024. These results demonstrate that SORL provides a practical, scalable, and general framework for stabilizing re- inforcement learning in multi-turn LLM agent training and asynchronous RL for mathematical reasoning.
Figures & tables
| Model | In-Domain | Out-of-Domain | Avg. | |||
|---|---|---|---|---|---|---|
| MedQA | MedMCQA | PubMedQA | MMLU-M | MedXpertQA | ||
| Inference-Based Methods (Base: Llama-3.1-8B-Instruct) | ||||||
| Direct Inference | 45.20 | 52.40 | 62.00 | 40.70 | 13.00 | 42.66 |
| CoT | 48.62 | 59.80 | 64.40 | 54.20 | 13.02 | 48.01 |
| MedLlama-3-8B (CoT) | 66.60 | 53.40 | 64.40 | 45.70 | 11.04 | 48.23 |
| Retrieval-Augmented and RL-Enhanced Methods (Base: Llama-3.1-8B-Instruct) | ||||||
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.