cs.LGSep 28, 2026

Learning Perturbation Robust Policies for LLM Agents with Stable Optimization

Authors: Pengxin Wang, Yuanzhe LI, Yuxin Ren, Huanrui Yang, Jingdi Chen

Organizations: Department of Electrical and Computer Engineering University of Arizona

Abstract

Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study how to improve perturbation robustness during policy optimization. We first introduce the notion of a perturbation robust policy and analyze conditions under which perturbed policy updates preserve stable monotonic improvement. Based on this analysis, we introduce Stable Perturbation-Robust Policy Optimization (SPrPO), which applies adaptive and sensitivity-aware perturbations during RL training. We evaluate SPrPO on ALFWorld and WebShop and conduct systematic experiments across multiple perturbation types and scales, showing improved perturbation robustness while maintaining stable policy optimization.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?

    May 29, 2026Stephane Hatgis-Kessell, Emma BrunskillLarge Language Model Policy OptimizationFrictive Policy Optimization

  2. Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization

    Nov 25, 2025Chenliang Li, Adel Elmahdy, Alex Boyd +7Large Language Model Reinforcement Learning