cs.LGMay 6, 2025

Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation: The Case of Multi-Armed Bandits

Authors: Max Qiushi Lin, Jincheng Mei, Matin Aghaei, Michael Lu, Bo Dai, Alekh Agarwal, Dale Schuurmans, Csaba Szepesvari, +1 more

Organizations: Simon Fraser University · Google DeepMind · Google Research · Google DeepMind and University of Alberta

Abstract

Policy gradient (PG) methods have played an essential role in the empirical successes of reinforcement learning. In order to handle large state-action spaces, PG methods are typically used with function approximation. In this setting, the approximation error in modeling problem-dependent quantities is a key notion for characterizing the global convergence of PG methods. We study Softmax PG with linear function approximation (referred to as Lin-SPG\texttt{Lin-SPG}) and demonstrate that the approximation error is irrelevant to the algorithm's global convergence even in the bandit setting. Consequently, we rethink the effect of approximation error in the standard stochastic multi-armed bandit problem. We first identify the conditions on the policy feature representation that can guarantee the asymptotic global convergence of Lin-SPG\texttt{Lin-SPG}. Under these feature conditions, we further prove that TT iterations of Lin-SPG\texttt{Lin-SPG} with a problem-specific learning rate result in an O(1/T)O(1/T) convergence to the optimal policy. Moreover, we prove that Lin-SPG\texttt{Lin-SPG} with an arbitrary constant learning rate can ensure asymptotic convergence to the optimal policy.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Vanishing L2 regularization for the softmax Multi Armed Bandit

    May 5, 2026Stefana-Lucia Anita, Gabriel TuriniciEntropy Regularized Reinforcement LearningPolicy Gradient

  2. Delightful Gradients Accelerate Corner Escape

    May 12, 2026Jincheng Mei, Ian OsbandPolicy GradientOptimal Policies