cs.LGSep 30, 2026

Fast Regularized Policy Mirror Descent with One-Step TD Updates

Authors: Qipei Chen, Wenye Li, Yule Sun, Ke Wei

Organizations: School of Mathematical Sciences, Fudan University · School of Data Science, Fudan University

Abstract

Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or increasingly accurate policy evaluation. We analyze PMD coupled with a persistent critic advanced by one temporal-difference (TD) update. For finite discounted MDPs, we establish global linear convergence in value for exact coordinate-wise Bellman updates, with any positive constant actor stepsize and arbitrary finite critic initialization. The proof combines a resolvent-based auxiliary distribution with a decaying Bellman-violation correction and a potential weighted by inverse coordinate weights. We then study stochastic TD-PMD with general strongly convex mirror maps under a single off-policy Markov trajectory. With suitably chosen constant stepsizes and a finite-batch TD update, the method achieves an expected value gap of εε after O~(1/((1−γ)5σ~bε))\widetilde{O}(1/((1-γ)^5 \widetildeσ_b ε)) transitions. The stochastic analysis relies on the trajectory-wise Lipschitz continuity of the regularizer, derived from uniform bounds on vertex Bregman divergences, together with a visitation-weighted resolvent estimate for signed critic-error propagation that yields an inverse-linear dependence on behavior coverage σ~b\widetildeσ_b. In contrast to many prior guarantees for regularized policy optimization, our sample-complexity guarantee holds without trajectory resets, generative-model access, or nested policy-evaluation loops. Numerical results are consistent with the theoretical convergence analysis.

Figures & tables

Explore similar work

CardsList
  1. Value Mirror Descent for Reinforcement Learning

    Apr 7, 2026Zhichao Jia, Guanghui LanMirror DescentValue Functions

  2. Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction

    May 16, 2026Xingguo Chen, Yuchen Shen, Shangdong Yang +3Temporal DifferenceMirror Descent

  3. Mirror Descent Beyond Euclidean Stability: An Exponential Separation in Initialization Sensitivity

    Jun 9, 2026Shira Vansover-Hager, Matan Schliserman, Ofir Schlisselberg +1Mirror Descent