Fast Regularized Policy Mirror Descent with One-Step TD Updates
Organizations: School of Mathematical Sciences, Fudan University · School of Data Science, Fudan University
Abstract
Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or increasingly accurate policy evaluation. We analyze PMD coupled with a persistent critic advanced by one temporal-difference (TD) update. For finite discounted MDPs, we establish global linear convergence in value for exact coordinate-wise Bellman updates, with any positive constant actor stepsize and arbitrary finite critic initialization. The proof combines a resolvent-based auxiliary distribution with a decaying Bellman-violation correction and a potential weighted by inverse coordinate weights. We then study stochastic TD-PMD with general strongly convex mirror maps under a single off-policy Markov trajectory. With suitably chosen constant stepsizes and a finite-batch TD update, the method achieves an expected value gap of after transitions. The stochastic analysis relies on the trajectory-wise Lipschitz continuity of the regularizer, derived from uniform bounds on vertex Bregman divergences, together with a visitation-weighted resolvent estimate for signed critic-error propagation that yields an inverse-linear dependence on behavior coverage . In contrast to many prior guarantees for regularized policy optimization, our sample-complexity guarantee holds without trajectory resets, generative-model access, or nested policy-evaluation loops. Numerical results are consistent with the theoretical convergence analysis.
Figures & tables
| Algorithm | Iteration complexity | Stochastic access | Sample complexity | |
| Cen et al. [2022] | NPG | — | — | |
| Cen et al. [2023] | TD–NPG | — | — | |
| Zhan et al. [2023] | PMD | — | — | |
| Lan [2023] | PMD | Generative model | ||
| Jia and Lan [2026] | VMD | Generative model | ||
| This work | TD–PMD | Off-policy Markov data |