cs.LGSep 28, 2026

An analysis of Mirror-Descent Soft Actor-Critic

Authors: Denis Zorba, Michal Valko

Organizations: School of Mathematics, University of Edinburgh, United Kingdom · INRIA, Paris, France

Abstract

Soft Actor-Critic (SAC) is widely used for entropy-regularised reinforcement learning with continuous action spaces, and practical implementations perform only a few actor steps towards an evolving target. In this work, we prove convergence guarantees when the target policy arises from policy mirror descent and compare it with the classical Gibbs target. We derive sufficient conditions for the strong convexity and smoothness of the actor objective, characterised by the curvature of the QQ-function estimate through the Legendre differential operator, and establish an O ⁣(N−15)\mathcal{O}\!\left(N^{-\frac{1}{5}}\right) best-iterate finite-time convergence rate up to actor and critic approximation errors. Moreover, the mirror-descent step size λλ directly controls the target drift and hence actor tracking error, whereas the analogous Gibbs bound contains a non-vanishing tracking term.

Figures & tables

Explore similar work

CardsList
  1. Refined Analysis of Entropy-Regularized Actor-Critic

    May 23, 2026Safwan Labbi, Paul Mangold, Daniil Tiapkin +1Entropy Regularized Reinforcement LearningSoft Actor-Critic

  2. Fast Regularized Policy Mirror Descent with One-Step TD Updates

    Sep 30, 2026Qipei Chen, Wenye Li, Yule Sun +1Mirror DescentTemporal Difference

  3. Generative Actor-Critic with Soft Bridge Policies

    May 9, 2026Ke He, Le He, Shunpu Tang +2Soft Actor-CriticOffline Reinforcement Learning