Almost Sure Convergence Rates of Stochastic Approximation and Reinforcement Learning via a Poisson-Moreau Drift
Organizations: University of Virginia
Abstract
Establishing almost sure convergence rates for stochastic approximation and reinforcement learning under Markovian noise is a fundamental theoretical challenge. We make progress towards this challenge for a class of stochastic approximation algorithms whose expected updates are contractive, a setting that arises in many reinforcement learning algorithms such as -learning and linear temporal difference learning. Specifically, for a power-law learning rate with , we obtain an almost sure convergence rate arbitrarily close to . For a harmonic learning rate , we obtain an almost sure convergence rate arbitrarily close to , which we argue is a strong result because it is close to the optimal rate given by the law of the iterated logarithm (for a special case of i.i.d. noise). Key to our analysis is a novel Lyapunov drift construction that applies a Poisson-equation based correction for Markovian noise to the well-established Moreau-envelope smoothing for the contractive mapping.