cs.LGSep 30, 2026

Sharp Stationary Gaussian Approximation for Constant-Stepsize SGD

Authors: Junghoon Seo

Organizations: PIT IN Corp., South Korea

Abstract

We prove a sharp Gaussian approximation for the invariant law of constant-stepsize SGD with bounded additive noise generated by an exogenous uniformly ergodic Markov chain. For a smooth, strongly convex objective with a Lipschitz Hessian and nondegenerate long-run noise covariance, the centered iterate normalized by the square root of the stepsize is O(α)O(\sqrtα)-close in 1-Wasserstein distance to its limiting Gaussian. The proof combines blockwise Gaussian comparison with long-run contraction. A four-state example gives a matching lower bound although the one-time noise marginal is symmetric and every nonzero-lag autocovariance vanishes. In this example, an adjacent third-order mixed moment produces the leading correction.

Explore similar work

Feb 15, 2026cs.LG

Constant-Stepsize Stochastic Approximation: Finite-Time Convergence, Gaussian Approximation, and Tail Bounds

Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency, yet the distribution of the iterates is typically intractable. Classical asymptotics results give Xk(α)≈X(α)≈x⋆+αYX_k^{(α)} \approx X^{(α)} \approx x^\star+\sqrtαY, where X(α)X^{(α)} is the steady state and YY is an appropriate Gaussian limit, by progressively taking the time k↑∞k\uparrow\infty and stepsize α↓0α\downarrow0. Such limit results, however, do not quantify finite-time, finite-stepsize errors. We develop an explicit pre-limit characterization for SA with i.i.d.\ and Markovian noise. We establish existence and uniqueness of the stationary law, a geometric Wasserstein convergence to stationarity, and almost-sure and L3L^3 convergence of the steady state to the root x⋆x^\star, identifying the scale α\sqrtα as first-order fluctuation. At this scale, we derive a higher-order quantitative Gaussian approximation with a Wasserstein error, using Stein's method and Poisson equation techniques. We further obtain non-uniform Berry--Esseen-type tail bounds, incorporating both steady-state approximation and finite-time convergence errors. We instantiate the theory for strongly convex smooth SGD, linear SA, and nonlinear contractive SA. Beyond strong convexity, for general convex SGD, we identify a Gibbs limiting law and prove a pre-limit Wasserstein approximation error under stability and Stein-equation hypothesis, which are validated numerically.
Oct 4, 2026stat.ML

Moment-Accurate Gaussian Mixtures for Constant-Step Stochastic Approximation

Local Gaussian models of constant-step learning predict output variability and expected losses, but weak convergence alone does not justify these moment predictions. We establish moment-accurate Gaussian mixtures by matching stationary energy with local Ornstein--Uhlenbeck limits, ruling out quadratic tail mass invisible to weak convergence. For step size aa, the second-order Wasserstein error is o(a)o(\sqrt a), uniformly over invariant laws, using each law's actual root weights. The assumptions combine confinement, descent, finitely many hyperbolic equilibria and root continuity with finite-variance innovations. The result yields observable covariances, expected objective gaps and first-order mean shifts, while allowing singular covariances, compatible saddles and weights without a limit. For additive noise given by a fixed invertible transform of independent standardized Student t3t_3 coordinates, symmetry gives an order-sharp a\sqrt a smooth-test bound. Numerical transport calculations demonstrate the value of root-specific covariances; controlled SGD studies assess observable predictions across step sizes, batch sizes and model geometries.
Jul 17, 2026cs.LG

Scaling Limits of Constant-Stepsize SGD at Flat Minima

For stochastic gradient descent (SGD) with a constant stepsize αα, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons. In the strongly convex case, this invariant law has the familiar α\sqrtα scaling and a Gaussian limit as α↓0α\downarrow 0. We show that this behavior changes fundamentally for convex objectives HH with flat minima and (sub)quadratic tails. More specifically, we study SGD with Markovian noise generated by a contractive driving chain. For every sufficiently small constant stepsize αα, we prove existence, uniqueness, and geometric convergence to an augmented invariant law in a Wasserstein distance induced by an αα-dependent metric. When the minimizer x⋆x_\star has local flatness exponent m≥2m\ge2, meaning that ∇2H(x)≍∥x−x⋆∥m−2Id\nabla^2 H(x)\asymp \lVert x-x_\star\rVert^{m-2} I_d as x→x⋆x\to x_\star, we obtain a contraction bound with factor 1−cαm−11-cα^{m-1}, where c>0c>0 is a constant. This recovers the factor 1−cα1-cα in the quadratic case m=2m=2. We then analyze the small-stepsize scaling limit. We show that the invariant law concentrates on the scale α1/mα^{1/m} and that the rescaled iterates converge weakly to the stationary distribution of the stochastic differential equation dYt=−h0(Yt) dt+Σ1/2 dBt,dY_t=-h_0(Y_t)\,dt+Σ^{1/2}\,dB_t , where h0h_0 is the limiting drift at the minimizer and ΣΣ denotes the asymptotic covariance. This recovers the Gaussian limit when m=2m=2 and gives generally non-Gaussian stationary limits in the flat case m>2m>2. Finally, we give corresponding results for coordinate-separable objectives with unequal flatness exponents.