stat.MLOct 4, 2026

Moment-Accurate Gaussian Mixtures for Constant-Step Stochastic Approximation

Authors: Xiaoli Li, Wei Biao Wu

Organizations: University of Chicago

Abstract

Local Gaussian models of constant-step learning predict output variability and expected losses, but weak convergence alone does not justify these moment predictions. We establish moment-accurate Gaussian mixtures by matching stationary energy with local Ornstein--Uhlenbeck limits, ruling out quadratic tail mass invisible to weak convergence. For step size aa, the second-order Wasserstein error is o(a)o(\sqrt a), uniformly over invariant laws, using each law's actual root weights. The assumptions combine confinement, descent, finitely many hyperbolic equilibria and root continuity with finite-variance innovations. The result yields observable covariances, expected objective gaps and first-order mean shifts, while allowing singular covariances, compatible saddles and weights without a limit. For additive noise given by a fixed invertible transform of independent standardized Student t3t_3 coordinates, symmetry gives an order-sharp a\sqrt a smooth-test bound. Numerical transport calculations demonstrate the value of root-specific covariances; controlled SGD studies assess observable predictions across step sizes, batch sizes and model geometries.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 30, 2026cs.LG

Sharp Stationary Gaussian Approximation for Constant-Stepsize SGD

We prove a sharp Gaussian approximation for the invariant law of constant-stepsize SGD with bounded additive noise generated by an exogenous uniformly ergodic Markov chain. For a smooth, strongly convex objective with a Lipschitz Hessian and nondegenerate long-run noise covariance, the centered iterate normalized by the square root of the stepsize is O(α)O(\sqrtα)-close in 1-Wasserstein distance to its limiting Gaussian. The proof combines blockwise Gaussian comparison with long-run contraction. A four-state example gives a matching lower bound although the one-time noise marginal is symmetric and every nonzero-lag autocovariance vanishes. In this example, an adjacent third-order mixed moment produces the leading correction.
Feb 15, 2026cs.LG

Constant-Stepsize Stochastic Approximation: Finite-Time Convergence, Gaussian Approximation, and Tail Bounds

Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency, yet the distribution of the iterates is typically intractable. Classical asymptotics results give Xk(α)≈X(α)≈x⋆+αYX_k^{(α)} \approx X^{(α)} \approx x^\star+\sqrtαY, where X(α)X^{(α)} is the steady state and YY is an appropriate Gaussian limit, by progressively taking the time k↑∞k\uparrow\infty and stepsize α↓0α\downarrow0. Such limit results, however, do not quantify finite-time, finite-stepsize errors. We develop an explicit pre-limit characterization for SA with i.i.d.\ and Markovian noise. We establish existence and uniqueness of the stationary law, a geometric Wasserstein convergence to stationarity, and almost-sure and L3L^3 convergence of the steady state to the root x⋆x^\star, identifying the scale α\sqrtα as first-order fluctuation. At this scale, we derive a higher-order quantitative Gaussian approximation with a Wasserstein error, using Stein's method and Poisson equation techniques. We further obtain non-uniform Berry--Esseen-type tail bounds, incorporating both steady-state approximation and finite-time convergence errors. We instantiate the theory for strongly convex smooth SGD, linear SA, and nonlinear contractive SA. Beyond strong convexity, for general convex SGD, we identify a Gibbs limiting law and prove a pre-limit Wasserstein approximation error under stability and Stein-equation hypothesis, which are validated numerically.
Oct 4, 2026stat.ML

Gaussian Limits for SGD Without Stationary Moments

Temporal dependence can separate the Gaussian approximation of stochastic gradient descent from its stationary moments. For unmodified least-squares SGD, we construct a design with standard Gaussian marginals whose stationary error has every positive moment infinite. Independent observations with the same marginals instead give finite stationary variance. Both regimes retain a Gaussian small-step limit. Our general theory establishes pathwise contraction from a finite second design moment, then uses score cancellation and localization to obtain stationary Gaussian and Ornstein--Uhlenbeck limits. Independent Gaussian regression errors yield an exact conditional Gaussian law and total-variation convergence under the same design integrability. Stronger design conditions identify a positive first-order total-variation constant and a deterministic covariance correction with o(a)o(a) error. A scalar coverage expansion translates this correction into its inference consequence. Experiments examine distributional error, coverage, and calibration with dependent scores. Together, these results establish precise probability-law approximation beyond moment-based stationary analysis.