Moment-Accurate Gaussian Mixtures for Constant-Step Stochastic Approximation
Organizations: University of Chicago
Abstract
Local Gaussian models of constant-step learning predict output variability and expected losses, but weak convergence alone does not justify these moment predictions. We establish moment-accurate Gaussian mixtures by matching stationary energy with local Ornstein--Uhlenbeck limits, ruling out quadratic tail mass invisible to weak convergence. For step size , the second-order Wasserstein error is , uniformly over invariant laws, using each law's actual root weights. The assumptions combine confinement, descent, finitely many hyperbolic equilibria and root continuity with finite-variance innovations. The result yields observable covariances, expected objective gaps and first-order mean shifts, while allowing singular covariances, compatible saddles and weights without a limit. For additive noise given by a fixed invertible transform of independent standardized Student coordinates, symmetry gives an order-sharp smooth-test bound. Numerical transport calculations demonstrate the value of root-specific covariances; controlled SGD studies assess observable predictions across step sizes, batch sizes and model geometries.
Figures & tables
| Work | Setting and moment control | Relevant conclusion |
|---|---|---|
| Chen et al. (2022) | Single-center stability; additive iid finite-variance noise | Stationary scaling limit; Gaussian identification under uniqueness conditions |
| Dieuleveut et al. (2020) | Strong convexity; higher regularity and moments | Stationary moment and bias expansions |
| Wei et al. (2025) , Thms. 3.4–3.5 | Strong convexity; regularity; joint long-time/small-step limit | CLT at ; convex-set bounds at |
| Wang et al. (2026) , Thm. 3.8 | Contractive setting; additive noise with finite third moment | Stationary rate; derived scaled third-moment bound |
| This paper | Finite hyperbolic roots; descent; second innovation moments and root continuity | Law-uniform all-root moments and actual-weight approximation |
| Model | Output variance | Objective gap | Output variance | Objective gap |
|---|---|---|---|---|
| Digits-256 | ||||
| Digits-1024 | ||||
| WDBC-256 |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Noise | per run | Left occupation | Raw | Completed | |
|---|---|---|---|---|---|
| Gaussian | 0.08 | 2,000,000 | 0 | 0 | |
| Gaussian | 0.04 | 2,000,000 | 0 | 0 | |
| Rademacher | 0.08 | 2,000,000 | 0 | 0 | |
| Rademacher | 0.04 | 2,000,000 | 0 | 0 | |
| Student t3 | 0.08 | 6,600,000 | 477 | 413 | |
| Student t3 | 0.04 | 51,500,000 | 522 | 400 |
| Problem | Full | Diagonal noise | Isotropic noise | Discrete linear | |
|---|---|---|---|---|---|
| Finite sum | 1 | – | |||
| Finite sum | 4 | – | |||
| Finite sum | 16 | – | |||
| Neural | 1 | ||||
| Neural | 4 | ||||
| Neural | 16 |
| Theory |
|---|
| Augmentation | ||||
|---|---|---|---|---|
| Tilt + dither | ||||
| Tilt only | ||||
| Dither only | ||||
| Neither |
| Model | Variance ratio | MARE (%) | Loss ratio | |
|---|---|---|---|---|
| Digits 256 | ||||
| Digits 256 | ||||
| Digits 256 | ||||
| Digits 256 | ||||
| Digits 1024 | ||||
| Digits 1024 |