Local Gaussian models of constant-step learning predict output variability and expected losses, but weak convergence alone does not justify these moment predictions. We establish moment-accurate Gaussian mixtures by matching stationary energy with local Ornstein--Uhlenbeck limits, ruling out quadratic tail mass invisible to weak convergence. For step size a, the second-order Wasserstein error is o(a), uniformly over invariant laws, using each law's actual root weights. The assumptions combine confinement, descent, finitely many hyperbolic equilibria and root continuity with finite-variance innovations. The result yields observable covariances, expected objective gaps and first-order mean shifts, while allowing singular covariances, compatible saddles and weights without a limit. For additive noise given by a fixed invertible transform of independent standardized Student t3 coordinates, symmetry gives an order-sharp a smooth-test bound. Numerical transport calculations demonstrate the value of root-specific covariances; controlled SGD studies assess observable predictions across step sizes, batch sizes and model geometries.
Finite hyperbolic roots; descent; second innovation moments and root continuity
Law-uniform all-root moments and actual-weight W2 approximation
Table 1: Gaussian and moment results for SA in their stated model, moment and time settings.
Figure 1: Root-specific moments and transport. (a) Scaled variance deviations from each root’s OU target; dotted lines show the fixed equal-weight pooled reference’s limiting errors. Bars are one standard error across 16 runs. (b) Second-moment trace deviations in the nongradient model; medians and interquartile ranges retain t3 dispersion. (c) Deterministic same-label scaled transport with alternating weights, using root-specific covariances or the same fixed pooled target.
Figure 2: Neural predictions across step sizes. Output-variance and actual objective-gap deviations share the [−10,110]% range; MARE starts at zero. All four steps, three batches and 16 paths per setting are included. Dashed lines and bands show the mean and one standard error of 64 matched stationary OU paths, whose loss is quadratic.
Model
Output variance B=1
Objective gap B=1
Output variance B=16
Objective gap B=16
Digits-256
1.012±0.005
1.030±0.005
0.998±0.003
0.997±0.003
Digits-1024
1.024±0.003
1.027±0.003
0.996±0.005
0.999±0.003
WDBC-256
1.003±0.003
1.012±0.005
0.997±0.003
1.003±0.004
Table 2: Small-step prediction ratios across models and batches at aλmax(H)=0.000625 . Entries are means ± one standard error across 16 independent paths; all use physical exposure 8000. Output ratios average the 12 fixed probes within each path.
Figure 3: Predictions beyond parameter covariance. Percentage deviations of (a) nonlinear output variance from (a/B)Df2(m)S1Df2(m)T and (b) actual objective gap from (a/4B)trC1(m) share the range [−5,5]% . (c) Signed mean-coefficient errors, with T=1600/a . Bars are one standard error across 16 paths in (a,b) and 32 in (c); the two outputs share paths. Zero denotes the prediction.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: A sharp symmetric finite-variance endpoint. Analytic computations for standardized t3 innovations. (a) The cosine gap divided by a approaches e−1/4/36 . (b) The symmetric tail modulus divided by a approaches 8/π . Dashed lines are the exact limits. The characteristic function uses 100 terms of a convergent logarithmic series at 80-digit arithmetic precision; the modulus uses ( 67 ). No trajectories or sampling error bars enter these calculations.
Figure 5: Local fluctuations with matched innovation covariance. Variance deviations in (a) are percentages relative to the OU target on a zero-centered axis. (a) Median variance deviation and interquartile range across 16 runs; the continuous OU prediction is zero. (b) Pooled first-coordinate Q–Q curves at a=0.0025 . The intervals describe between-run dispersion.
Figure 6: Global occupation and transition rates. (a) Per-run left occupation, with paired positions for a=0.08 and 0.04 ; circles start left and crosses start right. The dotted line is the heavy-tail limit ( 75 ). (b) Heavy-tail completed transitions divided by exposure to the last visited well core, compared with one-jump probabilities per iteration. Run lengths are reported in Table 3 .
Noise
a
N per run
Left occupation
Raw
Completed
Gaussian
0.08
2,000,000
0.500∗
0
0
Gaussian
0.04
2,000,000
0.500∗
0
0
Rademacher
0.08
2,000,000
0.500∗
0
0
Rademacher
0.04
2,000,000
0.500∗
0
0
Student t3
0.08
6,600,000
0.852(0.010)
477
413
Student t3
0.04
51,500,000
0.845(0.011)
522
400
Appendix
Table 3: Occupation experiment: 12 runs per row. N includes burn-in; counts sum post-burn-in events over runs. Heavy-tail occupation is the mean with one across-run standard error in parentheses. ∗ The value 0.500 is fixed by six starts in each well and no observed crossings; it is not a stationary-weight estimate.
Figure 7: Additional local measurements. Coverage in (b) is shown as percentage-point deviation from 95 percent on a zero-centered axis. (a) Bounded characteristic-function discrepancy ( 76 ), shown as median and interquartile range across 16 runs. (b) Nominal 95% local fluctuation coverage for the nonlinear observable, shown as mean and one standard error across runs. These finite-horizon measurements describe distributional agreement and coverage; the asymptotic rates are established by the analytic results.
Figure 8: Distributional shape and covariance at t=1 . Normalized cosine discrepancies (a), skew sine bias divided by a (b), and covariance error (c). Curves are deterministic characteristic-function calculations with analytic truncation error below 10−12 . The nonzero skew limit distinguishes shape error from covariance discretization.
Figure 9: Loss distributions beyond their means. (a) Actual nonlinear losses divided by a in the heterogeneous rectangles at a=0.0025 , with the shared 0.01χ22 limit. (b) Actual finite-sum losses multiplied by B/a at a=0.005 , with the generalized chi-square limit computed by deterministic quadrature. Bands are pointwise one standard error across 16 independent runs. The corresponding moment curves appear in Figures 1 and 3 .
Problem
B
Full
Diagonal noise
Isotropic noise
Discrete linear
Finite sum
1
2.32±0.37
52.40±0.37
34.70±0.30
–
Finite sum
4
1.74±0.23
51.58±0.36
34.03±0.30
–
Finite sum
16
1.34±0.26
52.36±0.27
34.79±0.23
–
Neural
1
6.97±0.47
228.29±4.15
184.92±3.49
6.98±0.47
Neural
4
6.31±0.36
233.11±3.50
189.61±2.99
6.31±0.37
Neural
16
6.09±0.36
227.63±2.49
185.48±2.16
6.09±0.36
Appendix
Table 4: Covariance-reference comparison. Output-variance MARE (percent), mean ± one standard error across 16 runs, at a=0.005 for the two-dimensional finite sum and aλmax(H)=0.000625 for the neural model. Each problem uses its fixed root Hessian and output Jacobian. The discrete-linear reference is evaluated for the neural model.
Figure 10: Small-step neural predictions. At aλmax(H)=0.000625 , (a) logit-variance and (c) actual objective-gap deviations use a common percentage scale. (b) Per-run MARE measures absolute relative error across the 12 probes. Bars summarize 16 paths; dashed lines and bands show the mean and one standard error of 64 matched stationary OU paths. The OU loss is quadratic. The main-text Figure 2 reports all four steps.
Figure 11: Finite-second-moment evidence in the nonlinear nongradient system. (a) All-root truncated moments at a=0.005 , with one standard error across 16 runs and a deterministic Gaussian reference. (b) Median relative covariance error and interquartile range. (c) All individual stable root-centered trace ratios; blue is Gaussian noise and orange is t3 . Raw quadratic measurements use dispersion summaries because t3 has no fourth moment.
a
(f1−f1(m))/a
(f2−f2(m))/a
0.04
−0.05988±0.00455
−0.028306±0.000644
0.02
−0.05769±0.00409
−0.027889±0.000539
0.01
−0.05470±0.00429
−0.027249±0.000498
Theory
−0.05560
−0.027255
Appendix
Table 5: Signed mean coefficients from 32 independent confirmation paths per step, with one standard error. Both outputs use the same paths.
Figure 12: Independent mean measurements and error sources. (a,b) Signed coefficients for both preselected observables. Gray paths use R=16,aT=400 ; the independent confirmation uses R=32,aT=1600 , with paired burn-in windows. Dashed lines are root-derived coefficients. (c,d) The three terms in ( 78 ) for the confirmation’s primary window. Bars are one standard error across independent paths; outputs and windows within a path are dependent.
Augmentation
B=1,as
B=1,al
B=16,as
B=16,al
Tilt + dither
0.986±0.012
1.473±0.011
0.985±0.007
1.009±0.006
Tilt only
0.986±0.012
1.507±0.015
0.986±0.007
1.008±0.006
Dither only
0.986±0.012
1.469±0.011
0.985±0.007
1.009±0.006
Neither
0.986±0.012
1.502±0.014
0.986±0.007
1.008±0.006
Appendix
Table 6: Neural augmentation ablation. Mean logit-variance ratios (one standard error across 16 trajectories) use each arm’s root-derived reference. The common steps are as=0.002603 and al=0.020825 .
Figure 13: Directional effect of dither. Its fraction of theoretical variance in six Hessian eigenvectors (a) and 12 fixed logit directions (b). Both panels use the same 0–100 percent scale; the output range is annotated. These are deterministic contributions to the local Lyapunov covariance.
Figure 14: Augmentation ablation across step and batch. Logit-variance deviations from prediction (top) and MARE in six Hessian directions (bottom). Each row shares its vertical range across batches. Each arm uses its own root-derived reference; bars are one standard error across 16 paths. Comparable small-step output accuracy coexists with substantial directional errors at the largest step and B=1 .
Figure 15: Neural finite-step decomposition on paired states. Actual and quadratic objective gaps (a), scaled drift residual and linear drift norm (b), and relative state-dependent covariance variation (c). Bars are one standard error across 16 paths at each step. The covariance norm is Frobenius; drift norms are Euclidean.
Figure 16: Observation-window sensitivity. Per-run logit MARE (a) and mean variance ratio (b) for paired short and full windows of 16 independent neural paths per setting. The stationary OU references use 64 paths at each window length. Bars are one standard error across paths; shared paths make the two neural estimates dependent.
Model
aλmax,B
Variance ratio
MARE (%)
Loss ratio
Digits 256
0.000625,1
1.012±0.005
3.018±0.225
1.030±0.005
Digits 256
0.000625,16
0.998±0.003
3.245±0.203
0.997±0.003
Digits 256
0.005,1
1.520±0.006
52.034±0.571
1.927±0.010
Digits 256
0.005,16
1.001±0.004
3.004±0.207
1.012±0.003
Digits 1024
0.000625,1
1.024±0.003
3.689±0.273
1.027±0.003
Digits 1024
0.000625,16
0.996±0.005
3.288±0.185
0.999±0.003
Appendix
Table 7: Independent dimension/data study. Each row uses 16 paths with retained physical exposure 8000. Values are mean ± one standard error. MARE is a percentage.
Figure 17: Moment predictions across dimension and data. Per-run logit MARE (top) and actual objective-gap ratio (bottom) for 16 independent paths per setting. Exact stationary OU references use 64 paths per case and step, matching the save and loss schedules and sharing the scaled reference across batches. All bars are one standard error across paths. The OU loss is quadratic.
Figure 18: Long-horizon communication at a smaller step. (a) Mean occupation and one standard error over 12 paths at nested horizons; the dashed line is the heavy-tail limit. (b) Total completed transitions against retained updates scaled by a3 . (c) Directional rates relative to the finite-step one-jump proxy, with 95% path-bootstrap intervals. Nested windows are dependent.
We prove a sharp Gaussian approximation for the invariant law of constant-stepsize SGD with bounded additive noise generated by an exogenous uniformly ergodic Markov chain. For a smooth, strongly convex objective with a Lipschitz Hessian and nondegenerate long-run noise covariance, the centered iterate normalized by the square root of the stepsize is O(α)-close in 1-Wasserstein distance to its limiting Gaussian. The proof combines blockwise Gaussian comparison with long-run contraction. A four-state example gives a matching lower bound although the one-time noise marginal is symmetric and every nonzero-lag autocovariance vanishes. In this example, an adjacent third-order mixed moment produces the leading correction.
Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency, yet the distribution of the iterates is typically intractable. Classical asymptotics results give Xk(α)≈X(α)≈x⋆+αY, where X(α) is the steady state and Y is an appropriate Gaussian limit, by progressively taking the time k↑∞ and stepsize α↓0. Such limit results, however, do not quantify finite-time, finite-stepsize errors. We develop an explicit pre-limit characterization for SA with i.i.d.\ and Markovian noise. We establish existence and uniqueness of the stationary law, a geometric Wasserstein convergence to stationarity, and almost-sure and L3 convergence of the steady state to the root x⋆, identifying the scale α as first-order fluctuation. At this scale, we derive a higher-order quantitative Gaussian approximation with a Wasserstein error, using Stein's method and Poisson equation techniques. We further obtain non-uniform Berry--Esseen-type tail bounds, incorporating both steady-state approximation and finite-time convergence errors. We instantiate the theory for strongly convex smooth SGD, linear SA, and nonlinear contractive SA. Beyond strong convexity, for general convex SGD, we identify a Gibbs limiting law and prove a pre-limit Wasserstein approximation error under stability and Stein-equation hypothesis, which are validated numerically.
Zedong Wang, Yuyang Wang, Ijay Narang +3
H. Milton Stewart School of Industrial & Systems Engineering, Georgia Institute of Technology, Atlanta, GA 30332, USA · Department of Mathematics, Brown University, Providence, RI 02912, USA · School of Computer Science, Georgia Institute of Technology, Atlanta, GA 30332, USA +2
Temporal dependence can separate the Gaussian approximation of stochastic gradient descent from its stationary moments. For unmodified least-squares SGD, we construct a design with standard Gaussian marginals whose stationary error has every positive moment infinite. Independent observations with the same marginals instead give finite stationary variance. Both regimes retain a Gaussian small-step limit. Our general theory establishes pathwise contraction from a finite second design moment, then uses score cancellation and localization to obtain stationary Gaussian and Ornstein--Uhlenbeck limits. Independent Gaussian regression errors yield an exact conditional Gaussian law and total-variation convergence under the same design integrability. Stronger design conditions identify a positive first-order total-variation constant and a deterministic covariance correction with o(a) error. A scalar coverage expansion translates this correction into its inference consequence. Experiments examine distributional error, coverage, and calibration with dependent scores. Together, these results establish precise probability-law approximation beyond moment-based stationary analysis.