stat.MLSep 24, 2026

Ordinary Nonconvex SGD under Distance-Dependent Moments: Finite-Horizon Stationarity and Nagaev Bounds

Authors: Wei Biao Wu

Abstract

Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic gradient descent for smooth, lower-bounded, possibly nonconvex objectives under distance-dependent conditional moments. Under second moments alone, a direct descent--displacement argument yields T−1/3T^{-1/3} expected average squared-gradient stationarity with a horizon-dependent stepsize. An explicit oracle-complexity corollary matches the known smooth Blum--Gladyshev (BG-0) lower bound, including the Lb2Δ3ε−6Lb_2Δ^3\varepsilon^{-6} and LΔσ2ε−4LΔσ^2\varepsilon^{-4} stochastic terms, where ΔΔ is the initial objective gap and σ2+b2∥x−x1∥2σ^2+b_2\|x-x_1\|^2 bounds the variance. Thus unchanged SGD attains the minimax stochastic complexity in this second-moment class. For p>2p>2, predictable localization and a Hilbert-space Fuk--Nagaev inequality yield a high-probability bound separating logarithmic variance and polynomial rare-shock contributions. The localization radius is derived from the recursion: no bounded-iterate assumption, clipping, normalization, momentum, or increasing batch size is needed. We also give increasing-confidence rates, an objective-gap-growth refinement recovering root-TT stationarity, and stochastic LpL^p-Lipschitz examples. The broad BG-0 optimality statement is distinguished from the smaller mean-square-smooth class, in which additional oracle structure permits faster algorithms.

Explore similar work

CardsList
  1. Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance

    Apr 17, 2026Arda Fazla, Ege C. Kaya, Antesh Upadhyay +1Stochastic OptimizationNonconvex Stochastic Optimization

  2. Beyond Bounded Variance: Variance-Reduced Normalized Methods for Nonconvex Optimization under Blum-Gladyshev Noise

    May 14, 2026Antesh Upadhyay, Arda Fazla, Abolfazl HashemiNonconvex Stochastic OptimizationVariance Reduction

  3. The Exact Time-Uniform Rate Frontier for Stochastic Gradient Descent on Smooth Convex Objectives

    Sep 8, 2026Ruijie Li, Kang Chen, Tianyu WangLearning Rate SchedulingLast-Iterate Convergence