cs.LGSep 24, 2026

Precise Convergence Speed of Clipped SGD

Authors: David A. R. Robin

Organizations: LAMSADE, Université Paris-Dauphine, PSL Research University

Abstract

We present a tightened convergence analysis of clipped gradient descent on (L0,L1)(L_0, L_1)-smooth functions, with quantitative constants. Building on the ideas of Koloskova et al (2023), we refactor several case disjunctions to reveal the central role of a control of the bias derived from fundamental properties of ℓ2\ell_2-projection, simplifying proofs. We also extend the domain of validity from η≤1/(9β)η\leq 1 / (9 β) to η<1/βη< 1 /β where β=L0+cL1β= L_0 + c L_1 for clipping constant cc, which matches the more traditional analysis of smooth functions. We strengthen the convergence criterion from (min⁡t<TE[∥∇f(xt)∥2])\left( \min_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right) to (1T∑t<TE[∥∇f(xt)∥2])\left( \frac{1}{T} \sum_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right) with matching speed, and lower the final achievable loss from O(min⁡(σ2/c,σ))\mathcal{O}(\min(σ^2/c, σ)) to the more precise 6min⁡(σ2/c,3σ)6 \min(σ^2 /c, 3 σ).

Explore similar work

May 4, 2026math.OC

Robust and Fast Training via Per-Sample Clipping

We propose a robust gradient estimator based on per-sample gradient clipping and analyze its properties both theoretically and empirically. We show that the resulting method, per-sample clipped SGD (PS-Clip-SGD), achieves optimal in-expectation convergence rates for non-convex optimization problems under heavy-tailed gradient noise. Moreover, we establish high-probability convergence guarantees that match the in-expectation rates up to polylogarithmic factors in the failure probability. We complement our theoretical results with multiple numerical experiments. In particular, we demonstrate that PS-Clip-SGD outperforms both vanilla SGD with momentum and standard gradient clipping when training AlexNet on the CIFAR-100 dataset, even after accounting for the additional computational time caused by per-sample clipping. We also empirically show that, in the presence of gradient accumulation, applying clipping at the mini-batch level can improve training performance while incurring virtually no additional computational cost. This finding is particularly interesting, as it contradicts the common practice of applying clipping only after all accumulation steps have been completed.
Sep 14, 2026cs.LG

Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise and H"{o}lder Smoothness

Classical convergence guarantees for stochastic gradient methods typically assume Lipschitz-smooth objectives and finite-variance gradient noise, both frequently violated in practice. In contrast, we study nonconvex stochastic optimization under the joint relaxation of these assumptions: objectives with (L,s)(L,s)-H"older continuous gradients, s∈(0,1]s\in(0,1], and gradient noise satisfying only a bounded α\alpha-th moment condition for α∈(1,2]\alpha\in(1,2]. We establish three convergence results. Firstly, that standard SGD converges at rate O(T−s/(1+s))O(T^{-s/(1+s)}) whenever α≥1+s\alpha\ge1+s, extending the classical nonconvex SGD rate to heavy-tailed noise and H"older smoothness simultaneously. Secondly, we analyze δ\delta-regularized gradient clipping (δ\delta-GClip), a provable trainer of wide and deep nets, and establish a stationarity rate of O(T−2s(α−1)/[(1+s)(2α−1)])O(T^{-2s(\alpha-1)/[(1+s)(2\alpha-1)]}) under the same condition. Thirdly, we analyze standard gradient clipping (G-Clip) and show that it recovers the above rate for α≥1+s\alpha\ge1+s while in the very heavy-tailed regime α<1+s\alpha<1+s, it has a convergence rate O(T−2s(α−1)/[(α−1)+s(2α−1)])O(T^{-2s(\alpha-1)/[(\alpha-1)+s(2\alpha-1)]}) --- the first convergence guarantee in this regime for any stochastic gradient based method.
Sep 10, 2026cs.LG

Almost Sure Convergence Analysis of Stochastic Gradient Methods with Clipping and Additive Noise

Stochastic gradient descent (SGD) with gradient clipping and additive noise has become a standard technique for training machine learning models, particularly in applications requiring robustness or privacy guarantees. However, clipping introduces a bias in stochastic gradients, while additive noise introduces additional variance, making the long-run behaviour of individual optimization trajectories difficult to characterize. In this work, we prove that SGD with clipping and additive Gaussian noise (SGD-CN) converges almost surely (a.s.) under smoothness and uniformly bounded stochastic-gradient noise assumptions, provided the step sizes satisfy some standard decaying conditions. Our analysis extends to momentum variants such as the stochastic heavy ball and Nesterov's accelerated gradient, where we show that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for understanding the pathwise behaviour of clipped stochastic gradient methods and suggest that, despite the bias and noise introduced by clipping and perturbation, the algorithm remains stable in both convex and nonconvex regimes.