cs.LGSep 24, 2026

Precise Convergence Speed of Clipped SGD

Authors: David A. R. Robin

Organizations: LAMSADE, Université Paris-Dauphine, PSL Research University

Abstract

We present a tightened convergence analysis of clipped gradient descent on (L0,L1)(L_0, L_1)-smooth functions, with quantitative constants. Building on the ideas of Koloskova et al (2023), we refactor several case disjunctions to reveal the central role of a control of the bias derived from fundamental properties of ℓ2\ell_2-projection, simplifying proofs. We also extend the domain of validity from η≤1/(9β)η\leq 1 / (9 β) to η<1/βη< 1 /β where β=L0+cL1β= L_0 + c L_1 for clipping constant cc, which matches the more traditional analysis of smooth functions. We strengthen the convergence criterion from (min⁡t<TE[∥∇f(xt)∥2])\left( \min_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right) to (1T∑t<TE[∥∇f(xt)∥2])\left( \frac{1}{T} \sum_{t < T} \mathbb{E}[\lVert \nabla f(x_t) \rVert_2] \right) with matching speed, and lower the final achievable loss from O(min⁡(σ2/c,σ))\mathcal{O}(\min(σ^2/c, σ)) to the more precise 6min⁡(σ2/c,3σ)6 \min(σ^2 /c, 3 σ).

Explore similar work

CardsList
  1. Robust and Fast Training via Per-Sample Clipping

    May 4, 2026Davide Nobile, Philipp GrohsStochastic Gradient DescentGradient Clipping

  2. Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise and H"{o}lder Smoothness

    Sep 14, 2026Misbah Uz Zaman, Anirbit MukherjeeStochastic Gradient DescentLipschitz Continuity

  3. Almost Sure Convergence Analysis of Stochastic Gradient Methods with Clipping and Additive Noise

    Sep 10, 2026Amartya Mukherjee, Jun LiuStochastic Gradient DescentGradient Clipping