Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise and H"{o}lder Smoothness
Abstract
Classical convergence guarantees for stochastic gradient methods typically assume Lipschitz-smooth objectives and finite-variance gradient noise, both frequently violated in practice. In contrast, we study nonconvex stochastic optimization under the joint relaxation of these assumptions: objectives with -H"older continuous gradients, , and gradient noise satisfying only a bounded -th moment condition for . We establish three convergence results. Firstly, that standard SGD converges at rate whenever , extending the classical nonconvex SGD rate to heavy-tailed noise and H"older smoothness simultaneously. Secondly, we analyze -regularized gradient clipping (-GClip), a provable trainer of wide and deep nets, and establish a stationarity rate of under the same condition. Thirdly, we analyze standard gradient clipping (G-Clip) and show that it recovers the above rate for while in the very heavy-tailed regime , it has a convergence rate --- the first convergence guarantee in this regime for any stochastic gradient based method.