math.OCOct 5, 2026

Last-Iterate Convergence Rate of Normalized Gradient Descent under Hölder Smoothness

Authors: Yuki Takezawa, Eduard Gorbunov

Organizations: Toyota Motor Corporation · MBZUAI

Abstract

Normalized gradient descent is a widely studied adaptive optimization method. Most existing analyses focus on the best iterate or a weighted average of the iterates, whereas practical implementations typically return the last iterate. In this paper, we study the last-iterate convergence of normalized gradient descent for convex, (ν,Mν)(ν,M_ν)-Hölder-smooth objectives. For a constant stepsize, we establish an upper bound of O((log⁡2(T)/T)(1+ν)/2)\mathcal{O}\bigl((\log^2(T)/T)^{(1+ν)/2}\bigr), which contains a logarithmic overhead relative to the known O(T−(1+ν)/2)\mathcal{O}\bigl(T^{-(1+ν)/2}\bigr) guarantees for the best and weighted-average iterates. For ν=0ν= 0, this overhead is known to be unavoidable. We complement this analysis with numerical results based on the performance estimation problem (PEP), investigating the finite-horizon worst-case behavior in the smooth setting and whether the logarithmic overhead reflects an intrinsic limitation of constant-step normalized gradient descent. We then show that a linearly decreasing stepsize yields a last-iterate guarantee of O(T−(1+ν)/2)\mathcal{O}\bigl(T^{-(1+ν)/2}\bigr), matching the order of the best-iterate/weighted-average guarantees without requiring knowledge of νν and MνM_ν.

Figures & tables

Explore similar work

CardsList
  1. A lower bound for stepsize-based acceleration of gradient descent

    Aug 11, 2026Jianhao Ma, Yuxin ChenStep AccuracyGradient Descent

  2. Learning from the Descent Direction: Adaptive Gradient Descent under One-Sided Hölder Regularity

    Jul 24, 2026Arzu Ahmadova, Ismail HuseynovGradient DescentDescent