cs.LGMay 17, 2025

On the O(dK1/4)O(\frac{\sqrt{d}}{K^{1/4}}) Convergence Rate of AdamW Measured by ℓ1\ell_1 Norm

Authors: Huan Li, Yiming Dong, Zhouchen Lin

Organizations: Institute of Robotics and Automatic Information Systems, College of Artificial Intelligence, Nankai University, Tianjin, China. · National Key Lab of General AI, School of Intelligence Science and Technology, Peking University, Beijing, China.

Abstract

As the default optimizer for training large language models, AdamW has achieved remarkable success in deep learning. However, its convergence behavior is not theoretically well-understood. This paper establishes the convergence rate 1K∑k=1KE[∣∣∇f(xk)∣∣1]≤O(dCK1/4)\frac{1}{K}\sum_{k=1}^KE\left[||\nabla f(x^k)||_1\right]\leq O(\frac{\sqrt{d}C}{K^{1/4}}) for AdamW measured by ℓ1\ell_1 norm, where KK represents the iteration number, dd denotes the model dimension, and CC matches the constant in the optimal convergence rate of SGD. Theoretically, we have ∣∣∇f(x)∣∣2≪∣∣∇f(x)∣∣1≤d∣∣∇f(x)∣∣2||\nabla f(x)||_2\ll ||\nabla f(x)||_1\leq \sqrt{d}||\nabla f(x)||_2 for any high-dimensional vector xx and E[∣∣∇f(x)∣∣1]≥2dπE[∣∣∇f(x)∣∣2]E\left[||\nabla f(x)||_1\right]\geq\sqrt{\frac{2d}π}E\left[||\nabla f(x)||_2\right] when each element of ∇f(x)\nabla f(x) is generated from Gaussian distribution N(0,1)\mathcal N(0,1). Empirically, our experimental results on real-world deep learning tasks reveal ∣∣∇f(x)∣∣1=Θ(d)∣∣∇f(x)∣∣2||\nabla f(x)||_1=\varTheta(\sqrt{d})||\nabla f(x)||_2. Both support that our convergence rate can be considered to be analogous to the optimal 1K∑k=1KE[∣∣∇f(xk)∣∣2]≤O(CK1/4)\frac{1}{K}\sum_{k=1}^KE\left[||\nabla f(x^k)||_2\right]\leq O(\frac{C}{K^{1/4}}) convergence rate of SGD in the ideal case. We also extend our result to NAdamW, an AdamW variant that employs a double-momentum mechanism, and demonstrate that it maintains the same convergence rate.

Figures & tables

Explore similar work

CardsList
  1. Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?

    Jun 22, 2026Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2AdamLong-Tailed Distribution

  2. The Convergence Behavior of Adam under Heavy-Tailed Noise

    Jul 29, 2026Yijiang PangAdamStochastic Convex Optimization

  3. Adam Converges in Nonsmooth Nonconvex Optimization

    Jun 21, 2026Zijian LiuStochastic Convex OptimizationAdam