On the O(K1/4d) Convergence Rate of AdamW Measured by ℓ1 Norm
Authors: Huan Li, Yiming Dong, Zhouchen Lin
Organizations: Institute of Robotics and Automatic Information Systems, College of Artificial Intelligence, Nankai University, Tianjin, China. · National Key Lab of General AI, School of Intelligence Science and Technology, Peking University, Beijing, China.
As the default optimizer for training large language models, AdamW has achieved remarkable success in deep learning. However, its convergence behavior is not theoretically well-understood. This paper establishes the convergence rate K1∑k=1KE[∣∣∇f(xk)∣∣1]≤O(K1/4dC) for AdamW measured by ℓ1 norm, where K represents the iteration number, d denotes the model dimension, and C matches the constant in the optimal convergence rate of SGD. Theoretically, we have ∣∣∇f(x)∣∣2≪∣∣∇f(x)∣∣1≤d∣∣∇f(x)∣∣2 for any high-dimensional vector x and E[∣∣∇f(x)∣∣1]≥π2dE[∣∣∇f(x)∣∣2] when each element of ∇f(x) is generated from Gaussian distribution N(0,1). Empirically, our experimental results on real-world deep learning tasks reveal ∣∣∇f(x)∣∣1=Θ(d)∣∣∇f(x)∣∣2. Both support that our convergence rate can be considered to be analogous to the optimal K1∑k=1KE[∣∣∇f(xk)∣∣2]≤O(K1/4C) convergence rate of SGD in the ideal case. We also extend our result to NAdamW, an AdamW variant that employs a double-momentum mechanism, and demonstrate that it maintains the same convergence rate.
Figures & tables
Figure 1 : Illustration of average training loss f(xk) for AdamW over epochs/steps, and at the initialization, f(x1)≤8 .
Figure 2 : Illustration of ∥∇f(xk)∥1=Θ(d)∥∇f(xk)∥2 for AdamW over epochs/steps. The gradient norm ratio shows ∥∇f(xk)∥2∥∇f(xk)∥1 , and d=4868 , 5060 , and 11136 , respectively.
Figure 3 : Illustration of ∥xk∥∞<λ1 for AdamW over epochs/steps. The model ℓ∞ norm shows ∥xk∥∞ , and λ=0.01 , 0.1 , and 0.05 , respectively.
Figure 4 : Illustration of small dσs2 over epochs/steps. The magnitude σs2 is approximated by ∥gk−∇f(xk)∥2 for AdamW without taking expectation, and d=2.37×107 , 2.56×107 , and 1.24×108 , respectively.
Figure 5 : Illustrations of k1∑t=1k∣∇f(xt)∣ (left) and xk (right) over steps on the toy example.