Organizations: National Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China · School of Artificial Intelligence, Nanjing University, Nanjing, China
This paper presents an improved analysis for sign-based methods with momentum updates. Traditional sign-based methods obtain a convergence rate of O(T−1/4) under the separable smoothness assumption, but they typically require large batch sizes or assume unimodal symmetric stochastic noise. To address these limitations, we demonstrate that signSGD with momentum can achieve the same convergence rate using constant batch sizes without additional assumptions. We also establish a convergence rate under the l2-smoothness condition, improving upon the result of prior work by a factor of O(d1/2), where d is the problem dimension. Furthermore, we explore sign-based methods in distributed settings and show that the proposed methods yield convergence rates of O(d1/2T−1/2+dn−1/2) and O(d1/4T−1/4), which outperform the previous results of O(dT−1/4+dn−1/2) and O(d3/8T−1/8), respectively. Numerical experiments also validate the effectiveness of the proposed methods.
Figures & tables
Figure 1: Results for CIFAR-10 dataset in the centralized environment.
Figure 2: Results for CIFAR-100 dataset in the distributed environment.
Method
SGDM
signSGD
EF-signSGD
AdamW
Signum
GPT-2
2.509±.005
2.236±.004
2.267±.003
2.183±.001
2.176±.002
Qwen3
1.677±.006
1.610±.003
1.609±.005
1.592±.002
1.592±.001
Table 1: Training losses of finetuning GPT-2 and Qwen3 on the Alpaca dataset.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
weight decay
0.1
batch size
256
max sequence length
512
gradient accumulation steps
1
epochs
1.0
learning rate schedule
cosine decay
Appendix
Table 2: Hyperparameter configurations.
Method
SGDM
signSGD
EF-signSGD
AdamW
Signum
lr
1e-1
1e-4
1e-0
5e-4
1e-4
β1
0.9
–
–
0.9
0.75
β2
–
–
–
0.95
–
Appendix
Table 3: Optimal hyperparameters for fine-tuning GPT-2 on Alpaca.
Method
SGDM
signSGD
EF-signSGD
AdamW
Signum
lr
5e-2
1e-5
5e-1
5e-5
1e-5
β1
0.9
–
–
0.9
0.75
β2
–
–
–
0.95
–
Appendix
Table 4: Optimal hyperparameters for fine-tuning Qwen3-0.6B on Alpaca.
Figure 3: Sensitivity results across different learning rates.
Figure 4: Sensitivity results across different momentum coefficients.
Figure 5: Stochastic gradient trajectory on the CIFAR-100 dataset.
Figure 6: Stochastic gradient trajectory on the Alpaca dataset.
State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing 210023, China · School of Artificial Intelligence, Nanjing University, Nanjing 210023, China