The Hidden Ratio in Adam: Stable Structure, Compression, and Sign Dynamics
Authors: Yihe Zhou, Tongtian Zhu, Yingxiao Huo, Satya Prakash Dash, Can Wang, Samuel Kaski, Mingfei Sun
Organizations: Department of Computer Science, The University of Manchester, Manchester, UK · College of Computer Science and Technology, Zhejiang University, Hangzhou, China
Adam is the default optimizer for training modern deep neural networks, yet its adaptive behavior remains poorly understood due to the complex interaction between its first- and second-moment exponential moving averages (EMAs). We study Adam in the tied-β regime, where the two EMA decay rates are equal, and show that its adaptive dynamics can be expressed through a transformed ratio with approximately scale-stable behavior. Empirically, this transformed ratio exhibits a stable, heavy-tailed distribution across tasks, model scales, and training stages, in contrast to the variability of raw moment magnitudes. This empirical stability has both practical and conceptual consequences. First, we derive a recurrence for the transformed ratio, yielding a reparameterization of Adam that replaces the second moment with a compressible state. Leveraging its stable distribution, we show that a fixed 4-bit codebook is sufficient in our experiments to store this state without auxiliary scaling, achieving performance competitive with full-precision Adam. Second, the transformed ratio view clarifies Adam's connection to sign-based methods: Adam reduces to sign-based momentum modulated by the transformed ratio, and replacing it with a constant recovers Signum as a limiting case. This perspective further provides a simple rule for transferring learning rates between the two methods. Together, these results suggest that tied-β Adam admits a simple and approximately stable ratio structure underlying its adaptive behavior and demonstrate its utility for both analysis and efficient implementation.
Figures & tables
c
β=0.90
β=0.92
β=0.95
β=0.97
0
3.36
3.31
3.25
3.21
0.1
3.33
3.28
3.22
3.18
0.2
3.24
3.19
3.13
3.09
Table 1 : yconsteff under Yc=2/(Z+c)2 .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
β
Adam
TR-Adam (FP32- mt , 4-bit yt )
TR-Adam (FP8- mt , 4-bit yt )
Δ FP32
Δ FP8
Llama-style 20M
0.90
3.694
3.706
3.710
+0.012
+0.016
Llama-style 20M
0.92
3.676
3.689
3.688
+0.013
+0.012
Llama-style 20M
0.95
3.672
3.672
3.670
-0.000
-0.002
Llama-style 20M
0.97
3.692
3.677
3.675
-0.014
-0.016
Pythia-style 160M
0.90
3.373
3.378
3.387
+0.004
+0.013
Pythia-style 160M
0.92
3.355
3.365
3.369
+0.010
+0.014
Appendix
Table 2: Pre-training validation loss averaged over the final three validation points. For Llama-style 20M, the learning rate is selected from the full learning-rate sweep; for Pythia-style 160M, it is selected from {1,3,5,8,10}×10−3 . Lower is better. Here Δ=TR-Adam−Adam .
Model
β
Adam
TR-Adam (FP32- mt , 4-bit yt )
TR-Adam (FP8- mt , 4-bit yt )
Δ FP32
Δ FP8
Pythia-1B
0.90
1.299
1.297
1.297
-0.002
-0.002
Pythia-1B
0.92
1.302
1.295
1.295
-0.007
-0.006
Pythia-1B
0.95
1.309
1.299
1.299
-0.011
-0.011
Llama-3.2-1B
0.90
1.075
1.072
1.072
-0.002
-0.002
Llama-3.2-1B
0.92
1.076
1.071
1.071
-0.005
-0.005
Llama-3.2-1B
0.95
1.082
1.073
1.073
-0.009
-0.009
Appendix
Table 3: SFT validation loss averaged over the final three validation points. Lower is better. Here Δ=TR-Adam−Adam .
β
Adam
TR-Adam (FP32- mt , 4-bit yt )
TR-Adam (FP8- mt , 4-bit yt )
Δ FP32
Δ FP8
0.90
0.788
0.825
0.813
+0.036
+0.025
0.92
0.800
0.824
0.819
+0.024
+0.019
0.95
0.808
0.817
0.811
+0.009
+0.003
Appendix
Table 4: ReMax validation reward averaged over the final three validation points. Higher is better. Here Δ=TR-Adam−Adam .