Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization
Organizations: University of California, Berkeley · Nubank
Abstract
Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of \emph{rounding space}: the coordinate in which a quantizer chooses between adjacent reconstruction levels. For the second moment, a local analysis of the quantization cell adjacent to zero shows that small mean state error need not imply small mean preconditioner error at the next step. A one-dimensional quadratic construction further shows qualitatively different optimization dynamics under state-space and preconditioner-space rounding. These results motivate Zero-Inclusive Preconditioner-space Stochastic Rounding (\textbf{ZIP-SR}), which retains zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space. As a complementary route, Zero-Excluding EDEN calibration (\textbf{ZE-EDEN}) uses a zero-excluding second-moment codebook and rescales the quantized second-moment block to mitigate the preconditioner distortion caused by the positive quantization floor. Both configurations use 4-bit NormalFloat (NF4) for the first moment, with targeted stochastic rounding of the LM-head first moment during the final 10% of training. Across GPT- and Llama-style pretraining experiments ranging from \textbf{130M} to \textbf{2.7B} parameters, both methods reduce TorchAO 4-bit AdamW's mean validation-loss gap to 32-bit AdamW at every evaluated model size, with the largest reported gap reduction reaching \textbf{70%}. In full-parameter supervised fine-tuning, both recipes achieve lower validation loss than TorchAO while remaining close to 32-bit AdamW on downstream tasks.
Figures & tables
| Rule | State error | Precond. error | |
|---|---|---|---|
| State-RTN | |||
| State-SR | |||
| Update-RTN | |||
| Update-SR |
| Configuration | First moment | Second moment | |||
|---|---|---|---|---|---|
| Non-LM-head | LM-head | Format | Rule | EDEN | |
| 32-bit | FP32 | FP32 | FP32 | — | |
| TorchAO 4-bit | SDyn4-RTN | SDyn4-RTN | Lin4-NZ | State-RTN | |
| ZIP-SR 4-bit | NF4-RTN | NF4-RTN/SR ∗ | Dyn4 | Update-SR | |
| ZE-EDEN 4-bit | NF4-RTN | NF4-RTN/SR ∗ | Dyn4-NZ | State-RTN | |
| Step | Setting | First moment | Second moment | Gap | vs. parent | |||
| Format | Format | Rule | EDEN | Gain | ||||
| 0 | TorchAO 4-bit | SDyn4 | Lin4-NZ | State-RTN | — | — | ||
| 1 | Dyn4-NZ control | SDyn4 | Dyn4-NZ | State-RTN | ||||
| 2 | NF4 first moment | NF4 | Dyn4-NZ | State-RTN | ||||
| 3a | ZE-EDEN | NF4 | Dyn4-NZ | State-RTN | ||||
| 3b | ZIP-SR | NF4 | Dyn4 | Update-SR | ||||
| Family | Size | 32-bit loss | 4-bit gap to 32-bit AdamW | Best reduction vs. TorchAO (%) | ||
|---|---|---|---|---|---|---|
| TorchAO | ZE-EDEN | ZIP-SR | ||||
| GPT-style | 162M | 32.8 | ||||
| 405M | 34.3 | |||||
| 1.4B | 70.1 | |||||
| 2.7B | +1.5435\smash{\makebox[0.0pt][r]{\raisebox{4.30554pt}{\dagger}\hskip-3.00003pt}} | — | ||||
| Llama-style | 130M | 16.6 | ||||
| Model | Configuration | Tulu val. loss | MMLU | GSM8K | HumanEval | IFEval strict |
|---|---|---|---|---|---|---|
| Qwen3- 8B-Base | 32-bit | |||||
| TorchAO 4-bit | ||||||
| ZE-EDEN 4-bit | ||||||
| ZIP-SR 4-bit | ||||||
| Llama- 3.2-3B | 32-bit | |||||
| TorchAO 4-bit |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Size | Heads | FFN | |||
|---|---|---|---|---|---|---|
| GPT-style | 162M | 12 | 768 | 12 | — | 3072 |
| 405M | 24 | 1024 | 16 | — | 4096 | |
| 834M | 24 | 1536 | 16 | — | 6144 | |
| 1.4B | 24 | 2048 | 16 | — | 8192 | |
| 2.7B | 32 | 2560 | 32 | — | 10240 | |
| Llama-style | 130M | 12 | 768 | 12 | 4 | 2304 |
| Family | Size | Candidate peak LRs | Selected LR | Val. loss |
| GPT-style | 162M | 3.135 | ||
| 405M | 2.818 | |||
| 834M | 2.641 | |||
| 1.4B | 2.517 | |||
| 2.7B | Not swept | — | ||
| Llama-style | 130M | 2.698 |
| Dyn4-NZ (State-RTN) | Dyn4 | ||
|---|---|---|---|
| Raw | + EDEN | Update-SR | |
| FP32 | — | ||
| SDyn4 | |||
| NF4 | |||
| Configuration | Validation loss | Gap to 32-bit |
|---|---|---|
| 32-bit AdamW | — | |
| TorchAO 4-bit | ||
| ZE-EDEN | ||
| ZIP-SR |
| LM-head | Non-head | Final train loss | Final val. loss |
|---|---|---|---|
| FP32 | NF4-RTN | 2.392 | 2.440 |
| NF4-RTN | FP32 | 2.529 | 2.579 |
| NF4-SR | FP32 | 2.391 | 2.440 |
| Diagnostic | NF4-RTN | NF4-SR |
|---|---|---|
| Relative RMS quantization error (%) | 9.21 | 14.75 |
| Nonzero-to-zero fraction (%) | 9.16 | 12.08 |
| Accumulated shadow-state error (%) | 20.95 | 32.18 |