How Bregman Divergences Shape Shampoo
Organizations: Zhejiang University · University of the Chinese Academy of Sciences · Peking University · Concordia University & Mila · University of Central Florida
Abstract
Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences can better compensate for finite-sample underestimation of the empirical second moment, helping explain the differing behavior of their corresponding Shampoo variants. We further validate this explanation through GPT-2 pretraining experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.
Figures & tables
| Optimizer | Method | Large-sample | 40-sample | EMA |
| AdamW | F | |||
| VN | ||||
| SQ | ||||
| KL | ||||
| Distributed Shampoo | F | |||
| VN |
| AdamW | Distributed Shampoo | |||
| Divergence | Few-sample | EMA | Few-sample | EMA |
| F | ||||
| VN | ||||
| SQ | ||||
| KL | ||||
| Method / width | 256 | 512 | 768 | 1024 | 1536 |
| F | 3.8281 | 3.5082 | 3.3993 | 3.3528 | 3.3111 |
| VN | 3.7573 | 3.4377 | 3.2876 | 3.1979 | 3.0972 |
| SQ | 3.7450 | 3.4173 | 3.2661 | 3.1824 | 3.0817 |
| KL | 3.7475 | 3.4180 | 3.2660 | 3.1769 | 3.0758 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Population objective | Factor updates discussed here |
| PSGD-Kron | SQ divergence | Original relative-gradient steps; other geometries are available |
| SQ-Shampoo | SQ divergence | Spectral EMA of SQ statistics |
| KL-Shampoo | KL divergence | Spectral EMA of KL statistics |
| Method | ||
| F-Shampoo | ||
| VN-Shampoo | ||
| SQ-Shampoo | ||
| KL-Shampoo |
| Setting | Additional parameters | ||
| Smooth long-tailed | scale , |
| Setting | Specification |
| Gradient pools | Four splits per experiment; FP32 mini-batch gradients per split, assigned by round-robin interleaving |
| Gradient collection | Batch size ; sequence length |
| Large-sample fit | All gradients in split |
| Few-sample fit | gradients from split ; four repeats |
| Fixed-point solver | Relaxation ; tolerance ; at most steps |
| Online EMA | disjoint chronological repeats from split , each with gradients; ; eigendecomposition every five steps |
| AdamW checkpoints | ||||
| Metric | Top (0–10%) | Middle (10–40%) | Lower (40–75%) | Tail (75–100%) |
| ( ) | ||||
| Signed | ||||
| Cross-fit | ||||
| Distributed Shampoo checkpoints | ||||
| Metric | Top (0–10%) | Middle (10–40%) | Lower (40–75%) | Tail (75–100%) |
| Width | Params (M) | LR | F | VN | SQ | KL |
| 256 | 22.3 | 0.0024 | 3.8281 | 3.7573 | 3.7450 | 3.7475 |
| 512 | 63.5 | 0.0024 | 3.5082 | 3.4377 | 3.4173 | 3.4180 |
| 768 | 123.6 | 0.0018 | 3.3993 | 3.2876 | 3.2661 | 3.2660 |
| 1024 | 202.5 | 0.0012 | 3.3528 | 3.1979 | 3.1824 | 3.1769 |
| 1536 | 417.0 | 0.0008 | 3.3111 | 3.0972 | 3.0817 | 3.0758 |