cs.LGOct 6, 2026

How Bregman Divergences Shape Shampoo

Authors: Bing Liu, Wenjie Zhou, Chengcheng Zhao, Hongtao Zhang, Boao Kong, Felix Dangel, Wu Lin

Organizations: Zhejiang University · University of the Chinese Academy of Sciences · Peking University · Concordia University & Mila · University of Central Florida

Abstract

Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences can better compensate for finite-sample underestimation of the empirical second moment, helping explain the differing behavior of their corresponding Shampoo variants. We further validate this explanation through GPT-2 pretraining experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization

    Sep 3, 2025Wu Lin, Scott C. Lowe, Felix Dangel +3Deep Learning OptimizationNeural Network Optimization

  2. Rethinking Bregman Divergences in Kronecker-Factored Optimizers

    May 30, 2026Bing Liu, Wenjie Zhou, Chengcheng ZhaoSecond-Order OptimizationAdaptive Gradient Methods

  3. DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers

    Feb 2, 2026Ionut-Vlad Modoranu, Philip Zmushko, Erik Schultheis +2Deep Learning OptimizationGPU Acceleration