cs.LGSep 30, 2026

Exact information accounting for SGD methods

Authors: Akshay Balsubramani

Abstract

As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis of stochastic gradient descent (SGD) and its variants. We show that a preconditioned SGD step is the posterior-mean update of a Gaussian Bayes model, and that its one-step regret splits into an intrinsic-time cost and a change in comparator information. The split extends to an identity for the objective itself. Convex convergence, strict-saddle-point escape, the link between flatness and generalization, the standard learning-rate schedules, adaptive optimizers, and the noisy, momentum, heavy-tailed, and gradient-free variants of SGD each correspond to a term or a special case of this identity. We measure its terms on synthetic and real training runs. On real networks it attributes the slack of classical convergence bounds to the terms their derivations drop and separates optimizers that reach the same training loss. That separation follows the number and consistency of their steps. Its relation to which of them generalizes better differs between networks. For gradient-free SGD the identity determines how a curvature preconditioner should enter the update. The sharpness-based generalization certificate it yields, with a data-independent isotropic prior, is vacuous at network scale unless the curvature spectrum is nearly flat across all parameters.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Information-Theoretic Generalization Bounds for Stochastic Gradient Descent with Predictable Virtual Noise

    Apr 30, 2026Mohammad PartohaghighiStochastic Gradient DescentGeneralization Bounds

  2. Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models

    Jul 6, 2026Stanislas Strasman, Sobihan Surendran, Sylvain Le CorffStochastic Gradient DescentGenerative Models