cs.LGMay 21, 2026

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics

Authors: Igor IgnashinAnna RadovskayaAndrew SemenovEgor LopatinStanislav PotapovAleksandr KovalenkoAndrey VeprikovAleksandr Shestakov+2 more

Organizations: Basic Research of Artificial Intelligence Laboratory (BRAIn Lab) · P.N. Lebedev Physical Institute of the Russian Academy of Sciences · Innopolis University

Abstract

Stochastic Gradient Descent (SGD) is commonly modeled as a Langevin process, assuming that minibatch noise acts as Brownian motion. However, this approximation relies on a continuous-time limit and a sqrt(eta) noise scaling that does not match the discrete SGD update at finite learning rate. In this work, we propose an alternative formulation of SGD as deterministic dynamics in a fluctuating loss landscape induced by minibatch sampling. Starting directly from the discrete update, we derive a master equation for the parameter distribution and obtain a discrete Fokker--Planck equation that differs from the standard Langevin form at order eta^2. Using this framework, we analyze SGD dynamics near critical points of the loss. We show that the behavior decomposes along the eigenbasis of the mean Hessian into qualitatively distinct regimes. In particular, nearly-flat directions do not admit a stationary distribution: the variance grows over time, corresponding to effective diffusion along valleys with a coefficient proportional to the learning rate. We provide empirical evidence supporting these predictions on neural network models in computer vision and natural language processing, observing a clear qualitative separation between confined and diffusive modes.

Explore similar work

CardsList
  1. High-dimensional Limit of SGD for Diagonal Linear Networks

    May 16, 2026Begoña García Malaxechebarría, Courtney Paquette, Maryam Fazel +1Stochastic Gradient DescentStochastic Differential Equations