cs.LGOct 8, 2026

Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD

Authors: Ruoyu Zhao, Mingxuan Zhang, Jianbo Dai, Jiaqi Wu, Chenyu Zhu, Tong Che

Organizations: City University of Hong Kong · Microsoft · Copula Lab · NVIDIA Research

Abstract

Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining. We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism. In a two-unit ReLU regression, a source task drives the two neurons toward positive proportionality and a target task rewards separating them. After source training for time TT, gradient flow recovers on the target in time linear in TT. Online SGD with batch size bb and step size ηη in both phases instead fails with high probability throughout a horizon of order ec/ηe^{c/η} once T≳log⁡(b/η)T \gtrsim \log(b/η), uniformly on an explicit set of initializations with Gaussian probability above one percent. For each fixed TT, small-step SGD still recovers, so the failure requires the joint limit of small steps and long pretraining. At the target clone, the population instability is carried entirely by inputs on which the two ReLU gates disagree. For units at angle δδ these inputs form a wedge of probability δ/πδ/π, and weight decay shrinks the angle exponentially during pretraining. On every other input both units receive the same random linear update, which contracts their separation in conditional expectation. Bounding the cumulative probability of sampling the wedge along the exact online recursion, without a diffusion approximation, shows that recovery with fixed probability from an identical source-gradient-flow checkpoint, within ec/ηe^{c/η} updates, requires Nb≳eλTNb \gtrsim e^{λT} target samples and batch size b≳ηeλTb \gtrsim ηe^{λT}, where NN counts updates and λλ is the weight decay. In simulations, recovery is approximately a function of the disagreement budget bδ/ηbδ/η and saturates in the horizon.

Figures & tables

Explore similar work

CardsList
  1. Awakening of the Buddha: Subspace Learning During Population-Loss Plateaus

    Sep 30, 2026Akash KumarRepresentation LearningReLU Neural Networks

  2. Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics

    May 21, 2026Igor Ignashin, Anna Radovskaya, Andrew Semenov +7Neural Network Training DynamicsGradient Descent Dynamics