Organizations: John A. Paulson School of Engineering and Applied Sciences, Harvard University, Cambridge, MA · Center for Brain Science, Harvard University, Cambridge, MA · Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University, Cambridge, MA
A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, yet its consequences for learning-curve and resource scaling remain unclear. We isolate this mechanism by studying normalized SGD in a random-feature model with power-law teacher and data covariance. Fixed-norm updates induce an effective learning rate that grows as gradients shrink. We derive a dynamical mean-field theory (DMFT) describing the joint dependence of the loss on training time, model width and batch size. Normalization initially accelerates SGD, mapping the power-law exponent rSGD<1 to 2rSGD/(1−rSGD), with exponential convergence at rSGD=1 and formal finite-time convergence for rSGD>1. At finite step size, however, the same feedback ultimately breaks the acceleration and leads to marginal stability. The late-time theory yields width-limited, edge-of-stochastic-stability (EoSS), and deterministic edge-of-stability (EoS) regimes. These phases determine when larger batches or wider models reduce serial training time at comparable compute. We quantify in which of these phases increased batch size or width can compensate the excess compute use per step by fewer optimization steps to target loss. Linearized ResNet experiments on CIFAR-5M support the predicted acceleration, breakdown, and resource-scaling trends. Together, these results connect normalization-induced acceleration, EoS effects, and width--batch allocation within a solvable theory.
Figures & tables
Figure 1: Three transient convergence regimes. Loss ( top ) and effective learning rate ( bottom ) at large α,ν . Before tstab , simulations and DMFT follow the continuous-time prediction: formal finite-time convergence for a>1/2 ( left ), exponential for a=1/2 ( center ), and accelerated power-law convergence for a<1/2 ( right ). SGD uses ηSGD=ηeff(0) to match the initial update norm. At tstab , ηeff reaches 2/λmax and the discrete dynamics depart from the flow prediction.
Figure 2: Three late-time regimes: width limited, EoSS, and EoS. Left: DMFT phase diagram in the (α,ν) plane. Color denotes ηeff∞ , solid curves phase boundaries, and perforated circles denote norm-SGD simulations; colors that blend well indicate a good match to the theory. Increasing ν removes the width bottleneck, while increasing α suppresses minibatch noise and drives the EoSS–EoS transition; within EoS, the stability threshold depends weakly on ν . Center: ηeff∞ versus ν , showing width-limited behavior and the phase boundary between batch-dependent EoSS branches to the common EoS branch as α increases. Right: ηefft at ν=1 . Dots are simulations, dashed curves are DMFT, and horizontal dotted lines are stationary predictions.
Figure 3: Spectral control of batch-size scaling and saturation. Left: Stationary effective learning rate ηeff∞ versus batch ratio α=B/d , at fixed width ratio ν=N/d . Solid curves show DMFT predictions and circles show simulations. Dashed lines indicate the small-batch linear approximation 2α/μ1 , with μ1=⟨λk⟩k , and dotted lines indicate the curvature ceiling 2/λmax(ν) . Stars mark the EoSS–EoS transition; the vertical line marks B=1 . Flatter spectra support a more extended increase in effective learning rate before saturation. Center: Training different batch sizes α on a=0.45 , b=0.2 data. With equal compute, a larger batch size achieves the same level of performance faster (fewer steps). The inset shows the loss as a function of compute with a nearly collapsed curve across all batch sizes within a certain range. Right: The excess loss above the loss floor in equation D.14 . Training curves for different batches collapse to one master curve when measures in time τ(t)=∑t′tηefft′ , implying that ηefft′,α together control compute vs. time tradeoff.
Figure 4: Spectral dependence of width efficiency. Stationary effective time per unit compute, quantified by Qν=ηeff∞/ν . Left: Width dependence of Qν at fixed b=3 , for different teacher exponents a . Solid curves show DMFT predictions before EoS, dashed curves indicate the EoS branch, and markers show simulations. Increasing Qν corresponds to superlinear growth of the effective learning rate with width; a maximum identifies an optimal width according to this stationary efficiency measure. Center: Dependence of Qν on the teacher and data exponents at fixed ν=0.1 . The background shows DMFT predictions, while circles show simulation estimates using the same logarithmic color scale. Dotted reference lines indicate the fixed width and data exponent linking the two panels. Both panels use d=2000 , B=200 , and fixed bare learning rate η=10−5 . Right: Loss as a function of steps (inset: compute) showing an example where wider networks train faster, while requiring approximately the same compute as narrower networks.
Figure 5: Theory predicts training curves exponents for training with linearized networks. Loss ( top ) and effective learning rate ( bottom ). Before ηeff hits its ceiling ηeff∞ , the loss decreases with an accelerated power law, after the ceiling is reached the power law degrades to the standard SGD power law. Throughout the narrowest width and smallest batch size are optimal, as predicted in the previous section for these a,b . ( left ) Loss as a function of images seen, Bt , for different batch sizes and fixed w=128 , η=0.1 (proportional to compute). ( right ) Loss as a function of compute, calculated as 6NBt , for different widths and fixed B=1,000 , η=0.01
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Critical batch ratio αc versus width ratio ν . Solid curves show DMFT predictions, the dashed line shows the leading top-mode estimate ν/d , and markers show the large-dimension approximation ν/d+d1∑k=2dkb−11 . The remaining spectrum raises the saturation threshold substantially for smaller b .
Figure 7: Width efficiency and spectral dependence. Left: Qν=ηeff∞/ν versus width for b=0.5 and different teacher exponents a . DMFT predictions (solid; dashed in EoS) are compared with simulations (symbols) and anchored Phase-1 scalings (dotted), which predict Qν∝ν−1/2 for a>1/2 , with a logarithmic correction at a=1/2 . Right: Qν across (a,b) at ν=0.1 , with theory and simulation markers sharing the color scale. The thick white curve marks the EoS boundary; the dark contour marks Qν=1 . For these parameters, no interior efficiency maximum is observed for ν<1 : Qν(ν) decreases to a minimum before rising toward a peak near the interpolation threshold ν=1 .
Figure 8: Width efficiency at b=1 . Left: Qν=ηeff∞/ν versus width for different teacher exponents a . DMFT predictions (solid; dashed in EoS) agree with simulations (symbols). Dotted curves show anchored Phase-1 scalings, including the logarithmic corrections at b=1 ; for a>1/2 , Qν∝[νlog(1/ν)]−1/2 . Efficiency initially decreases, reaches a minimum, and rises toward a peak near ν=1 . Right: Qν across (a,b) at ν=0.1 , with theory and simulation markers sharing the color scale. The thick white curve marks the EoS boundary, and the dark contour marks Qν=1 . The vertical and horizontal dotted reference lines indicate ν=0.1 and b=1 , respectively.
Figure 9: Complimentary figure to 4 right panel. Here we train a larger range of widths to convergence, showing that the width speedup can sometimes be subsumed by the loss floor, making larger widths for compute and speed efficient.
Normalization renders large parts of neural networks effectively scale invariant, inducing a hidden feedback loop in which learning-rate schedules and weight decay interact through the parameter norm to control the effective step taken by the optimizer. We show that this interaction is governed by an exact discrete-time law: a single scalar quantity captures all schedule and decay forcing, while norm growth induces an opposing geometric self-quenching effect. This yields a sharp boundary that cleanly separates contraction- and expansion-dominated effective learning rate regimes. To understand the underlying mechanism, we provide exact analysis of a fully solved normalized regression model where the dynamics reduce to two dimensions and show that the balance point is intrinsically unstable, implying that constant learning rate with weight decay cannot stably maintain an interior equilibrium and instead produces recurrent behavior driven by discrete-time Jacobian structure. We further extend this perspective across optimizers through unified homogeneous-optimizer framework that reveals a structural dichotomy in self-quenching strength, providing a first-principles explanation for why adaptive methods exhibit systematically weaker stabilization under normalization. Across dynamical systems and neural networks (MLP, CNN, GPT2 / MNIST, CIFAR, wikiText, OpenWebText), the predicted law holds with high precision and enables direct control of training via the identified scalar, with performance peaking sharply at the predicted boundary. Together, these results isolate a single governing quantity for scale-invariant optimization, providing a precise and actionable lens on training dynamics, optimizer behavior, and schedule design in modern deep learning. Code is available in https://github.com/shasanamin/normalized-optimization-dynamics.
We study scaling laws of signSGD under a power-law random features (PLRF) model that accounts for both feature and target decay. We analyze the population risk of a linear model trained with one-pass signSGD on Gaussian-sketched features. We express the risk as a function of model size, training steps, learning rate, and the feature and target decay parameters. Comparing against the SGD risk analyzed by Paquette et al. (2024), we identify a drift-normalization effect and a noise-reshaping effect unique to signSGD. We then obtain compute-optimal scaling laws under the optimal choice of learning rate. Our analysis shows that the noise-reshaping effect can make the compute-optimal slope of signSGD steeper than that of SGD in regimes where noise is dominant. Finally, we observe that the widely used warmup-stable-decay (WSD) schedule further reduces the noise term and sharpens the compute-optimal slope, when feature decay is fast but target decay is slow.
Jihwan Kim, Dogyoon Song, Chulhee Yun
Seoul National University & KAIST InnoCORE LLM · University of California, Davis · KAIST
The cooldown phase of a warmup-stable-decay (WSD) learning-rate schedule, now a default in large-model pretraining, lowers the final training loss in some settings and does nothing in others. We give a provable account of which case obtains, and it turns on two properties together: the structure of the gradient noise and whether the optimizer normalizes its update. On a strongly convex objective with multiplicative (gradient-proportional) noise, stochastic gradient descent contracts geometrically at a constant learning rate, so cooldown has nothing to improve. Under the same objective and noise, sign-based and normalized methods, the standard surrogates for adaptive optimizers, settle on a noise floor of order η2 and reach the minimizer only as the learning rate is driven to zero; any additive noise then reinstates a floor for every method. The mechanism is elementary: an SGD step shrinks in proportion to the gradient and so anneals itself, whereas a normalized step keeps unit scale and cannot. We solve the signSGD stationary law on the quadratic exactly and obtain the floor constant in closed form, prove a local form of the dissociation under (L0,L1)-smoothness, extend the floor to normalized SGD in dimension d>1 by a scale-invariance argument, and establish robustness to momentum and heavy-tailed noise. Simulation confirms every prediction, and we demonstrate the resulting noise-regime diagnostic on a real classification task with directly measured gradient noise. The mechanism explains whether cooldown helps; the interior cooldown fraction used at scale lies outside stationary landscape-and-noise geometry.
Subham Singh, Ashutosh Mishra, Subha Raut
Department of Mathematics and Statistics, Mississippi State University · Independent Researcher