Organizations: John A. Paulson School of Engineering and Applied Sciences, Harvard University, Cambridge, MA · Center for Brain Science, Harvard University, Cambridge, MA · Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard University, Cambridge, MA
A recurring design principle in modern optimizers is to decouple update magnitude from the raw gradient norm, yet its consequences for learning-curve and resource scaling remain unclear. We isolate this mechanism by studying normalized SGD in a random-feature model with power-law teacher and data covariance. Fixed-norm updates induce an effective learning rate that grows as gradients shrink. We derive a dynamical mean-field theory (DMFT) describing the joint dependence of the loss on training time, model width and batch size. Normalization initially accelerates SGD, mapping the power-law exponent rSGD<1 to 2rSGD/(1−rSGD), with exponential convergence at rSGD=1 and formal finite-time convergence for rSGD>1. At finite step size, however, the same feedback ultimately breaks the acceleration and leads to marginal stability. The late-time theory yields width-limited, edge-of-stochastic-stability (EoSS), and deterministic edge-of-stability (EoS) regimes. These phases determine when larger batches or wider models reduce serial training time at comparable compute. We quantify in which of these phases increased batch size or width can compensate the excess compute use per step by fewer optimization steps to target loss. Linearized ResNet experiments on CIFAR-5M support the predicted acceleration, breakdown, and resource-scaling trends. Together, these results connect normalization-induced acceleration, EoS effects, and width--batch allocation within a solvable theory.
Figures & tables
Figure 1: Three transient convergence regimes. Loss ( top ) and effective learning rate ( bottom ) at large α,ν . Before tstab , simulations and DMFT follow the continuous-time prediction: formal finite-time convergence for a>1/2 ( left ), exponential for a=1/2 ( center ), and accelerated power-law convergence for a<1/2 ( right ). SGD uses ηSGD=ηeff(0) to match the initial update norm. At tstab , ηeff reaches 2/λmax and the discrete dynamics depart from the flow prediction.
Figure 2: Three late-time regimes: width limited, EoSS, and EoS. Left: DMFT phase diagram in the (α,ν) plane. Color denotes ηeff∞ , solid curves phase boundaries, and perforated circles denote norm-SGD simulations; colors that blend well indicate a good match to the theory. Increasing ν removes the width bottleneck, while increasing α suppresses minibatch noise and drives the EoSS–EoS transition; within EoS, the stability threshold depends weakly on ν . Center: ηeff∞ versus ν , showing width-limited behavior and the phase boundary between batch-dependent EoSS branches to the common EoS branch as α increases. Right: ηefft at ν=1 . Dots are simulations, dashed curves are DMFT, and horizontal dotted lines are stationary predictions.
Figure 3: Spectral control of batch-size scaling and saturation. Left: Stationary effective learning rate ηeff∞ versus batch ratio α=B/d , at fixed width ratio ν=N/d . Solid curves show DMFT predictions and circles show simulations. Dashed lines indicate the small-batch linear approximation 2α/μ1 , with μ1=⟨λk⟩k , and dotted lines indicate the curvature ceiling 2/λmax(ν) . Stars mark the EoSS–EoS transition; the vertical line marks B=1 . Flatter spectra support a more extended increase in effective learning rate before saturation. Center: Training different batch sizes α on a=0.45 , b=0.2 data. With equal compute, a larger batch size achieves the same level of performance faster (fewer steps). The inset shows the loss as a function of compute with a nearly collapsed curve across all batch sizes within a certain range. Right: The excess loss above the loss floor in equation D.14 . Training curves for different batches collapse to one master curve when measures in time τ(t)=∑t′tηefft′ , implying that ηefft′,α together control compute vs. time tradeoff.
Figure 4: Spectral dependence of width efficiency. Stationary effective time per unit compute, quantified by Qν=ηeff∞/ν . Left: Width dependence of Qν at fixed b=3 , for different teacher exponents a . Solid curves show DMFT predictions before EoS, dashed curves indicate the EoS branch, and markers show simulations. Increasing Qν corresponds to superlinear growth of the effective learning rate with width; a maximum identifies an optimal width according to this stationary efficiency measure. Center: Dependence of Qν on the teacher and data exponents at fixed ν=0.1 . The background shows DMFT predictions, while circles show simulation estimates using the same logarithmic color scale. Dotted reference lines indicate the fixed width and data exponent linking the two panels. Both panels use d=2000 , B=200 , and fixed bare learning rate η=10−5 . Right: Loss as a function of steps (inset: compute) showing an example where wider networks train faster, while requiring approximately the same compute as narrower networks.
Figure 5: Theory predicts training curves exponents for training with linearized networks. Loss ( top ) and effective learning rate ( bottom ). Before ηeff hits its ceiling ηeff∞ , the loss decreases with an accelerated power law, after the ceiling is reached the power law degrades to the standard SGD power law. Throughout the narrowest width and smallest batch size are optimal, as predicted in the previous section for these a,b . ( left ) Loss as a function of images seen, Bt , for different batch sizes and fixed w=128 , η=0.1 (proportional to compute). ( right ) Loss as a function of compute, calculated as 6NBt , for different widths and fixed B=1,000 , η=0.01
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Critical batch ratio αc versus width ratio ν . Solid curves show DMFT predictions, the dashed line shows the leading top-mode estimate ν/d , and markers show the large-dimension approximation ν/d+d1∑k=2dkb−11 . The remaining spectrum raises the saturation threshold substantially for smaller b .
Figure 7: Width efficiency and spectral dependence. Left: Qν=ηeff∞/ν versus width for b=0.5 and different teacher exponents a . DMFT predictions (solid; dashed in EoS) are compared with simulations (symbols) and anchored Phase-1 scalings (dotted), which predict Qν∝ν−1/2 for a>1/2 , with a logarithmic correction at a=1/2 . Right: Qν across (a,b) at ν=0.1 , with theory and simulation markers sharing the color scale. The thick white curve marks the EoS boundary; the dark contour marks Qν=1 . For these parameters, no interior efficiency maximum is observed for ν<1 : Qν(ν) decreases to a minimum before rising toward a peak near the interpolation threshold ν=1 .
Figure 8: Width efficiency at b=1 . Left: Qν=ηeff∞/ν versus width for different teacher exponents a . DMFT predictions (solid; dashed in EoS) agree with simulations (symbols). Dotted curves show anchored Phase-1 scalings, including the logarithmic corrections at b=1 ; for a>1/2 , Qν∝[νlog(1/ν)]−1/2 . Efficiency initially decreases, reaches a minimum, and rises toward a peak near ν=1 . Right: Qν across (a,b) at ν=0.1 , with theory and simulation markers sharing the color scale. The thick white curve marks the EoS boundary, and the dark contour marks Qν=1 . The vertical and horizontal dotted reference lines indicate ν=0.1 and b=1 , respectively.
Figure 9: Complimentary figure to 4 right panel. Here we train a larger range of widths to convergence, showing that the width speedup can sometimes be subsumed by the loss floor, making larger widths for compute and speed efficient.