stat.MLMay 21, 2026

Uniform-in-Time Weak Propagation-of-Chaos in Shallow Neural Networks

Authors: Margalit GlasgowJoan Bruna

Organizations: Massachusetts Institute of Technology · Courant Institute School of Mathematics, Computing and Data Science, New York University

Abstract

We consider one-hidden layer neural networks trained in the feature-learning regime using gradient descent, and relate the output of the finite-width network fρ^tmf_{\hatρ_t^m} to its infinite-width counterpart fρtMFf_{ρ_t^{MF}}, which evolves in the mean-field dynamics. While constant-time horizon bounds for fρtMFfρ^tm\|f_{ρ_t^{MF}} - f_{\hatρ_t^m}\| may be obtained via standard Grönwall estimates, the long-time behavior of the fluctuation is a more delicate matter. Uniform-in-time bounds often rely on (local) strong convexity in the landscape or Logarithmic Sobolev inequalities present in noisy gradient dynamics. In this work, we establish non-asymptotic weak propagation-of-chaos that holds uniformly in time, obtained by exploiting instead the convergence rate of the mean-field deterministic Wasserstein-gradient-flow dynamics. Specifically, denoting by LtL_t the mean-field excess MSE loss at time tt and mm the number of neurons, under standard regularity assumptions and the condition 0Lt1/2dt=O(logd)\int_0^\infty L_t^{1/2} dt =O(\log d), we obtain the uniform in time bound fρtMFfρ^tm2poly(d)mmin(1,c/6)\|f_{ρ_t^{MF}}- f_{\hatρ_t^m}\|^2 \lesssim \text{poly}(d) m^{-\min(1,c/6)} whenever LttcL_t \lesssim t^{-c}. Our result holds in a noiseless setting and does not make any assumptions on the geometry of the landscape near the optimum, and extends seamlessly to other forms of discretization, including finite number of samples and time discretization. A key takeaway of our result is that whenever the convergence rate of the mean-field, population-loss dynamics is faster than t2t^{-2}, we can attain a loss of εε with only poly(d/ε)\text{poly}(d/ε) neurons, training samples, and GD steps.

Explore similar work

CardsList