cs.LGSep 30, 2026

Awakening of the Buddha: Subspace Learning During Population-Loss Plateaus

Authors: Akash Kumar

Organizations: Department of Computer Science & Engineering University of California-San Diego

Abstract

Population loss can remain nearly constant while a neural network learns a substantially more predictive representation. We establish this separation for two-layer ReLU and leaky-ReLU networks trained on Gaussian inputs by simultaneous fixed-step population gradient descent on all parameters. For structured additive teachers whose links are positive mixtures of Gaussian-damped cubics in H1(γ)H^1(γ), we give explicit conditions under which small IID Gaussian initialization yields a high-probability guarantee: at a checkpoint during a high-loss plateau, minimum alignment between the rank-rr teacher subspace and the leading rr-dimensional eigenspace of the predictor's average gradient outer product (AGOP) increases by at least 1/21/2, and the minimum refit MSE under unchanged coefficient budgets decreases by more than 0.3990.399, both relative to initialization. The same trajectory subsequently attains a trained loss below every value in the plateau window. A complementary result treats unequal-weight cubic teachers and small additive Sobolev perturbations using projected-feature refits. For SwiGLU networks with an exactly fitted intercept, we prove leading-AGOP alignment during a loss plateau at fixed width and dimension as Gaussian initialization vanishes, for square-integrable teachers with nonzero Hermite content of degree one, two, or three. A rank-one cubic specialization also gives simultaneous unrestricted-refit gains at a prescribed width. An approximation lower bound further shows that certain interaction targets retain nonzero error when ridge neurons are restricted to shared orthogonal axes within the teacher subspace. Population-moment experiments with ReLU students across 21 teachers and 50 initializations per teacher complement the analysis.

Figures & tables

Appendix figures & tables34 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Removing spurious minima for planar features by skip connections

    Oct 1, 2026Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin +1Flat MinimaSkip

  2. When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks

    Jun 3, 2026Berk Tinaz, Changzhi Xie, Mahdi SoltanolkotabiRectified Linear Unit NetworksGradient Descent

  3. Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization

    May 28, 2025Etienne Boursier, Matthew Bowditch, Matthias Englert +1Anisotropic Loss LandscapesOverparameterization