Population loss can remain nearly constant while a neural network learns a substantially more predictive representation. We establish this separation for two-layer ReLU and leaky-ReLU networks trained on Gaussian inputs by simultaneous fixed-step population gradient descent on all parameters. For structured additive teachers whose links are positive mixtures of Gaussian-damped cubics in H1(γ), we give explicit conditions under which small IID Gaussian initialization yields a high-probability guarantee: at a checkpoint during a high-loss plateau, minimum alignment between the rank-r teacher subspace and the leading r-dimensional eigenspace of the predictor's average gradient outer product (AGOP) increases by at least 1/2, and the minimum refit MSE under unchanged coefficient budgets decreases by more than 0.399, both relative to initialization. The same trajectory subsequently attains a trained loss below every value in the plateau window. A complementary result treats unequal-weight cubic teachers and small additive Sobolev perturbations using projected-feature refits. For SwiGLU networks with an exactly fitted intercept, we prove leading-AGOP alignment during a loss plateau at fixed width and dimension as Gaussian initialization vanishes, for square-integrable teachers with nonzero Hermite content of degree one, two, or three. A rank-one cubic specialization also gives simultaneous unrestricted-refit gains at a prescribed width. An approximation lower bound further shows that certain interaction targets retain nonzero error when ridge neurons are restricted to shared orthogonal axes within the teacher subspace. Population-moment experiments with ReLU students across 21 teachers and 50 initializations per teacher complement the analysis.
Figures & tables
Figure 1: Features improve before loss falls; gated teachers retain mixed neurons. Fifty initializations of one ReLU-student configuration per normalized teacher ( s=X1 , t=X2 ). Read down each column: training loss stays near one, AGOP subspace alignment rises, and a same-budget readout refit improves before loss release. Shading marks the loss-only plateau shared by all fifty seeds; the labeled boundaries are the earliest individual plateau endpoints. Horizontal position is log(1+n/100) : early time is expanded, ticks show actual updates, and initialization and all 20,000 updates remain visible. For SiLU, mean Amin rises 0.028→0.986 and refit MSE falls 0.829→0.455 by update 200 while loss remains near one; refit improvement is not monotone. Row four projects each weight’s teacher-plane component onto its nearest axis in the indicated frame, retaining outside-plane components and biases, then refits. Bottom-row color shows mean importance share of projected neuron directions within the teacher plane: 0∘/90∘ are its two axes; interior angles mix their coordinates. This is not the angle to the plane and does not measure the outside-plane component. Curves use all fifty seeds; light bands are pointwise 10th–90th percentiles, not confidence intervals. Dotted alignment bounds retain unresolved spectra; red crosses mark snapped-refit solver tolerance misses. Diagnostics are recorded every 100 updates (early dots); connecting lines add no observations. Appendix A gives numerical brackets, the plateau rule and eighteen further teachers.
Figure 2: Subspace recovery before loss release at increasing teacher rank. ReLU students ( d=64 , m=256 ), ten initializations per column, learn y=r−1/2∑ih2(uiTX) at r=2,4,8,16 . Rows show training MSE, AGOP alignment, bounded-head refit MSE, neuron angle to the teacher subspace, and projected neuron angle to its nearest axis . The last two rows distinguish subspace entry from individual-axis alignment. Unlike the bottom row of Figure 1 , the bottom row here uses a projected nearest-axis angle at arbitrary rank. Means and 10th–90th percentile bands retain all ten seeds; heatmaps average separately normalized importance-weighted histograms. The dotted r/d line is an isotropic reference for Amean . Shading marks the common 5% loss-ratio prefix, with a 1% comparison boundary; the loss also remains in [0.95,1.05] . Grey heatmap regions have no shared saved state; hatching marks geometrically impossible angles. The quadratic teacher is invariant to rotations inside its subspace, so its axes are not identifiable: this is a symmetry control for coordinate mixing. Appendix B defines the angles, protocol, numerical checks and additional teacher comparisons.
Figure 3: SwiGLU students: twenty initializations per teacher. All seeds 0–19; d=16 , m=64 , ε=.05 . Loss and numerical unrestricted-refit MSE are divided by teacher variance; rank-one AGOP alignment is Atop=Amin . Navy, teal and purple consistently denote loss, AGOP alignment and refit MSE. Solid means and light 10th–90th percentile bands use all twenty seeds on their shared support; later individual tails remain visible. Shading marks the shared loss-only prefix ∣ℓ−ℓ(0)∣≤10−3 . The axis counts adaptive Euler updates, with early updates expanded. Diagnostic interpolation is for display only; bands show seed variation, not confidence intervals or integration-error bounds. Crosses, if present, mark the last finite point of censored or failed trajectories.
Figure 4: SwiGLU subspace learning and gate geometry across ranks and links. Ten initializations per setting, d=64 , m=64 ; additive quadratic, absolute-value and Gaussian-bump teachers vary rank and link. Loss and numerical refit risk are divided by teacher variance. The first three rows retain the loss, AGOP and refit conventions; the dotted r/d line references isotropic Amean . The fourth row shows effective gate angles obtained from weighted and median squared-cosine scores; these are saved scalar summaries, not per-neuron angle distributions. The last row reports the fraction of gates within 25.8∘ of a teacher axis and the fraction of axes covered by such gates. Quadratic-teacher axes remain unidentifiable. Shading uses a common 5% loss-ratio prefix, with a 1% comparison boundary; this is broader than the ∣ℓn−ℓ0∣≤10−3 prefix retained in Figure 3 . All ten seeds contribute to means and 10th–90th percentile bands; individual tails and diagnostic gaps remain visible. Gate entry and axis specialization are distinct observations; neither alone establishes a prediction gain. Appendix B gives definitions, numerical diagnostics, threshold sensitivity and rank-16 limitations.
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
Teacher
Mean ΔAmin
Mean refit decrease
Final MSE
∣z∣
[−0.013,0.727]
−0.021
0.147±0.106
st
0.972
0.769
0.003±0.001
g1/2
[0.952,0.972]
0.697
0.099±0.053
g1/3
[0.912,0.972]
0.473
0.173±0.058
g2/3
0.972
0.783
0.032±0.003
GELU(s)t
0.972
0.240
0.017±0.005
Appendix
Table 1: Baseline results for all 21 teachers and all 50 declared seeds per teacher. Endpoint gains use each seed’s exact loss-prefix endpoint. Brackets are mean lower and upper bounds, not confidence intervals; unresolved AGOPs contribute [0,1] before taking gain differences. Refits retain numerical lower and feasible upper risks. Final MSE is mean ± seed standard deviation.
Teacher
E [range]
ΔAmin
ΔAmean
Raw additive SiLU
128 [122,137]
0.222
0.576
Centered additive SiLU
203 [190,227]
0.000411
0.476
Additive ReLU
113 [108,120]
0.115
0.534
Additive sine
203.5 [191,219]
0.0231
0.487
Appendix
Table 2: Separate four-teacher followup, fifty seeds each. Endpoint columns report medians of individual-run endpoint changes, not changes at an averaged endpoint. E is the plateau endpoint update (median and range); ΔA=A(E)−A(0) for each alignment score; ΔR=R0−RE−10−5 is the conservative numerical refit gain. Later loss is the original model at update 20,000.
Teacher
ΔR
L20000
Raw additive SiLU
0.43
0.054
Centered additive SiLU
0.589
0.00056
Additive ReLU
0.313
1.2e-06
Additive sine
0.607
0.062
Appendix
Table 7
Figure 5: Additive damped and mixture links, and the ReLU gate. The two additive-family targets and the ReLU product use five diagnostic rows. Each teacher uses fifty fixed initialization seeds with identical optimization parameters. Means, numerical bounds, percentile bands and loss-prefix shading follow Figure 1 .
Figure 6: Other gated teachers. Exact GELU and tanh products, together with the bilinear rotation control, use the same five rows. Each teacher uses fifty fixed initialization seeds with identical optimization parameters. Means, numerical bounds, percentile bands and loss-prefix shading follow Figure 1 .
Figure 7: Raw and centered additive SiLU: denser followups. Each teacher has paired linear early (0–500 updates) and full (0–20,000 updates) views, with fifty seeds at d=64 , m=32 , s=0.01 . Rows show original loss, minimum and mean full-AGOP alignment, and constrained refit MSE. Bands are empirical 10th–90th percentiles; the gray trace reports the resolved fraction. Refit lower/upper numerical bounds remain separate. Shading marks the loss-prefix intersection across all fifty seeds. The common configuration was chosen after the pilot did not establish reliable full-subspace learning, as disclosed in Section A.2 . The enlarged early view exposes partial alignment and refit changes; the averages do not establish full-subspace recovery for every seed.
Figure 8: Quadratic additive teacher. The three rows display original loss, full-current-head AGOP alignment and same-budget refit MSE. Each teacher uses fifty fixed initialization seeds with identical optimization parameters. Means, numerical bounds, percentile bands and loss-prefix shading follow Figure 1 .
Figure 9: Additional positive-mixture links. Damped links and the three-scale mixture use the same three-row format. Each teacher uses fifty fixed initialization seeds with identical optimization parameters. Means, numerical bounds, percentile bands and loss-prefix shading follow Figure 1 .
Figure 10: Perturbed cubic and fourth-Hermite teachers. These standalone perturbation magnitudes and the fourth Hermite are empirical cases, not claims of the small-neighborhood theorem. Each teacher uses fifty fixed initialization seeds with identical optimization parameters. Means, numerical bounds, percentile bands and loss-prefix shading follow Figure 1 .
Figure 11: Additive ReLU and sine: denser followups. The paired early and full views use the same fifty-seed configuration, three rows, linear update axes, numerical uncertainty conventions and loss-prefix shading as Figure 7 . Both teacher means and all Hermite components are retained. Partial geometric learning and refit gains are distinct from recovery of every teacher direction; Table 2 reports continuous endpoint changes.
Figure 12: Absolute-value teacher: original baseline. This panel retains the original fifty seeds at r=2 , d=16 , m=32 , s=10−4 and initial force step h=0.05 ; it was not rerun. The three rows show original loss, full AGOP alignment and constrained refit MSE over 20,000 updates. Means, percentiles, unresolved spectral intervals and shared loss-prefix shading follow the original baseline protocol. It is separated here so the four new followups do not displace the absolute-value control.
Figure 13: ReLU teacher-link comparison at rank eight. Ten seeds each for quadratic, absolute-value and Gaussian-bump teachers, d=64,m=256 . Loss and bounded refit are in units of VY ; target means are retained and fitted by the intercept. The 5% shared plateaus and 1% boundaries use ( 26 ). Means, bands and support conventions are those defined above.
Figure 14: ReLU subspace entry and teacher-axis specialization. The same rank-eight runs as Figure 13 . Importance-weighted angle histograms distinguish entry into the teacher subspace from alignment with a single projected teacher axis. For the nonquadratic links, axis concentration develops later than subspace entry; the quadratic column is a rotation-invariant control. Five-degree histograms average the ten separately normalized seed distributions. Grey means no shared saved state, and hatching marks angles above the nearest-axis geometric limit.
Figure 15: SwiGLU rank sweep, including the weak rank-sixteen case. Quadratic teachers at ranks 2,4,8,16 , with d=64,m=64 , head-rate multiplier 0.01 and ten seeds each. The rank-sixteen column does not establish robust all-direction recovery during the plateau. All seeds and later diagnostic gaps remain visible; large mean alignment does not replace the weakest-direction and same-checkpoint requirements.
Figure 16: SwiGLU teacher-link comparison at rank eight. Quadratic, absolute-value and Gaussian-bump links use the same ten-seed protocol with head-rate multiplier 0.01 . Loss and numerical unrestricted-refit risks are divided by VY . These are empirical trajectories under declared loss tolerances, not guarantees for every link or initialization.
Figure 17: Slower SwiGLU heads help some runs without resolving the rank limit. Quadratic teachers at ranks eight and sixteen compare head-rate multipliers 0.01 and 0.001 , with ten paired seeds per setting. The curves compare how head-update speed affects loss and alignment; the rank-sixteen setting remains a limitation. The settings remain distinct cohorts; outcomes are not pooled or selected by performance.
Figure 18: SwiGLU gate geometry across ranks. The same standard-head rank sweep as Figure 15 . Gate-summary angles from ( 29 ), together with the dotted value-weight angle in ( 30 ), can improve even where the weakest AGOP direction is not recovered during the plateau. They describe different quantities, and weighted gate alignment alone is not an all-direction recovery result. The quadratic teacher has no identifiable internal axes.
Figure 19: SwiGLU gate entry and axis specialization for different links. The same rank-eight teacher comparison as Figure 16 . Effective gate and value angles and thresholded gate-axis counts use saved scalar summaries; no per-neuron angle distributions are inferred. For the nonquadratic teachers, the fraction of near-axis gates rises later than the decline in the effective subspace angle. The weighted angle and unweighted gate counts must be interpreted together and relative to the stated loss window; they do not classify every neuron or establish that coordinate mixing is necessary.
Figure 20: Fixed representative seed 0 for the four rank-one teacher links. Rows show original normalized loss, leading AGOP alignment and numerical frozen-feature refit risk. These are adaptive Euler flow approximations, with d=16,m=64,ϵ=.05 . The display includes the subsequent measured loss decrease.
Figure 21: All three predeclared seeds for every rank-one link, using the same scales and diagnostics as Figure 20 . Each curve is an individual run; no seed is selected by its outcome.
Figure 22: Fixed seed 0 for balanced and unbalanced rank-two cubic teachers. In the alignment row, solid curves are Atop and dashed curves are Amin . Strong leading-direction alignment can coexist with weak recovery of the least aligned teacher direction.
Figure 23: All three seeds for each rank-two coefficient vector. The weakest-direction statistic is retained for every seed, including cases with poor simultaneous recovery.
Figure 24: Width-256 rank-one reference reproduction ( d=16,ϵ=.1 , seed 0). The refit panel reports attained prediction risk from the unregularized numerical head solve. The triangle is the separately validated reference checkpoint 66, whose corrected risk is .232108 ; it does not change the primary prefix. The width exceeds (d+1)(d+2)/2=153 . Quadrature and spectral-cutoff checks accompany these finite-initialization values; the display does not assert an exact zero-risk fit.
Figure 25: Fixed mean-field steps Δt=.3,.1,.05 versus the adaptive flow approximation, all with the same rank-one cubic seed 0 initialization. Parameter-space learning rates are mΔt . The x-axis is accumulated mean-field time, not optimizer step count. The right column retains numerically divergent controls on their full loss scale, while omitting alignment/refit after the reference safeguard activates; the left column displays finite controls and the adaptive reference. The nominal time horizon is 250, checked on the accumulated update clock with at most one-step overshoot (actual endpoint times are retained); any trajectory reaching that horizon is censored, and no long-time convergence is claimed. These controls do not convert measured release into a fixed-step release theorem.
Figure 26: The two predeclared reference smoke checks ( d=4,m=6 , seed 0), retained for completeness. The rank-one smoke uses ϵ=.1 and the rank-two smoke ϵ=.05 .
Case (seed)
tP
ℓ0→ℓP
ℓf
Smoke r=2 (0)
445.6
1.000000→0.999418
0.8994
Smoke r=1 (0)
235.4
1.000000→0.999461
0.8977
h3 (0)
226.4
1.000000→0.999039
0.4999
h3 (1)
266.2
1.000000→0.999403
0.4997
h3 (2)
311.6
1.000000→0.999401
0.5000
h3 , m=256 (0)
55.5
1.000000→0.999037
0.8994
Appendix
Table 3: Every declared SwiGLU trajectory at the loss-selected prefix with δ=10−3 . P is the last diagnostic inside the prefix, so tP may precede the exact last admissible update (both clocks are retained in the audit). R is the attained prediction risk of the equilibrated 10−12 eigencutoff solve, not a certified unrestricted infimum. -- means the weakest-direction statistic duplicates Atop at rank one. The reference smokes use d=4,m=6 ; the scientific runs use d=16,m=64,ϵ=.05 , except m=256,ϵ=.1 . The three GD controls share seed 0; † marks the stated time/step analysis horizon reached before the loss stopping target, rather than convergence; ‡ denotes the retained numerical divergence of the Δt=.3 control. Loss values are divided by the teacher variance.
Case (seed)
Atop0→AtopP
Amin0→AminP
R0→RP
Smoke r=2 (0)
0.193→0.995
0.121→0.256
1.000→0.961
Smoke r=1 (0)
0.113→0.987
–
1.000→0.960
h3 (0)
0.128→0.998
–
1.000→0.836
h3 (1)
0.022→0.994
–
1.000→0.629
h3 (2)
0.000→0.994
–
1.000→0.642
h3 , m=256 (0)
0.150→0.995
–
0.818→0.043
Appendix
Table 31
Case (seed)
tP
ℓ0→ℓP
ℓf
Smoke r=2 (0)
432.5
1.000000→0.999930
0.8994
Smoke r=1 (0)
224.9
1.000000→0.999938
0.8977
h3 (0)
221.5
1.000000→0.999945
0.4999
h3 (1)
262.4
1.000000→0.999927
0.4997
h3 (2)
307.6
1.000000→0.999930
0.5000
h3 , m=256 (0)
52.7
1.000000→0.999945
0.8994
Appendix
Table 4: Every declared SwiGLU trajectory at the loss-selected prefix with δ=10−4 . P is the last diagnostic inside the prefix, so tP may precede the exact last admissible update (both clocks are retained in the audit). R is the attained prediction risk of the equilibrated 10−12 eigencutoff solve, not a certified unrestricted infimum. -- means the weakest-direction statistic duplicates Atop at rank one. The reference smokes use d=4,m=6 ; the scientific runs use d=16,m=64,ϵ=.05 , except m=256,ϵ=.1 . The three GD controls share seed 0; † marks the stated time/step analysis horizon reached before the loss stopping target, rather than convergence; ‡ denotes the retained numerical divergence of the Δt=.3 control. Loss values are divided by the teacher variance.
Case (seed)
Atop0→AtopP
Amin0→AminP
R0→RP
Smoke r=2 (0)
0.193→0.981
0.121→0.137
1.000→0.985
Smoke r=1 (0)
0.113→0.910
–
1.000→0.988
h3 (0)
0.128→0.986
–
1.000→0.933
h3 (1)
0.022→0.974
–
1.000→0.706
h3 (2)
0.000→0.977
–
1.000→0.859
h3 , m=256 (0)
0.150→0.945
–
0.818→0.098
Appendix
Table 33
Case (seed)
tf
ℓf
Status
Smoke r=2 (0)
455.2
0.8994
loss stop
Smoke r=1 (0)
242.4
0.8977
loss stop
h3 (0)
238.0
0.4999
loss stop
h3 (1)
277.7
0.4997
loss stop
h3 (2)
323.8
0.5000
loss stop
h3 , m=256 (0)
57.3
0.8994
loss stop
Appendix
Table 5: Terminal observations under the stopping rules in Section U.6 , for all 24 conditions. “Loss stop” means the prescribed ℓstop was reached; differences below the target reflect update-grid undershoot, not comparable unconstrained final losses. tf,ℓf use the complete update history inside the stated analysis budget; D is the latest saved diagnostic within that budget. For balanced rank-two seed 2, tf=471.9 is update 3000, whereas the last diagnostic is update 2928 at tD=471.5 ; the raw 60-update tail remains archived and dotted in the plot. Failed-control post-safeguard metrics are omitted. Refit risks are numerical attained risks, subject to the quadrature limitations in the text; in particular the fixed .1 endpoint risk changes by .01133 on doubling quadrature.
Case (seed)
tD
AtopD
AminD
RD
Smoke r=2 (0)
455.2
1.000
0.463
0.856
Smoke r=1 (0)
242.4
1.000
–
0.845
h3 (0)
238.0
1.000
–
0.445
h3 (1)
277.7
1.000
–
0.291
h3 (2)
323.8
1.000
–
0.332
h3 , m=256 (0)
57.3
1.000
–
0.077
Appendix
Table 35
Teacher
ΔAtop
ΔR
h3
0.962 [0.745, 0.996]
0.222 [0.146, 0.491]
h3+.3h2
0.956 [0.742, 0.996]
0.475 [0.320, 0.672]
ReLU
0.927 [0.719, 0.997]
0.519 [0.414, 0.683]
tanh
0.942 [0.709, 0.983]
0.476 [0.299, 0.676]
Appendix
Table 6: Twenty-seed rank-one outcomes. Alignment and refit gains are median [minimum, maximum] at the final saved in-prefix diagnostic, using all twenty values unless a parenthesized resolved count n is shown.
Figure 27: All twenty recorded trajectories per rank-one teacher. Individual curves retain their own saved grids and terminal times. The three rows and numerical meanings are those of Figure 3 ; all outcomes remain visible.
h3
h3+.3h2
ReLU
tanh
Seed
ΔA
ΔR
ΔA
ΔR
ΔA
ΔR
ΔA
ΔR
0
0.870
0.164
0.869
0.426
0.862
0.635
0.852
0.518
1
0.972
0.371
0.969
0.397
0.957
0.647
0.961
0.676
2
0.994
0.358
0.993
0.558
0.997
0.428
0.956
0.460
3
0.981
0.235
0.979
0.483
0.969
0.551
0.955
0.459
4
0.745
0.239
0.742
0.360
0.719
0.486
0.709
0.438
Appendix
Table 7: Every seed’s endpoint changes in leading AGOP alignment and normalized numerical refit MSE. Positive ΔR means lower error. Values are rounded numerical diagnostics, not error certificates; -- indicates an unresolved value.
Department of Electrical and Computer Engineering University of Southern California Los Angeles, CA, USA · Department of Computer Science University of Southern California Los Angeles, CA, USA
INRIA LMO, Université Paris-Saclay Orsay, France · Mathematics Institute University of Warwick Coventry, UK · Department of Computer Science University of Warwick Coventry, UK