cs.LGJun 2, 2026

Neural Networks Provably Learn Spectral Representations for Group Composition

Authors: Jianliang HeLeda WangFengzhuo ZhangSiyu ChenZhuoran Yang

Organizations: Department of Statistics and Data Science, Yale University

Abstract

Understanding how structured internal structure emerges during neural network training is central to the study of deep learning. We investigate this phenomenon through the group composition task, where a two-layer neural network is trained to predict g1g2g_1 \star g_2 for elements of a finite group GG. By lifting the projected gradient flow to the Fourier domain, we demonstrate that the training dynamics are governed by a Riemannian gradient ascent on a representation-theoretic energy functional. We prove that, under random initialization, this flow drives each neuron to converge almost surely toward a single irreducible representation, while the cross-layer Fourier coefficients achieve a rotational rank-one alignment. This framework provides a representation-theoretic account of feature learning and characterizes a novel low-rank compression phenomenon for matrix-valued group representations. Moreover, for Abelian groups, we provide a complete population-level description: random initialization promotes uniform diversification across nontrivial representations and induces Haar-uniform phases, jointly approximating the indicator via a majority-vote mechanism. We further prove that both phase alignment and representation competition emerge with exponential convergence rates.

Explore similar work

Jun 8, 2026cs.LG

Beyond Neural Collapse: Task-Intrinsic Geometry Governs Neural Representations in Modular Arithmetic

While neural collapse (NC) predicts that a KK-class-balanced classifier should organize terminal representations as a (K1)(K-1)-dimensional simplex equiangular tight frame (ETF), modular addition consistently enters a different regime: networks compress to a two-dimensional cyclic geometry in which both classifier weights and token embeddings lie on circles. We refine the explanation of this phenomenon in three directions. First, we formalize a layerwise non-uniform training mechanism: downstream classifier weights are driven by dense cross-entropy gradients into a rank-2 equiangular configuration before upstream embeddings fully reorganize, and once this classifier plane forms, backpropagated feature gradients constrain embedding motion to the same plane while weight decay suppresses orthogonal components. Second, after this subspace locking, the induced in-plane dynamics admit an entropy-regularized transport interpretation on S1S^1; combined with modular-addition labels, this reduces embedding formation to phase alignment, whose minimizers are single-frequency characters of Z/PZ\mathbb{Z}/P\mathbb{Z} and hence equal-angle points on a circle. Third, we quantify why this solution prevails over NC: a simplex ETF gains only an O(1)O(1) advantage in cross-entropy, whereas the cyclic rank-2 solution enjoys a Θ(K)Θ(K) advantage under Schatten or weight-decay surrogates, yielding a critical threshold λcrit=Θ(1/K)λ_{\mathrm{crit}} = Θ(1/K). Our results explain both why classifier weights move first and why embeddings subsequently align with them, showing that grokking on modular arithmetic is governed not by maximal separation alone but by a task-structured trade-off between separation, symmetry, and complexity.
Hu Tan, Kuo Gai, Shihua Zhang
Sep 14, 2026cs.LG

The Rank the Task Demands: A Causal Rank Law for Matrix Memories Trained on Group Composition

Matrix-valued memories make rank the natural budget of a learned representation: the number of independent directions a state spans bounds what it can bind, compose, and track. We report causal evidence, on a group-composition testbed trained under a hard single-state bottleneck with a fixed decoder that cannot launder rank, that gradient descent recruits precisely the rank the task's algebra demands. A companion paper [Larson, 2026a] establishes the analogous recruitment and causal necessity pattern on a KK-pair associative-binding testbed, where exact recovery provably requires state rank at least KK; this paper inherits that instrument and extends the rank law from a scalar capacity bound to a representation-theoretic one. We train toward chosen minimal faithful reference representations embedded in larger matrices. On group-composition state tracking over five finite groups spanning the solvable/non-solvable divide, the recruited rank equals the group's minimal faithful real representation dimension dmind_{\min} (Spearman ρ=0.9747\rho = 0.9747, the design's tie-capped maximum), the dimension-matched solvable/non-solvable pair S4S_4/A5A_5 is statistically equivalent under a pre-registered test, and a pre-registered force-rank test separates a guaranteed similarity ceiling from empirical recovery at the target dimension: one rank below dmind_{\min}, cosine similarity is capped by the target's tied unit spectrum at (dmin1)/dmin0.894\sqrt{(d_{\min}{-}1)/d_{\min}} \le 0.894, below the 0.90.9 threshold in every group by construction, with observed cells at 86-95% (mean 91%) of that ceiling; at dmind_{\min}, not guaranteed a priori, recovery clears the pre-registered anchor-relative bar at four seeds per group in all five groups. Within this testbed, measured effective rank tracks representation dimension; the matched-dimension S4S_4/A5A_5 comparison establishes equivalence within the pre-registered tolerance.
Samuel Larson
Jun 26, 2026cond-mat.dis-nn

Spectral phase transitions and trainability in neural network learning dynamics

The emergence of low-dimensional structures in the spectra of neural network weight matrices is a common empirical feature of trained models, but the dynamical origin of this phenomenon during learning remains an open problem. We formulate neural network training as the stochastic evolution of an initially random matrix ensemble, driven by stochastic gradient descent (SGD) updates that reshape the spectral bulk while amplifying signal strength. This induces a Baik-Ben Arous-Péché (BBP) transition during training, where isolated eigenvalues detach from the random bulk distribution, providing a dynamical framework for representation formation in high-dimensional learning dynamics. We demonstrate this in a solvable linear teacher-student model, where spectral evolution is analytically tractable and a phase diagram of trainability governed by the step size (or learning rate) and initial weight variance is obtained, and subsequently extend our formalism beyond the linear regime to nonlinear and stochastic settings. Numerical simulations in realistic settings support this picture, showing robust emergence of spectral alignment during training. Our results suggest that spectral analysis may provide a unified perspective of stochastic learning dynamics, linking trainability, optimisation hyperparameters, spectral phase transitions, and representation learning in neural networks.
Chanju Park, Dario Bocchi, Francesco D'Amico +2