stat.MLMay 18, 2026

Canonical Regularisation of Wide Feature-Learning Neural Networks

Authors: George WhittlePranav VaidhyanathanJuliusz ZiomekNatalia AresMaike A. Osborne

Abstract

Wide neural networks in the feature-learning regime drive modern deep learning, and yet they remain far less studied than their kernel-regime counterparts. We consider a critical yet under-explored difference between these two regimes: the regulariser and prior implied by gradient flow training. This canonical regularisation property is well-studied in kernel regime networks -- of all the infinite global minima, gradient flow selects exactly the vanishing ridge solution -- and underpins the celebrated NN-GP correspondence, precisely allowing the modelling of noise during training. However, we prove ridge regularisation biases gradient flow in feature-learning regime networks, even in the infinitesimal limit of vanishing regularisation. Over training, ridge distorts the inductive bias of the network, with a particular damage done to pretrained networks where the implicit prior is informative. We resolve this by axiomatising the canonical regulariser as a regime-agnostic function-space energy and lift, which uniquely identifies ridge in the kernel regime, and crucially generalises to the feature-learning regime. By studying the Riemannian geometry of feature-learning networks, we derive geodesic ridge from our framework, generalising ridge to the feature-learning regime. Correspondingly, we prove the canonical function-space prior is a Riemannian Gibbs Process, generalising the more familiar Gaussian Process. As a practical contribution, we propose arc ridge as a minimax-robust, scalable surrogate to geodesic ridge, revealing a deep relationship between early stopping and canonical regularisation across learning regimes. Finally, we demonstrate the consequences of our theory empirically on both image processing and NLP transfer-learning problems.

Explore similar work

May 23, 2026cs.LG

Feature Learning in Wide Neural Networks under μμP: Identifiability and Sparse-Dictionary Decomposition of the Mean-Field Limit

We establish four structural results for feature learning in wide two-layer neural networks under the Maximal Update Parametrization (μμP). First, we prove global existence and uniqueness of the mean-field limit of noisy gradient descent under μμP, identifying the maximal admissible weight ww^* on the moment sequence of the initialization as the reciprocal parameter-moment-growth boundary, and hence the largest weighted moment class propagated by the flow. The finite-particle approximation has uniform-in-time squared-Wasserstein rate O(N1)O(N^{-1}). Second, we characterize identifiability of the mean-field limit: two admissible parameter measures induce the same network function in L2L^2 exactly when their active components agree modulo the finite-rank realization symmetry of the architecture. The orbit depth DorbD^*_{\mathrm{orb}} is separated from the moment-variety depth DvarD^*_{\mathrm{var}}. Third, under the Barron-Hermite target condition the active support of the long-time limit measure admits a sparse-dictionary decomposition: it is supported on at most SS^* atoms modulo finite-rank realization symmetry, with SS^* bounded by an explicit coefficient-threshold number. Fourth, we derive the total feature-learning-error decomposition into statistical, optimization, propagation-of-chaos, and sparse-residual components, with a target-dependent Hermite/Barron tail replacing any initialization-only residual. The four results are tied together by an architectural identity: the triple (w,Dorb,S)(w^*, D^*_{\mathrm{orb}}, S^*) -- the maximal admissible weight, the orbit identifiability depth, and the sparse-dictionary depth at which the target is realizable -- is the natural learning cell of the architecture-data pair (σ,ρ)(σ, ρ). The proofs are self-contained except for standard results from μμP and mean-field Langevin theory.
Akmal Xodarev
May 18, 2026cs.LG

Pointwise Generalization in Deep Neural Networks

We address the fundamental question of why deep neural networks generalize by establishing a pointwise generalization theory for fully connected networks. This framework resolves long-standing barriers to characterizing the rich nonlinear feature-learning regime and builds a new statistical foundation for representation learning. For each trained model, we characterize the hypothesis via a pointwise Riemannian Dimension, derived from the eigenvalues of the learned feature representations across layers. This establishes a principled framework for deriving hypothesis-dependent, representation-aware generalization bounds. These bounds offer a systematic upgrade over approaches based on model size, products of norms, and infinite-width linearizations, yielding guarantees that are orders of magnitude tighter in both theory and experiment. Analytically, we identify the structural properties and mathematical principles that explain the tractability of deep networks. Empirically, the pointwise Riemannian Dimension exhibits substantial feature compression, decreases with increased over-parameterization, and captures the implicit bias of optimizers. Taken together, our results indicate that deep networks are mathematically tractable in practical regimes and that their generalization is sharply explained by pointwise, feature-spectrum-aware complexity.
Shaojie Li, Yunbei Xu
May 11, 2026stat.ML

Sharp feature-learning transitions and Bayes-optimal neural scaling laws in extensive-width networks

We study the information-theoretic limits of learning a one-hidden-layer teacher network with hierarchical features from noisy queries, in the context of knowledge transfer to a smaller student model. We work in the high-dimensional regime where the teacher width kk scales linearly with the input dimension dd -- a setting that captures large-but-finite-width networks and has only recently become analytically tractable. Using a heuristic leave-one-out decoupling argument, validated numerically throughout, we derive asymptotically sharp characterizations of the Bayes-optimal generalization error and individual feature overlaps via a system of closed fixed-point equations. These equations reveal that feature learnability is governed by a sequence of sharp phase transitions: as data grows, teacher features become recoverable sequentially, each through a discontinuous jump in overlap. This sequential acquisition underlies a precise notion of \textit{effective width} kck_c -- the number of learnable features at a given data budget nn -- which unifies two distinct scaling regimes: a feature-learning regime in which the Bayes-optimal generalization error εBO\varepsilon^{\rm BO} scales as n1/(2β)1 n^{1/(2β)-1}, and a refinement regime in which it scales as n1n^{-1}, where β>1/2β>1/2 is the exponent of the power-law feature hierarchy. Both laws collapse to the single relation εBO=Θ(kcd/n)\varepsilon^{\rm BO}=Θ(k_c d/n). We further show empirically that a student trained with \textsc{Adam} near the effective width kck_c achieves these optimal scaling laws (up to a small algorithmic gap), and provide an information-theoretic account of the associated scaling in model size.
Minh-Toan Nguyen, Jean Barbier