cs.LGSep 29, 2026

Grokking through the Lens of Minimum-Norm Interpolation

Authors: Gil Kur, Ileana Rugina, Clémentine Carla Juliette Dominé, Marco Mondelli

Organizations: ETH Zürich · Institute of Science and Technology Austria · Harvard University

Abstract

Grokking shows that fitting the training data and learning the underlying signal can occur at very different stages. However, existing theories offer limited quantitative insight into how this delayed generalization depends on inductive bias and signal structure. Our work addresses the gap by developing a statistical theory that characterizes how regularization geometry and signal sparsity govern generalization near interpolation. In particular, we focus on the prototypical setting of high-dimensional regression and identify regimes in which sparsity-promoting regularization makes exact interpolation much more accurate than approximate fitting. In strongly overparameterized noiseless problems, we prove a zero--one generalization law and construct a family of convex norms whose interpolators transition from the trivial risk of the all-zero predictor to exact recovery, while keeping the training error equal to 00. Furthermore, when feature dimension and sample size are proportional, we provide a precise characterization of training and generalization errors along ℓr\ell_r-regularization paths. This in turn allows us to quantify the generalization gain that remains near interpolation: we show that this gain increases as the norm becomes more sparsity-promoting and as the target becomes sparser, with a sharp drop in generalization reached for noiseless data and ℓ1\ell_1 regularization. Experiments on diagonal linear networks and transformers trained on modular arithmetic demonstrate the generality of our theoretical predictions. Finally, beyond grokking, our work reveals a statistical instability in minimum-norm interpolation: small perturbations in the regularization strength can lead to drastically different generalization, while preserving small training error.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 20, 2024math.ST

Generalization error of min-norm interpolators in transfer learning

This paper establishes the generalization error of pooled min-ℓ2\ell_2-norm interpolation in transfer learning, where data from diverse distributions are available. Min-norm interpolators arise naturally as implicit regularized limits of modern machine learning algorithms. Prior work has characterized their out-of-distribution risk when samples from the test distribution are unavailable during training. In many applications, however, limited test samples may be available at training time, yet properties of min-norm interpolation in this regime remain poorly understood. We address this gap by characterizing the bias and variance of pooled min-ℓ2\ell_2-norm interpolation under both covariate shift and model shift. Our results yield several important implications. In certain cases under model shift, we show that adding data always hurts when the signal-to-noise ratio (SNR) is low. At higher SNR levels, transfer learning is beneficial provided the shift-to-signal ratio falls below a threshold that we characterize explicitly. Under covariate shift, we find that when the source sample size is small relative to the dimension, greater heterogeneity between domains reduces risk, and vice versa. While our model shift results are initially established for Gaussian designs, we extend them to more general designs through a universality argument. To illustrate the broader applicability of our technical tools beyond interpolation learning, we characterize the risk of a bias-corrected estimator that uses the pooled interpolator as an initialization and corrects the resulting bias with target data. On the technical side, we develop a novel anisotropic local law and a Lindeberg-swapping argument, yielding tools that may be of independent interest in random matrix theory and universality analysis. Finally, we supplement our theory with simulations demonstrating the finite-sample efficacy of our results.
Jun 1, 2026cs.LG

Fast Generalization after Interpolation via Critically Damped Momentum Optimization

A central problem in machine learning is that models can achieve near-perfect training performance while generalizing substantially less well to unseen examples. This gap is especially acute in high-dimensional, low-sample regimes, where many interpolating solutions exist and optimization must implicitly select among minima with different generalization properties. Following recent theoretical advances on optimization dynamics near the interpolation threshold, we note that the two-regime structure of risk minimization, with loss minimization followed by complexity minimization, motivates a biphasic optimization schedule. We thus theoretically demonstrate that GROKtimizer, a biphasic strategy that combines rapid convergence to interpolation with Critically Damped Momentum (CDM)-based post-interpolation norm minimization, offers a natural solution for selecting low-norm interpolating solutions. Under a local quadratic model of the post-interpolation basin, GROKtimizer provides a quadratic speedup over classical gradient descent, with provable optimality among first-order optimizers. To showcase the applicability of our method, we evaluate GROKtimizer on several synthetic benchmarks common in the classical grokking literature and on various real-world datasets. Finally, we reconcile our findings with the flat-minima hypothesis, highlighting the importance of post-interpolation dynamics in the construction of high-quality, generalizing models.
Sep 27, 2026cs.LG

Benign Overfitting for General Norms and Distributions

Understanding why predictors can generalize despite interpolating noisy training data is a central puzzle in machine learning. Most work on such "benign overfitting" studies minimum-2-norm linear regression, reflecting the inductive bias of gradient descent. However, modern optimizers such as Adam and Muon use non-Euclidean update geometries, favoring solutions associated with other norms. Analyzing regression for non-Euclidean norms is substantially more difficult, with known results essentially limited to Gaussians. In this paper, we develop a method to analyze benign overfitting in linear regression for general norms and general (sub-Gaussian) distributions. As a special case, we prove that minimum-p-norm interpolation with p>1 can benignly overfit even for non-Gaussian distributions, under suitable conditions. Perhaps surprisingly, for the 1-norm, benign overfitting does not hold in general for well-behaved (but non-Gaussian) distributions, showing that existing positive 1-norm results rely crucially on Gaussianity. Our proof analyzes the geometry of the dual optimization problem, using concentration and central limit tools to show it is approximately Euclidean in many high-dimensional cases.