cs.LGSep 29, 2026

Grokking through the Lens of Minimum-Norm Interpolation

Authors: Gil Kur, Ileana Rugina, Clémentine Carla Juliette Dominé, Marco Mondelli

Organizations: ETH Zürich · Institute of Science and Technology Austria · Harvard University

Abstract

Grokking shows that fitting the training data and learning the underlying signal can occur at very different stages. However, existing theories offer limited quantitative insight into how this delayed generalization depends on inductive bias and signal structure. Our work addresses the gap by developing a statistical theory that characterizes how regularization geometry and signal sparsity govern generalization near interpolation. In particular, we focus on the prototypical setting of high-dimensional regression and identify regimes in which sparsity-promoting regularization makes exact interpolation much more accurate than approximate fitting. In strongly overparameterized noiseless problems, we prove a zero--one generalization law and construct a family of convex norms whose interpolators transition from the trivial risk of the all-zero predictor to exact recovery, while keeping the training error equal to 00. Furthermore, when feature dimension and sample size are proportional, we provide a precise characterization of training and generalization errors along ℓr\ell_r-regularization paths. This in turn allows us to quantify the generalization gain that remains near interpolation: we show that this gain increases as the norm becomes more sparsity-promoting and as the target becomes sparser, with a sharp drop in generalization reached for noiseless data and ℓ1\ell_1 regularization. Experiments on diagonal linear networks and transformers trained on modular arithmetic demonstrate the generality of our theoretical predictions. Finally, beyond grokking, our work reveals a statistical instability in minimum-norm interpolation: small perturbations in the regularization strength can lead to drastically different generalization, while preserving small training error.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Generalization error of min-norm interpolators in transfer learning

    Jun 20, 2024Yanke Song, Kenneth Gu, Sohom Bhattacharya +1Stochastic InterpolantsEmpirical Risk Minimization

  2. Benign Overfitting for General Norms and Distributions

    Sep 27, 2026Daniel Barzilai, Ohad ShamirOverfittingSpectral Norm