cs.LGAug 26, 2026

Mapping the Emergence of Regularization-Driven Dynamics in Grokking

Authors: Yiming Lin, Yuxuan Wang

Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences

Abstract

For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight decay (WD) perturbations across the pre-generalization plateau and measure how they shift later generalization time. Across three grokking tasks, these shifts are unordered early in the plateau but later form a stable dose ordering before visible generalization, with stronger WD increases leading to earlier generalization and stronger WD decreases leading to later generalization. Test-loss barriers between perturbed and baseline generalization checkpoints collapse toward zero while the ordered timing effects persist. A similar response reorganization is observed under ℓ1\ell_1 regularization in the grokking setting of Junior et al. (2025). Drawing on Waddington's developmental landscape as an analogy, we call this combination of increasingly constrained solution selection and persistent dose-ordered timing shifts the canalization of grokking solution selection. Together, our response maps and loss-barrier measurements reveal a dynamical reorganization before visible generalization that is consistent with the theoretical picture of regularization-driven motion along a stable slow manifold (Boursier et al., 2025).

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries

    Jul 27, 2026Taeyoung KimGrokkingWeight Decay

  2. A Stochastic--Geometric Theory of Scaling Laws in Grokking

    Jun 29, 2026Róisín Luo, Christian Gagné, Jonas Ngnawé +2GrokkingScaling Laws

  3. A Theoretical Analysis of Generalization Dynamics in Neural Networks under Gradient Descent with Weight Decay

    Sep 7, 2026Yuqing Wang, Ioannis G. Kevrekidis, Mikhail BelkinTwo-Layer Neural NetworksGeneralization Bounds