cs.LGMay 7, 2026

Weight-Decay Turns Transformer Loss Landscapes Villani: Functional-Analytic Foundations for Optimization and Generalization

Authors: Abhijit DasSayantan Dutta

Organizations: GE HealthCare, Bangalore 560066, Karnataka, India

Abstract

Weight decay is widely used as a regularizer in large language models, yet its precise role in shaping Transformer loss landscapes remains theoretically underexplored. This paper provides the first rigorous functional-analytic characterization of the standard Transformer objective--cross-entropy loss with L2L^2 regularization--by proving it satisfies Villani's criteria for coercive energy functions. Specifically, we show that the regularized loss F\mathcal{F} is infinitely differentiable, grows at least quadratically, has Gaussian-integrable tails, and satisfies the differential growth condition ΔF+1sF2-Δ\mathcal{F} + \tfrac{1}{s}\|\nabla\mathcal{F}\|^{2} \to \infty as θ\|θ\| \to \infty for all s>0s>0. From this structure, we derive explicit log-Sobolev and Poincaré constants CLSλ1+d/λ2C_{\mathrm{LS}} \leq λ^{-1} + d/λ^{2}, linking the regularization strength λλ and model dimension dd to finite-time convergence guarantees for noisy stochastic gradient descent and PAC-Bayesian generalization bounds that tighten with increasing λλ. To validate our theory, we introduce a scalable Villani diagnostic Ψs(θ)=ΔF+s1F2Ψ_s(θ) = -Δ\mathcal{F} + s^{-1}\|\nabla \mathcal{F}\|^2 and estimate it efficiently using Hutchinson trace probes in models with over 100M parameters. Experiments on GPT-Neo-125M across Penn Treebank and WikiText-103 confirm the predicted quadratic growth of ΨsΨ_s, spectral inflation of the Hessian, and exponential convergence behavior consistent with our log-Sobolev analysis. These results demonstrate that weight decay not only improves generalization empirically but also establishes the mathematical conditions required for fast Langevin mixing and theoretically grounded curvature-aware optimization in deep learning.

Explore similar work

May 28, 2025cs.LG

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization

The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint. While weight decay is standard practice in modern training procedures, most theoretical analyses focus on unregularized settings. In this work, we investigate the loss landscape of the 2\ell_2-regularized training loss for two-layer ReLU networks. We show that the landscape becomes benign -- i.e., free of spurious local minima -- under large overparametrization, specifically when the network width mm satisfies mmin(nd,2n)m \gtrsim \min(n^d, 2^n), where nn is the number of data points and dd the input dimension. More precisely in this regime, almost all constant activation regions contain a global minimum and no spurious local minima. We further show that this level of overparametrization is not only sufficient but also necessary via the example of orthogonal data. Finally, we demonstrate that such loss landscape results primarily hold relevance in the large initialization regime. In contrast, for small initializations -- corresponding to the feature learning regime -- optimization can still converge to spurious local minima, despite the global benignity of the landscape.
Etienne Boursier, Matthew Bowditch, Matthias Englert +1
May 19, 2026cs.LG

Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics

Transformers trained on modular arithmetic exhibit sharp transitions between memorization, generalization, and collapse. We show that weight decay acts as a scalar empirical control parameter for these regimes, and introduce two cheap online diagnostics, mean pairwise attention-head cosine similarity and entropy standard deviation, that track training dynamics from attention activations alone and complement loss-landscape diagnostics at lower compute cost. Across eleven experimental conditions and three model scales (0.82M to 85M parameters), the weight-decay axis separates memorization, developmental grokking, and collapse. A near-transition logistic fit localizes the memorization-to-developmental boundary at λc=0.0158λ_c=0.0158 (95% CI [0.0109, 0.0200], N=210); a power-law fit gives an empirical exponent ν=0.757ν=0.757 (CI [0.725, 0.799]). Reference exponents ν=1/2ν=1/2 and 3D Ising ν0.63ν\approx 0.63 lie outside this empirical CI under our four-bin grid, so we report νν as empirical and defer universality-class identification to denser finite-size-scaling work. A horizon-matched multi-task replication (n=280, four modular operations) preserves the weight-decay control pattern; a paired attention-head re-initialization experiment at λ=0.05λ=0.05 changes Phase-2 amplitude (Cohen's d=1.190d=-1.190, n=10, pt=4.5×103p_t=4.5 \times 10^{-3}), while matched weight-norm clipping does not. Three cross-architecture probes (4L MLP, 4L LSTM, and 4L Mamba; each n=70) replicate the weight-decay-controlled transition with architecture-specific λcλ_c values. Main diagnostic claims are scoped to modular arithmetic in small transformer attention models; the non-attention experiments are scope probes, and architecture-wide, language-model, and universality-class claims are out of scope.
Lucky Verma
May 11, 2026cs.LG

Neural Weight Norm = Kolmogorov Complexity

Why does weight decay work? We prove that, in any fixed-precision regime, the smallest weight norm of a looped neural network outputting a binary string equals the Kolmogorov complexity of that string, up to a logarithmic factor. This implies that weight decay induces a prior matching Solomonoff's universal prior, the optimal prior over computable functions, up to a polynomial factor. The result is norm-agnostic: in fixed precision, every weight norm collapses to the non-zero parameter count up to constants, so the same sandwich bound holds for any norm used as a regulariser. The proof has two short reductions: any program for a universal Turing machine can be encoded into neural weights at unit cost per program bit, and any fixed-precision network can be described by enumerating its non-zero parameters with logarithmic addressing overhead. Both bounds are tight up to constants, with the logarithmic factor realised by permutation encodings: a network whose parameters encode a permutation produces a string whose Kolmogorov complexity is the non-zero parameter count times its logarithm. The fixed-precision assumption is essential: with infinite precision, neural networks can encode non-computable functions and the weight norm loses its relevance.
Tiberiu Musat