stat.MLMay 22, 2026

Asymmetric Scaling Laws from Sparse Features

Authors: John SousMichael Winer

Organizations: Department of Applied Physics, Yale University, New Haven, Connecticut 06511, USA · Energy Sciences Institute, Yale University, West Haven, Connecticut 06516, USA · Institute for Advanced Study, Princeton, NJ 08540, USA · Alignment Research Center

Abstract

We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never observed in the training input. This mechanism induces a novel bottleneck absent from dense models. We derive the asymptotic population loss in both the underparameterized and overparameterized regimes, and show that the loss exhibits a double-descent peak near the interpolation threshold -- where the number of parameters is just sufficient to fit the training data -- resulting in a loss curve governed by two distinct scaling exponents -- one for the overparameterized regime and one for the underparameterized regime -- with a gap determined by the degree of sparsity. Additionally, we derive a compute-optimal frontier that favors increasing dataset size over model capacity under fixed compute budgets. We also analyze gradient-descent dynamics and identify a scaling law for the probability that fixed-step gradient descent becomes unstable. We further show that the sparsity-induced effect persists under nonlinear activations.

Explore similar work

CardsList
  1. Skaling: Chinchilla's Exponents Meet Kaplan's Coupling

    Aug 7, 2026Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz +1Scaling LawsExponents

  2. Prescriptive Scaling Laws for Data Constrained Training

    May 2, 2026Justin Lovelace, Christian Belardi, Srivatsa Kundurthy +2Scaling LawsCapacity