cs.LGJul 1, 2026

How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size

Authors: Fabian Schaipp

Organizations: Inria, ´Ecole Normale Sup´erieure, PSL Research University, Paris

Abstract

We propose a scaling law that takes into account model size and training data while explicitly splitting the latter into training steps and batch size (called three-term law). Fitting the proposed law on a large set of training runs, we find that it correctly recovers the scaling of the optimal batch size. Moreover, because it makes use of training runs with suboptimal batch size, our proposed law can be robustly fit with a significantly smaller amount of training runs. We further show that the three-term law can be used to derive scaling laws for suboptimal batch sizes, and that it matches previous empirical findings related to the critical batch size.

Explore similar work

CardsList
  1. Prescriptive Scaling Laws for Data Constrained Training

    May 2, 2026Justin Lovelace, Christian Belardi, Srivatsa Kundurthy +2Scaling LawsCapacity

  2. Smooth Scaling Laws Hide Stepwise Token Learning

    Jun 29, 2026Pingjie Wang, Zechen Hu, Peiru Yang +2Scaling LawsLarge Models