hep-exJun 18, 2026

Towards Engineering Scaling Laws with Pretraining Data Composition

Authors: Jan-Lucas UsluKevin GreifDaniel WhitesonBenjamin Nachman

Abstract

Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established for large language models, these relationships are emerging for large models in particle physics. As with language, empirical studies show that the performance scales as a power law. However, unlike natural language or image domains, fundamental physics has high-fidelity simulators that produce synthetic data cheaply. This favors scaling regimes where additional data is cheaper than additional parameters, and allows the pretraining dataset itself to be engineered to influence the scaling. For the task of classifying hadronic jets produced in collisions of high-energy particle beams, we show that the scaling behavior can be engineered towards requiring more data rather than larger models by inclusion of pretraining data which is more diverse and better aligned with the downstream classification task.

Explore similar work

CardsList
  1. Bridging Compute- and Data-Optimal Pretraining

    Jul 28, 2026Tian Qin, Kimia Hamidieh, David Alvarez-MelisScaling LawsBest-Epoch Metrics

  2. Neural Scaling Laws for Jet Generation

    May 27, 2026Oz Amram, Darius A. Faroughy, Tjarko Gerdes +5Scaling LawsHigh Energy Physics