cs.LGSep 30, 2026

From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes

Authors: Yichen Wang, Fanghui Liu, Yudong Chen

Organizations: Department of Computer Sciences, University of Wisconsin–Madison, USA. · School of Mathematical Sciences, Institute of Natural Sciences and MOE-LSC, Shanghai Jiao Tong University, China.

Abstract

Power-law learning curves are often treated as fixed properties of a model and its data, although learning-rate and batch-size schedules can change the observed loss. We study this dependence in noisy online SGD with linear random features. Conditional on the representation, an exact Volterra equation separates two response components: a forcing term that propagates unresolved target error and a memory kernel that propagates stochastic-error injections. We prove that either component follows a power law if and only if its cumulative weighted spectral mass has the corresponding low-spectrum scaling; individual eigenvalues and target coefficients need not obey coordinatewise power laws. Under a joint schedule, intrinsic time Tt=∑s<tηsT_t=\sum_{s<t}η_s controls optimization progress, while rt=Bt/ηtr_t=B_t/η_t controls noise injection. Their interaction yields sharp conditions under which a schedule preserves, changes, or destroys the clean power law, together with a memory ceiling on noise reduction. The power-law random-feature model realizes this mechanism in 3+3(+2)3+3(+2) propagation regimes with phase-dependent compute rates. Controlled nanoGPT experiments show that (1) learning-rate and batch-size schedules with matched B/ηB/η paths are nearly equivalent in intrinsic time, (2) a forcing-memory surrogate accurately predicts loss across schedules, and (3) its fitted exponents across real-world datasets identify the regime of LLMs in 3+3(+2)3+3(+2) map.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Optimal Learning Rate Schedules under Functional Scaling Laws: Power Decay and Warmup-Stable-Decay

    Feb 6, 2026Binghui Li, Zilin Wang, Fengling Chen +3BatchScaling Laws

  2. Towards joint scaling laws with optimal batch size schedules

    Jul 30, 2026Jiaxiang Li, Zhiqi Bu, Shiyun XuBatchLarge Language Model Training

  3. On the Nonlinearity of Learning Rate Scaling for LLM Training

    Jun 28, 2026Zaiwen Yang, Huaqing Zhang, Jing Xu +1BatchLarge Language Model Training