From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes
Organizations: Department of Computer Sciences, University of Wisconsin–Madison, USA. · School of Mathematical Sciences, Institute of Natural Sciences and MOE-LSC, Shanghai Jiao Tong University, China.
Abstract
Power-law learning curves are often treated as fixed properties of a model and its data, although learning-rate and batch-size schedules can change the observed loss. We study this dependence in noisy online SGD with linear random features. Conditional on the representation, an exact Volterra equation separates two response components: a forcing term that propagates unresolved target error and a memory kernel that propagates stochastic-error injections. We prove that either component follows a power law if and only if its cumulative weighted spectral mass has the corresponding low-spectrum scaling; individual eigenvalues and target coefficients need not obey coordinatewise power laws. Under a joint schedule, intrinsic time controls optimization progress, while controls noise injection. Their interaction yields sharp conditions under which a schedule preserves, changes, or destroys the clean power law, together with a memory ceiling on noise reduction. The power-law random-feature model realizes this mechanism in propagation regimes with phase-dependent compute rates. Controlled nanoGPT experiments show that (1) learning-rate and batch-size schedules with matched paths are nearly equivalent in intrinsic time, (2) a forcing-memory surrogate accurately predicts loss across schedules, and (3) its fitted exponents across real-world datasets identify the regime of LLMs in map.
Figures & tables
| Regime and branch | ||
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Limit regime | Results in this paper |
| Exact finite width: fixed, | Exact finite-bulk kernel and noisy–clean gap ( Lemmas C.8 and C.9 ) |
| Fixed-width terminal: fixed, | Stable terminal-risk transfer and the noise-induced plateau ( Theorems A.9 , E.3 and E.5 ) |
| Fixed-time DE: , , fixed | Positive spectral representation and DE risk recursion for varying schedules ( Lemmas A.1 , C.1 and C.5 ) |
| Fixed infinite-spectrum system: | Schedule response, full classification, and gap-to-ratio identification ( Theorems 5.1 , A.6 , D.13 , D.16 , D.19 and D.23 ) |
| Joint width–time: , | Spectral iff criteria, uniform transfer, and LM/IM/FB risk laws ( Theorems 3.1 , A.10 , B.1 , D.4 , C.14 , C.15 , C.18 and C.21 ) |
| Resource limit: or ; vary | Optimal ratio control and phasewise data/compute rates ( Theorems C.22 , 5.2 and C.23 ) |
| Analytical level | Formal results |
| Exact conditional finite-width dynamics | Exact finite-bulk kernel and noisy–clean gap ( Lemmas C.8 and C.9 ). |
| Deterministic equivalent | Positive DE spectral representation, varying-schedule DE recursion, and conditional transfer to realized SGD ( Lemmas A.1 , C.1 and C.21 ). |
| Power-law asymptotics | Spectral iff criteria, uniform Tauberian and schedule transfer, and LM/IM/FB risk laws ( Theorems A.4 , B.1 , D.4 , C.14 and C.18 ). |
| Symbol | Meaning | Symbol | Meaning |
|---|---|---|---|
| Model and spectrum | |||
| input dimension | feature width | ||
| spectral-decay exponent | target-regularity exponent | ||
| target-energy tail exponent | covariate | ||
| population covariance | population eigenvalue | ||
| teacher coefficients | teacher predictor | ||