cs.LGSep 28, 2026

Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons

Authors: Mana Sakai, Masaaki Imaizumi

Organizations: The University of Tokyo · RIKEN Center for Advanced Intelligence Project · Kyoto University

Abstract

Hyperparameter transfer across model width can substantially reduce the cost of tuning large neural networks, but its behavior when the training horizon grows with width is not fully understood. Building on the framework of fast hyperparameter transfer (Ghosh et al., 2026), which formalizes when transfer is effective, we investigate conditions that ensure fast transfer in the growing-horizon regime. Specifically, we study learning-rate transfer in a shallow linear network with a single trainable hidden matrix, trained by full-batch gradient descent. Under additional spectral assumptions, our main results are threefold. (i) We prove fast learning-rate transfer as n,T→∞n,T\to\infty whenever T=o(n)T=o(\sqrt{n}). (ii) We characterize the transfer rates through the finite-width perturbation scale, the first-order sensitivities of the loss and its learning-rate derivative to finite-width perturbations, and the local loss curvature. (iii) We derive limiting distributions for the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix. These results clarify how spectral structure and local loss sensitivities govern learning-rate transfer at growing horizons.

Explore similar work

CardsList
  1. Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks

    Jul 8, 2026Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara +1BatchSequential Scaling

  2. Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks

    May 29, 2026Tianyu Pang, Vignesh Kothapalli, Shenyang Deng +3BatchTwo-Layer Neural Networks

  3. Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate

    May 20, 2026Dayal Singh Kalra, Maissam BarkeshliHyperparameterLarge Language Model Training