cs.LGSep 30, 2026

How Does Local Landscape Geometry Evolve in Language Model Pre-Training?

Authors: Zhanpeng Zhou, Yuhan Sun, Bingrui Li, Jinbo Wang, Huaijin Wu, Lei Wu, Junchi Yan

Organizations: Shanghai Jiao Tong University · Tsinghua University · Peking University

Abstract

The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts from sharp to flatter regions early in training. This dynamic explains the necessity of LR warmup and further suggests that larger peak LRs require proportionally longer warmup periods. In Phase II, the local landscape is governed by the gradient noise scale. Our theory identifies a depth flatness trade-off: high noise from smaller batches widens the loss basin, whereas reduced noise from larger batches deepens it. This theory motivates a dynamic batch-size (BS) scheduler that begins with a small BS and increases it late in training. Together, we provide a unified view of loss landscape evolution, which translates into actionable tuning strategies for large-scale pre-training.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Towards joint scaling laws with optimal batch size schedules

    Jul 30, 2026Jiaxiang Li, Zhiqi Bu, Shiyun XuBatchLarge Language Model Training

  2. Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

    Aug 28, 2026Niccolò Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan +4BatchScaling Laws

  3. The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training

    May 26, 2026Hongtao Zhang, Wenjie Zhou, Chenxi Jia +2Large Language Model PretrainingPretraining