How Does Local Landscape Geometry Evolve in Language Model Pre-Training?
Organizations: Shanghai Jiao Tong University · Tsinghua University · Peking University
Abstract
The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts from sharp to flatter regions early in training. This dynamic explains the necessity of LR warmup and further suggests that larger peak LRs require proportionally longer warmup periods. In Phase II, the local landscape is governed by the gradient noise scale. Our theory identifies a depth flatness trade-off: high noise from smaller batches widens the loss basin, whereas reduced noise from larger batches deepens it. This theory motivates a dynamic batch-size (BS) scheduler that begins with a small BS and increases it late in training. Together, we provide a unified view of loss landscape evolution, which translates into actionable tuning strategies for large-scale pre-training.
Figures & tables
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Terminology | General Meaning | Usage in This Paper |
| Sharpness | A measure of curvature in the loss landscape, often characterized via the Hessian. Different works may define it differently. | We define sharpness as the curvature along the sharpest direction of the loss landscape. Mathematically, it is presented as the largest eigenvalue of the Hessian or of the preconditioned curvature matrix . |
| Flat/sharp minimum | A minimum is a point where the gradient vanishes and the loss does not decrease in a small neighborhood. A sharp minimum has large curvature; a flat minimum has small curvature. | We use these terms sparingly and follow the standard definitions from the sharpness/flat-minima literature. |
| Wide/deep basin | A loss basin is a region of the landscape surrounding a minimum. A wide basin rises loss slowly in most directions, whereas a deep basin has a significantly lower minimum value compared to its surroundings. | We use these terms to establish the depth–flatness trade-off: large noise scales tend to find wide basins, while small noise scales tend to find deeper regions with lower loss. |
| Acronym | Size | n head | depth | ||
| GPT-2 (small) | 124M | 768 | 3072 | 12 | 12 |
| LLaMA (93M) | 93M | 512 | 2048 | 16 | 8 |
| LLaMA (170M) | 170M | 768 | 3072 | 12 | 8 |
| LLaMA (270M) | 270M | 1024 | 4096 | 16 | 8 |
| LLaMA (530M) | 530M | 1536 | 6144 | 24 | 8 |