cs.AIJun 16, 2026

Small Initialization Matters for Large Language Models

Authors: Liangkai HangJunjie YaoZhiyu LiFeiyu XiongHongkang YangZhi-Qin John Xu

Organizations: School of Mathematical Sciences, Shanghai Jiao Tong University, Shanghai, 200240, China. · MemTensor (Shanghai) Technology Co., Ltd.. for Advanced Algorithms Research, Shanghai, China. · Institute of Natural Sciences, Shanghai Jiao Tong University, Shanghai, 200240, China.

Abstract

Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered. Although progress is usually attributed to scale, data and architecture, we show that parameter initialization is a gene-like determinant of training and, in particular, of model capacity. Reducing the initialization scale consistently improves pretraining, with the largest gains on reasoning-demanding tasks. We identify two widely used empirical settings that restrain the advantage of small initialization, and show how relaxing them restores favorable scaling. We further uncover a critical initialization that balances the reasoning and training. Mechanistically, small initialization drives a distinct developmental trajectory: parameters first condense into low-complexity structures and later expand into richer representations, giving concrete form to the idea that compression is intelligence. Token-level analyses show that the gains concentrate on non-trivial, context-constrained predictions rather than all tokens uniformly. These results motivate a simple γγ-initialization rule: expose initialization rage as an explicit knob and use small initialization by default, an almost cost-free intervention that improves pretraining and strengthens reasoning across model scales.

Explore similar work

CardsList
  1. Small LLMs: Pruning vs. Training from Scratch

    Jun 12, 2026Yufeng Xu, Taiming Lu, Kunjun Li +3Small Language ModelsPruning