Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoretically. We study fine-tuning from pretrained weights through the lens of sparse linear regression and two-layer diagonal linear networks. In our setting, pretraining provides information through the support (and signs) of the initialization predictor, which may contain coordinates relevant to the downstream task. We show how pretrained information reshapes the implicit bias and training dynamics, and can thereby reduce the sample complexity of recovering the target parameters and support. In particular, for a clean initialization with correctly inherited signs, we show that the required sample size is comparable to that of a weighted Lasso estimator that explicitly exploits the pretrained support through a suitably chosen regularizer. Our results thus show how information encoded in pretrained weights can be implicitly exploited by gradient-based fine-tuning, reducing the amount of data needed to recover a downstream task.
Figures & tables
Figure 1: Schematic dual trajectories under non-zero initialization. Missing and wrongly signed true coordinates travel dual distances 1 and 2 , respectively, before entering with the correct sign.
Early-stopped S2S
Weighted Lasso
Sample size
n≳s+f0+m2+mf0+mlog(d)
n≳mmisslog(d)+(s−mmiss)log(sinit)
Support guarantee
S⋆⊆Sfinal⊆S⋆∪F0
supp(βWL)=S⋆
Table 1: Noisy recovery guarantees under the respective assumptions: Theorem 3 with B=m+f0 for S2S, and Proposition 1 with σ/mini∈S⋆∣βi⋆∣=O(1) for weighted Lasso. Numerical constants and confidence dependence are suppressed.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Early stopping along S2S trajectories on two independent instances. Left: test loss for d=300 , ntrain=220 , ntest=1200 , s=40 , and σ=0.09 . The test loss reaches its oracle minimum at saddle k=33 and then increases sharply as the trajectory continues. Right: null-gradient stopping rule for d=32 , n=220 , s=6 , and σ=0.05 . The four missing true coordinates are proposed before any null coordinate; the first null proposal falls below the stopping threshold and is therefore rejected.
Figure 3: Empirical probability of exact support recovery as a function of the number m of missing true coordinates for a clean initialization ( d=260 , n=140 , s=20 , σ=0.005 ). Probabilities are estimated over 200 independent trials. S2S succeeds with high probability when few true coordinates remain to be learned, while its recovery probability decreases sharply as m increases. Weighted Lasso degrades more gradually.
Pretraining and fine-tuning are central stages in modern machine learning systems. In practice, feature learning plays an important role across both stages: deep neural networks learn a broad range of useful features during pretraining and further refine those features during fine-tuning. However, an end-to-end theoretical understanding of how choices of initialization impact the ability to reuse and refine features during fine-tuning has remained elusive. Here we develop an analytical theory of the pretraining fine-tuning pipeline in diagonal linear networks, deriving exact expressions for the generalization error as a function of initialization parameters and task statistics. We find that different initialization choices place the network into four distinct fine-tuning regimes that are distinguished by their ability to support feature learning and reuse and therefore by the task statistics for which they are beneficial. In particular, a smaller initialization scale in earlier layers enables the network to both reuse and refine its features, leading to superior generalization on fine-tuning tasks that rely on a subset of pretraining features. We demonstrate empirically that the same initialization parameters impact generalization in ResNets trained on CIFAR-100 and SVHN as well as Transformers trained on modular arithmetic tasks. Overall, our results demonstrate an alytically how data and network initialization interact to shape fine-tuning generalization, highlighting an important role for the relative scale of initialization across different layers in enabling continued feature learning during fine-tuning.
Nicolas Anguita, Francesco Locatello, Andrew M. Saxe +4
Department of Engineering, University of Cambridge · Institute of Science and Technology, Austria (ISTA) · Gatsby Computational Neuroscience Unit and Sainsbury Wellcome Centre, UCL +1
Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has m+1 features, m of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only O(d) fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.
Etienne Boursier, Nicolas Flammarion
INRIA, LMO, Universit´e Paris-Saclay, Orsay, France · EPFL, Lausanne, Switzerland
Finetuning pretrained models occurs in a low-dimensional subspace of the full parameter space. Prior work has focused on characterizing this optimization subspace, but largely ignored the complementary question: why do certain directions remain unexplored during finetuning? Are these stable directions irrelevant to downstream tasks, or do they already encode task-relevant structure that requires no further adjustment? Answering this question is central to understanding how pretrained knowledge transfers. Through systematic spectral analysis across vision and language models, we show that the leading singular vectors of pretrained weight matrices remain highly stable under finetuning and are shared across unrelated downstream tasks, revealing that pretraining establishes a reusable spectral coordinate system. Models pretrained on larger datasets exhibit greater spectral stability under distribution shift or task change, directly linking pretraining scale to geometric transferability. Motivated by these findings, we propose a parameter-efficient method that freezes pretrained singular vectors and optimizes only leading spectral coefficients, achieving competitive performance on GLUE with 0.2% trainable parameters. Our results reveal that the stable directions encode transferable structure rather than irrelevant noise: successful pretraining discovers spectral bases that downstream tasks inherit and operate within.
Junjie Yu, Yue Wang, Zihan Deng +3
Department of Biomedical Engineering, Southern University of Science and Technology · Department of Psychology, The University of Hong Kong