Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoretically. We study fine-tuning from pretrained weights through the lens of sparse linear regression and two-layer diagonal linear networks. In our setting, pretraining provides information through the support (and signs) of the initialization predictor, which may contain coordinates relevant to the downstream task. We show how pretrained information reshapes the implicit bias and training dynamics, and can thereby reduce the sample complexity of recovering the target parameters and support. In particular, for a clean initialization with correctly inherited signs, we show that the required sample size is comparable to that of a weighted Lasso estimator that explicitly exploits the pretrained support through a suitably chosen regularizer. Our results thus show how information encoded in pretrained weights can be implicitly exploited by gradient-based fine-tuning, reducing the amount of data needed to recover a downstream task.
Figures & tables
Figure 1: Schematic dual trajectories under non-zero initialization. Missing and wrongly signed true coordinates travel dual distances 1 and 2 , respectively, before entering with the correct sign.
Early-stopped S2S
Weighted Lasso
Sample size
n≳s+f0+m2+mf0+mlog(d)
n≳mmisslog(d)+(s−mmiss)log(sinit)
Support guarantee
S⋆⊆Sfinal⊆S⋆∪F0
supp(βWL)=S⋆
Table 1: Noisy recovery guarantees under the respective assumptions: Theorem 3 with B=m+f0 for S2S, and Proposition 1 with σ/mini∈S⋆∣βi⋆∣=O(1) for weighted Lasso. Numerical constants and confidence dependence are suppressed.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Early stopping along S2S trajectories on two independent instances. Left: test loss for d=300 , ntrain=220 , ntest=1200 , s=40 , and σ=0.09 . The test loss reaches its oracle minimum at saddle k=33 and then increases sharply as the trajectory continues. Right: null-gradient stopping rule for d=32 , n=220 , s=6 , and σ=0.05 . The four missing true coordinates are proposed before any null coordinate; the first null proposal falls below the stopping threshold and is therefore rejected.
Figure 3: Empirical probability of exact support recovery as a function of the number m of missing true coordinates for a clean initialization ( d=260 , n=140 , s=20 , σ=0.005 ). Probabilities are estimated over 200 independent trials. S2S succeeds with high probability when few true coordinates remain to be learned, while its recovery probability decreases sharply as m increases. Weighted Lasso degrades more gradually.
Department of Engineering, University of Cambridge · Institute of Science and Technology, Austria (ISTA) · Gatsby Computational Neuroscience Unit and Sainsbury Wellcome Centre, UCL +1