stat.MLSep 28, 2026

Two-Timescale Fine-tuning Provably Learns New Features for Two-Layer ReLU Networks

Authors: Etienne Boursier, Nicolas Flammarion

Organizations: INRIA, LMO, Universit´e Paris-Saclay, Orsay, France · EPFL, Lausanne, Switzerland

Abstract

Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has m+1m+1 features, mm of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only O(d)\mathcal{O}(d) fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A Theory of How Pretraining Shapes Inductive Bias in Fine-Tuning

    Feb 23, 2026Nicolas Anguita, Francesco Locatello, Andrew M. Saxe +4Model Fine-TuningFeature Learning

  2. Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks

    Sep 28, 2026Alexandre Declèves, Etienne Boursier, Nicolas FlammarionPretrainingModel Fine-Tuning

  3. The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge

    May 13, 2026Ryoya Awano, Taiji SuzukiFeature LearningStochastic Gradient Descent