stat.MLOct 8, 2026

σσTransfer: Uncertainty Transfer from Small to Large Networks under μPμ\mathrm{P}

Authors: Richard Bergna, Fernando Ruiz Mazo, Nicolò Felicioni, José Miguel Hernández-Lobato, Kamil Ciosek

Organizations: University of Cambridge · Spotify

Abstract

Reliable predictive uncertainty in Laplace approximations depends critically on the prior precision, yet selecting it requires a posterior sweep that is prohibitively expensive for neural networks with billions of parameters. Under the Maximal Update Parametrization (μPμ\mathrm{P}), we derive a rescaling of the prior covariance that makes the selected precision stable as model width grows. This leads to σTransferσ\mathrm{Transfer}: we select the precision on a smaller model and zero-shot transfer it to the much larger model, i.e., without searching for the precision on the larger model at all. We show convergence of the prior kernel, posterior covariance, selected precision, and posterior-derived decisions under explicit conditions, and verify σTransferσ\mathrm{Transfer} across regression, image classification, and Transformer readouts. For example, measured precision-sweep speedups reach ∼5000×\sim 5000\times when transferring from width 128 to 4096 on MNIST, at a target-NLL degradation of 0.0020.002; transferring from a public 1B to 7B model gives a median search speedup of ∼2.3×\sim 2.3\times (up to ∼330×\sim 330\times), with a mean measured target-NLL increase below 10−410^{-4} across ten tasks. The same posterior stability also enables transfer of acquisition, OOD-detection, and abstention decisions without constructing a target posterior.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scalable AI Uncertainty Quantification via Generalized Laplace Active Subspaces

    Oct 8, 2026Wouter N. Edeling, Peter V. CoveneyUncertainty QuantificationUncertainty Calibration

  2. Optimality of Sub-network Laplace Approximations: New Results and Methods

    May 9, 2026Swarnali Raha, Kshitij Khare, Rohit K PatraBayesian Neural NetworksUncertainty Quantification

  3. μμpscaling small models: Principled warm starts and hyperparameter transfer

    Feb 11, 2026Yuxin Ma, Nan Chen, Mateo Díaz +3Neural Network OptimizationHyperparameter Transfer