The Impact of Likelihood Tempering on the Limiting Predictive Moments of Variational Bayesian Linear Neural Networks
Organizations: Department of Statistical Sciences, University of Toronto · Dunlap Institute for Astronomy and Astrophysics
Abstract
In wide Bayesian neural networks, Gaussian mean-field variational inference is prone to "prior dominance": the Kullback-Leibler (KL) regularization term of the ELBO outweighs the expected log-likelihood, and the variational predictive distribution collapses to the prior predictive as the width grows. Tempering the likelihood, by raising it to the power for a temperature , is equivalent to scaling the KL term by . We ask in this paper how fast must decrease with to counteract this degeneracy and strike a good balance between the two terms. For single-hidden-layer linear networks with isotropic Gaussian priors, we derive the limiting predictive distribution under schedules of the form , with constants , as and compare it with the untempered neural network Gaussian process (NNGP) posterior, the infinite-width limit of the exact posterior. Our main result is that the predictive expectation and variance undergo phase transitions at different scales: the limiting expectation leaves its prior value at , once falls below an explicit threshold, and equals the least-squares prediction for , whereas the limiting variance keeps its prior value for , matches the NNGP's for , and vanishes for . With suitable choices of , one can recover either the NNGP posterior expectation or its variance.
Figures & tables
| Condition | Limit | Behaviour |
|---|---|---|
| Prior | ||
| , | Prior | |
| , | NNGP mean iff | |
| LS mean, prior variance | ||
| NNGP variance iff | ||
| Point mass |