cs.LGMay 28, 2026

Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime

Authors: Paolo BaglioniChristian KeupVincenzo ZimbardoRosalba PacelliAlessandro VezzaniRaffaella BurioniPietro Rotondo

Organizations: 1INFN, Sezione di Milano Bicocca, Piazza della Scienza 3, 20126, Milano, Italy · 2INFN, Gruppo Collegato di Parma, Parco Area delle Scienze 7/A, 43124 Parma, Italy · 3Dipartimento di Scienze Matematiche, Fisiche e Informatiche, Universit`a degli Studi di Parma, Parco Area delle Scienze, 7/A 43124 Parma, Italy · 4INFN, sezione di Padova, Via Marzolo 8, 35131 Padova, Italy · 5Istituto dei Materiali per l’Elettronica ed il Magnetismo (IMEM-CNR), Parco Area delle Scienze, 37/A-43124 Parma, Italy

Abstract

The scaling limit where both the size of the training set PP and the width NN of a deep neural network grow at the same rate, the so-called proportional-width regime, has been intensely studied for shallow, single-hidden-layer networks. However, extending these non-perturbative results from shallow architectures to deep non-linear networks has proven very challenging. Here we present an effective approximate approach to predict the generalization performance of Bayesian multi-layer perceptrons (MLPs) of fixed depth LL on arbitrary high-dimensional data. We propose an equivalent Wishart Ansatz to capture the dominant stochastic fluctuations of the hierarchical empirical kernels of MLPs. This allows us to perform a large deviation analysis for the partition function of MLPs in the proportional limit, expressed in terms of a renormalized NNGP kernel. In this description, even strong representation learning in the proportional limit is encoded in at most LL scalar order parameters, determined self-consistently. Extending the approach to convolutional architectures (CNNs), we identify a hierarchical local kernel renormalization mechanism, which allows to quantify more complex data-dependent transformations of the large-width kernel in CNNs due to finite-width effects. We test our effective theory against sampling experiments from the Bayesian posterior of finite deep neural networks with depths LO(10)L \sim O(10) and PO(103)P\sim O(10^3) on classic benchmark datasets, finding overall very good agreement together with two distinct types of systematic deviations.

Explore similar work

CardsList