Scalable AI Uncertainty Quantification via Generalized Laplace Active Subspaces
Authors: Wouter N. Edeling, Peter V. Coveney
Organizations: Scientific Computing Group, CWI, Science park 123, Amsterdam, 1098XG, the Netherlands. · Faculty of Electrical Engineering, Mathematics and Computer Science, University of Twente, Hallenweg 15, Enschede, 7500AE, the Netherlands. · Centre for Computational Science, University College London, 20 Gordon Street, London, WC1H 0AJ, UK. · Advanced Research Computing Centre, University College London, Gordon Street, London, WC1E 6BT, UK.
Reliable uncertainty quantification (UQ) is essential for deploying neural networks in scientific and high-stakes applications, but full Bayesian inference over the network parameters is computationally infeasible. We propose a low-rank generalized Laplace approximation for neural-network UQ based on a small number of data-informed curvature directions. Starting from a generalized Bayesian posterior defined through an empirical loss, we construct a local Gaussian approximation around a pretrained set of weights in this active curvature subspace. The posterior variances in the retained subspace are available in closed form, and the prior variance is calibrated by an empirical Bayes procedure. The generalized Bayesian formulation allows us to compare two posterior scalings: the standard Bayesian scaling associated with the summed negative log likelihood, and a mean-loss scaling in which the empirical loss is normalized by the number of data. A central finding is that the standard scaling induces a data-size dependent contraction of the posterior variance in the leading active directions. In regression problems, this can force the low-rank framework to retain additional weak-curvature directions in order to achieve nominal coverage of calibration data. When posterior samples are propagated through the non-linear network, these additional directions can degrade the coherence of the predictive intervals and shift the posterior predictive mean away from the pretrained model. In contrast, the generalized mean-loss scaling yields a more stable, lower dimensional active subspace and produces calibrated, coherent predictive confidence intervals. These results indicate that generalized Laplace active subspaces provide a practical and scalable route to calibrated uncertainty quantification in neural networks.
Figures & tables
Figure 1: The matrix-vector product Gv∈RD with D=2701 , computed using the matrix-free method outlined in Section 4.1.5 , and by explicitly forming matrix G via ( 28 ). The absolute error ∣Gv−(Gv)explicit∣ is plotted on the right vertical axis.
Figure 2: The d leading eigenvalues λj and leading eigenvector p1 of G , computed using the matrix-free method outlined in Section 4.1.5 , and by explicitly forming matrix G via ( 28 ). The absolute errors are plotted on the right vertical axis.
Figure 3: 90% confidence intervals for the, a) generalized Bayesian scaling, and b) standard Bayesian scaling. Both CIs were generated by drawing 1000 posterior samples, and propagating these through the network.
Figure 4: The results for the standard Bayesian scaling β=Nσ−2 with N=1000 and d=43 . The CIs were generated by drawing 1000 posterior samples, and propagating these through the network.
Figure 5: The results for the standard Bayesian scaling β=Nσ−2 combined with the linearized Laplace predictive posterior distribution, computed using f(x;θ0)±2σf .
Figure 6: The coverage of the calibration procedure as outlined in Section 4.1.7 , versus the active subspace dimension d , for both the standard Bayesian scaling β=Nσ−2 and the generalized scaling β=σ−2 .
Figure 7: The parity plot of the mean compressive strength of concrete, with 90% CIs. We show the standard and generalized scaling results, both computed with sampled Laplace.
Figure 8: The cumulative (normalized) activity scores for our concrete compression strength feed-forward neural network with approximately D=106 weights.
Figure 9: The inner products ⟨p1,p1ref⟩ as a function of Nsub∈{500,1000,1500,2000,2500,3000} , replicated 25 times per Nsub . The reference vectors p1ref were computed using Nsub=10.000 SMILES strings.
baseline
Laplace (std. scaling)
Laplace (gen. scaling)
d=1
8692
8709
6863
d=2
8692
8753
5279
d=3
8692
8732
4435
Table 1: Number of valid SMILES strings (out of 10k).
Figure 10: Kernel density estimates of the sequence length for the standard and generalized scaling with d∈{1,2,3} . The maximum allowed sequence length is 256.
Quantity
Description
Molecular weight
Molecular size based on atomic masses, large values indicate large molecules.
SA score
Synthetic Accessibility score, higher values indicate more difficult synthesis.
SlogP
Estimate of how fat-soluble a molecule is.
Table 2: Molecular quantities of interest used to compare generated SMILES distributions within REINVENT.
Figure 11: Kernel density estimates of the QoIs for the standard and generalized scaling with d∈{1,2,3} .
Figure 12: The cumulative (normalized) activity scores for the REINVENT LSTM with approximately 1.6M weights. Computed using d=2 .
Figure 13: The contraction ratio ( 10 ) and posterior variances ( 22 ) for the generalized (left column) and standard (right column) Bayesian scaling, for both N=50 and N=1000 . These results were computed using the analytic regression function ( 7 ).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Index
Token
Description
Index
Token
Description
0
$
Stop/end token
17
Br
Bromine atom
1
^
Start token
18
C
Aliphatic carbon atom
2
#
Triple bond
19
Cl
Chlorine atom
3
%10
Two-digit ring closure
20
F
Fluorine atom
4
(
Branch opening
21
N
Aliphatic nitrogen atom
5
)
Branch closing
22
O
Oxygen atom
Appendix
Table 3: Vocabulary used by the autoregressive SMILES generator. The vocabulary contains start and stop tokens, bond symbols, branch symbols, ring-closure indices, atom tokens, and charged or aromatic atom tokens.
Existing methods for quantifying predictive uncertainty in neural networks are either computationally intractable for large language models or require access to training data that is typically unavailable. We derive a lightweight alternative through two approximations: a first-order Taylor expansion that expresses uncertainty in terms of the gradient of the prediction and the parameter covariance, and an isotropy assumption on the parameter covariance. Together, these yield epistemic uncertainty as the squared gradient norm and aleatoric uncertainty as the Bernoulli variance of the point prediction, from a single forward-backward pass through an unmodified pretrained model. We justify the isotropy assumption by showing that covariance estimates built from non-training data introduce structured distortions that isotropic covariance avoids, and that theoretical results on the spectral properties of large networks support the approximation at scale. Validation against reference Markov Chain Monte Carlo estimates on synthetic problems shows strong correspondence that improves with model size. We then use the estimates to investigate when each uncertainty type carries useful signal for predicting answer correctness in question answering with large language models, revealing a benchmark-dependent divergence: the combined estimate achieves the highest mean AUROC on TruthfulQA, where questions involve genuine conflict between plausible answers, but falls to near chance on TriviaQA's factual recall, suggesting that parameter-level uncertainty captures a fundamentally different signal than self-assessment methods.
Nils Grünefeld, Jes Frellsen, Christian Hardmeier
1IT University of Copenhagen · 3Pioneer Centre for Artificial Intelligence · 2Technical University of Denmark
Although the Laplace approximation offers a simple route to uncertainty quantification in deep neural networks, its reliance on inverting large Hessian matrices has motivated a range of computationally feasible low-dimensional or sparse approximations. A prominent class of such methods - sub-network Laplace approximations, constructs surrogates by restricting attention to a small subset of parameters. Existing approaches in this family typically rely on diagonal, layer-wise, or other architectural heuristics for subset selection, which ignore cross-parameter interactions and lack formal optimality guarantees. In this paper, we provide a rigorous theoretical analysis of the sub-network Laplace paradigm. We prove that all sub-network Laplace methods systematically underestimate the predictive variance of the full Laplace posterior, and that this bias decreases monotonically as the retained sub-matrix expands. Leveraging this insight, we propose two principled, analytically grounded sub-network Hessian approximations: \textit{Gradient-Laplace} selects parameters with the largest average squared gradients of the model output with respect to the parameters over a reference dataset; while \textit{Greedy-Laplace} iteratively refines this selection by accounting for off-diagonal interactions in the precision matrix. We establish theoretical guarantees characterizing their optimality properties and show that Gradient-Laplace provably outperforms existing heuristic approaches. Extensive numerical studies across diverse settings indicate that these methods perform strongly relative to existing benchmarks.
Swarnali Raha, Kshitij Khare, Rohit K Patra
Department of Statistics University of Florida Gainesville, FL · Linkedin Inc. New York, NY
Epistemic uncertainty quantification (UQ) for deep neural networks (DNNs) is a requirement for safe adoption of AI in mission-critical settings. Several leading methods for UQ linearize DNNs to form Bayesian Generalized Linear Models (GLMs), where epistemic uncertainty is modeled via the predictive posterior distribution. Linearizing around the parameters of the final connected layer of a DNN is a commonly used approximation for reducing the computational burden of such GLMs, though it is often believed to come at the cost of degraded performance. In this work, we compare GLMs arising from full-network and last-layer linearization using both theoretical and empirical approaches. We first employ tools from random matrix theory to conduct a theoretical comparison; this analysis reveals no meaningful improvement in the UQ capabilities of full linearization. Coupled with a large-scale empirical evaluation across a range of modern machine learning tasks, we arrive at the following conclusion: a last-layer approximation yields comparable UQ performance while offering substantially improved computational efficiency.
Joseph Wilson, Chris van der Heide, Liam Hodgkinson +1
School of Mathematics and Physics, University of Queensland, Australia · ARC Training Centre for Information Resilience (CIRES), Brisbane, Australia · Department of Electrical and Electronic Engineering, University of Melbourne, Australia +1