As robustness verification methods are becoming more precise, training certifiably robust neural networks is becoming ever more relevant. To this end, certified training methods compute and then optimize an upper bound on the worst-case loss over a robustness specification. Curiously, training methods based on the imprecise interval bound propagation (IBP) consistently outperform those leveraging more precise bounding methods. Still, we lack an understanding of the mechanisms making IBP so successful. In this work, we thoroughly investigate these mechanisms by leveraging a novel metric measuring the tightness of IBP bounds. We first show theoretically that, for deep linear models, tightness decreases with width and depth at initialization, but improves with IBP training, given sufficient network width. We, then, derive sufficient and necessary conditions on weight matrices for IBP bounds to become exact and demonstrate that these impose strong regularization, explaining the empirically observed trade-off between robustness and accuracy in certified training. Our extensive experimental evaluation validates our theoretical predictions for ReLU networks, including that wider networks improve performance, yielding state-of-the-art results. Interestingly, we observe that while all IBP-based training methods lead to high tightness, this is neither sufficient nor necessary to achieve high certifiable robustness. This hints at the existence of new training methods that do not induce the strong regularization required for tight IBP bounds, leading to improved robustness and standard accuracy.
Figures & tables
Figure 1: Comparison of exact ( ), optimal box ( ), and IBP ( ) propagation through a one layer network. We show the concrete points maximizing the logit difference y2−y1 as a black × and the corresponding relaxation as a red × .
Figure 2: Mean relative error between local tightness ( Definition 3.6 ) and true tightness computed with MILP for a CNN3 trained with PGD or IBP at ϵ=0.05 on MNIST .
Figure 3
Figure 5: Effect of network depth (right) and width (left) on tightness and training set IBP -certified accuracy.
Figure 6: Tightness, standard, and certified accuracy for CNN3 on CIFAR-10, depending on training method and perturbation magnitude ϵ used for training and evaluation.
Figure 6
Figure 9: Effect of a 4 -fold width increase on certified and standard accuracy for MNIST at ϵ=0.3 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: g(n) and g2(n) visualized.
Figure 11: Monte-Carlo estimations of Theorem 3.10 . Result bases on 10000 samples for each d . Left: c plotted against d in log scale. Right: E(δi∗) plotted against k for d=2000 (blue), together with the theoretical predictions (orange).
Figure 12: Accuracies and tightness of a CNN7 for CIFAR-10 ϵ=2552 depending on regularization strength with STAPS .
Figure 13: Tightness over propagation region size ξ for SABR and MNIST .
Figure 14: Mean relative error between local tightness ( Definition 3.6 ) and true tightness computed with MILP for a CNN3 trained with PGD or IBP at ϵ=0.005 on CIFAR-10.
Figure 15: Tightness and inverse IBP loss at initialization depending on width and depth.
Figure 16: Tightness (right) and inverse IBP loss (left and center) after IBP training depending on training perturbation size ϵ , evaluated at training ϵ (left) or constant ϵ=10−3 (center).
Dataset
ϵ
Method
Width
Accuracy
Certified
Tightness
MNIST
0.1
IBP
1×
85.70
67.71
0.871
2×
88.42
73.77
0.857
4×
90.31
79.89
0.803
Appendix
Table 3: Certified and standard accuracy depending on network width.
Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations. Neural network verification provides mathematically rigorous robustness guarantees, yet at substantial computational cost. To mitigate this, certified training techniques optimise for verifiable robustness during training, typically inducing a trade-off between natural and certified accuracy controlled by method-specific hyperparameters. Because these metrics are inherently conflicting, the common practice of reporting a single configuration is problematic: it can mislead conclusions about overall performance and prevents unbiased assessments of the state of the art. We address this by evaluating certified training methods via Pareto front comparisons over the natural--certified accuracy trade-off. To enable fair, method-agnostic comparisons, we perform efficient automated multi-objective hyperparameter optimisation to identify a set of Pareto-optimal configurations for each method. This approach often uncovers substantial undertuning in previously reported configurations, yielding superior performance and establishing a new state of the art. Leveraging these fronts, we present the first comprehensive multi-objective comparison of certified training approaches, showing that prior advancements are less pronounced than assumed and revealing previously unreported performance complementarities.
Konstantin Kaulen, Hadar Shavit, Holger H. Hoos
Chair of AI Methodology, RWTH Aachen University, Aachen, Germany · LIACS, Leiden University, Leiden, The Netherlands
Certified training aims to produce models whose predictions can be formally verified against adversarial perturbations, typically by optimising upper bounds on the worst-case loss over an allowed perturbation set. For neural networks, certified training methods based purely on tight relaxation bounds produce networks that are amenable to certification, but sacrifice standard accuracy. Conversely, adversarial training often yields stronger empirical robustness and standard accuracy, but the resulting models are generally difficult to certify with neural network verifiers. Recently, the literature has shown that better standard-certified accuracy trade-offs can be achieved by combining adversarial training objectives with loose over-approximations based on Interval Bound Propagation (IBP), effectively interpolating between lower and upper bounds of the worst-case loss. Building on this, we introduce AD-CERT, a certified training objective that combines adversarial distillation with an IBP upper bound. We show that distilling adversarial information over the logit space from an empirically robust teacher provides an effective lower bound surrogate for certified training, with AD-CERT achieving state-of-the-art certified performance on several robustness benchmarks. Furthermore, in a unified setup, distilling adversarial information at the logit-level is shown to improve certified accuracy over a robust feature-space distillation objective by up to 5.40 percentage points.
Matteo Melis, Jesus Martinez Del Rincon, Vishal Sharma
School of EEECS Queen’s University Belfast · School of EEECS2026 Queen’s University Belfast
The Lipschitz property of a deep neural network provides a direct measure of its sensitivity to input perturbations and, when explicitly controlled, offers a principled way to limit the propagation of errors and improve robustness. Over the past decade, Lipschitz-bounded layers have been incorporated into increasingly expressive and high-performing deep models, narrowing the gap between empirical robustness and formal, by-design guarantees of stability. This article introduces the fundamental concepts underlying Lipschitz-bounded neural networks, explaining the principles behind Lipschitz-constrained layers, the mechanisms used to enforce their bounds, and how they yield robustness certificates at the cost of a single forward pass. The tutorial concludes by discussing emerging and open directions, highlighting Lipschitz control as a general framework for offering guaranteed, by-design stability.
Fabio Brau, Giorgio Piras, Maura Pintor +1
Department of Electrical and Electronic Engineering, University of Cagliari, Cagliari, Italy