The Lipschitz property of a deep neural network provides a direct measure of its sensitivity to input perturbations and, when explicitly controlled, offers a principled way to limit the propagation of errors and improve robustness. Over the past decade, Lipschitz-bounded layers have been incorporated into increasingly expressive and high-performing deep models, narrowing the gap between empirical robustness and formal, by-design guarantees of stability. This article introduces the fundamental concepts underlying Lipschitz-bounded neural networks, explaining the principles behind Lipschitz-constrained layers, the mechanisms used to enforce their bounds, and how they yield robustness certificates at the cost of a single forward pass. The tutorial concludes by discussing emerging and open directions, highlighting Lipschitz control as a general framework for offering guaranteed, by-design stability.
Figures & tables
Strategy
Robustify
Certify
Employment Time
Adversarial Training [ 13 ]
✓
✗
Train
Formal Verification [ 14 ]
✗
✓
Test
Randomized Smoothing [ 15 ]
✓
(✓)
Train/Test
Lipschitz-Bounded Networks [ 16 ]
✓
✓
Train/Test
TABLE I: Comparison of robustness-related strategies (i.e., defenses). Adversarial training increases robustness without providing guarantees, while Randomized smoothing certifies only in probabilistic terms. The Employment column indicates whether the method is applied at test time, requires training or both.
Fig. 1: Taxonomy of 1 -Lipschitz layer constructions, organized by enforcement of the Lipschitz constraints.
Fig. 2: The weight Wi lives in the ambient space, off O(d) , the set of admissible orthogonal weight matricesA plain gradient step on the classification loss ℓ takes it to W~i+1 , in a direction that has no reason to respect the constraint; the penalty term Rβ then pulls the iterate towards O(d) , producing the updated weight Wi+1 . Because the two components compete through the weight β , the iterate approaches the manifold but does not guarantee to reach it.
Fig. 3: Weights Wi lives in the manifold L , while the learnable parameters θi lives in the flat unconstrained parameter space Θ . The update of the parameter happens through a gradient step that moves θi to θi+1 . The map Φ then sends each parameter to a weight matrix that lies in O(d) by construction, so both the updated iterate Wi+1=Φ(θi+1) is admissible.
Layer
Dense: Φ(θ)
Convolution: Φ(θ)
Kernel Parameterization
AOL [ 30 ]
AOL(θ)=θ(D(θ)∑jθ⊤θij)−1/2
AOL(θ)=θ(D(θ)∑j,pθ⊤⋆θij,p)−1/2
BCOP [ 20 ]
BB(θ)=Ak , Ai+1=Ai(I+21(I−Ai⊤Ai))
BCOP(θ)=[BB(θ0),…,BB(θ2k−2)]†
SOC [ 31 ]
exp(A)=∑j≥0j!Aj , A=θ−θ⊤
exp⋆(A)=∑j≥0j!A⋆j , A=θ−θ⊤
Cholesky [ 37 ]
Chol(θ)=θL−⊤ , θ⊤θ=LL⊤
—
Spectral Parameterization
TABLE II: The parameterization map Φ for dense ( θ∈Rc×c ) and convolutional ( θ∈Rc×c×k×k ) parameters. For kernels, θ⊤ denotes the adjoint kernel and F the spatial DFT, so that Fθ is a collection of c×c complex blocks to which the dense map is applied blockwise (with ∗ in place of ⊤ ). σ denotes the ReLU, b a bias.
Fig. 4: The weights Wi starts on the manifold O(d) . The update is performed not in the ambient space (i.e., using the Euclidean gradient ∇ℓ(Wi) ), since it generally points off the manifold. However, the update is projected onto the tangent space TwiO(d) (shaded), giving the Riemannian gradient g , the steepest direction along the manifold. The step is then taken by the geodesic exponential map , which travels along the shortest path on O(d) leaving Wi with initial direction −αg , tracing the curve to Wi+1=expWi(−αg) ; the iterate has moved and is still admissible.
Fig. 5: In scale representation of the margin landscape of a naive ResNet-56 and a Lipschitz-bounded classifier [ 24 ] around a CIFAR-10 image on the 2D plane spanned by −∇M(x,y) (the boundary direction) and a random orthogonal direction. Distances are measured in ℓ2 in pixel space, where the bold curve is the decision boundary M=0 . For the ResNet-56, the margin drops from 10.2 to 0 within a surprisingly small (and unpredictable) distance from the boundary of ≈0.07 , while for the 1-Lipschitz model it decreases slowly with a bounded rate, offering a certified radius Cf(x)≈0.24 .
Fabio Brau is an Assistant Professor at the University of Cagliari, conducting research on AI security and trustworthiness. He holds a Master’s in Mathematics from the University of Pisa and a Ph.D. in Embedded Systems from the Scuola Superiore Sant’Anna, where he worked on robustness certification for deep neural networks. His research spans Adversarial Machine Learning and Trustworthy AI for Computer Vision and NLP systems, with contributions to Lipschitz-constrained architectures for certifiable robustness. He has published in top-tier venues including NeurIPS, CVPR, AAAI, and IEEE TPAMI, and is involved in European research projects on trustworthy AI.
Table 8
Giorgio Piras is a Postdoctoral Researcher at the sAIfer Lab, Department of Electrical and Electronic Engineering, University of Cagliari, Italy, working on adversarial machine learning and large language model security. He received his BSc (110/110, 2019) and MSc with honors (2021) from the University of Cagliari, and his PhD with honors (2025) from Sapienza University of Rome (National PhD-AI program), with the thesis Adversarial Pruning: Improving Evaluations and Methods , supervised by Prof. Battista Biggio. He was a visiting student at Karlsruhe Institute of Technology. He reviews for NeurIPS, ICLR, AAAI, USENIX, ACM CCS, Pattern Recognition, Neurocomputing, Machine Learning, and IEEE TIFS.
Table 9
Maura Pintor (Member, IEEE) is an Assistant Professor at the PRA Lab, Department of Electrical and Electronic Engineering, University of Cagliari, Italy, where she received her PhD in Electronic and Computer Engineering in 2022. Her research focuses on adversarial machine learning, particularly evaluating, debugging, and improving the robustness of machine learning systems to make AI more reliable, secure, and trustworthy. She is an Area Chair for NeurIPS, Associate Editor for Pattern Recognition, and Consulting Associate Editor for IEEE TIFS. She is a member of the IEEE Information Forensics and Security Technical Committee, IAPR, and ELLIS, and maintains the open-source library SecML-Torch.
Table 10
Battista Biggio (Fellow, IEEE) (MSc 2006, PhD 2010) is Full Professor at the University of Cagliari, Italy. He has provided pioneering contributions in machine learning security, playing a leading role in this field. His seminal paper on “Poisoning Attacks against Support Vector Machines” won the prestigious 2022 ICML Test of Time Award. His work on “Wild Patterns” won the 2021 Best Paper Award and Pattern Recognition Medal from Elsevier Pattern Recognition. He has managed more than 10 research projects, and serves as a PC member of ICML and USENIX Security, and as Area Chair of NeurIPS. He chaired IAPR TC1 (2016-2020), and served as Associate Editor for IEEE TNNLS, IEEE CIM, and Elsevier PRJ. He is now Associate Editor-in-Chief for PRJ. He is also Fellow of IEEE, Senior Member of ACM, and member of IAPR and ELLIS.
Lipschitz-based robustness certification bounds a network's sensitivity through concrete numerical computation rather than symbolic reasoning, and so scales efficiently. It is increasingly used even where verifiable guarantees matter. Yet, as with most prior work on robustness certification and verification, soundness is typically proved against a semantic model assuming exact real arithmetic. Deployed networks instead execute in floating-point, creating a gap between certified properties and executed behaviour. As motivating evidence, we give counterexamples showing that real arithmetic robustness guarantees can fail under floating-point execution, even for previously verified certifiers. We then develop a formal, compositional theory relating real arithmetic Lipschitz-based sensitivity bounds to floating-point execution under standard rounding-error models for feed-forward ReLU networks. We derive sound conditions for floating-point robustness, including bounds on certificate degradation and sufficient conditions for the absence of overflow. We also give an efficient floating-point Gram iteration algorithm for Lipschitz bounds and prove that it never under-estimates the true norm. Separately, when a model is certified pre-deployment, we show how measuring its actual deviation against a high-precision execution can substantially reduce certificate degradation. We formalise the theory and its soundness, and implement an executable certifier, evaluated across dense networks spanning image, tabular, and many-class classification. To our knowledge, ours is the first method for soundly accounting for floating-point effects in Lipschitz-based robustness certification, and, done efficiently, the first floating-point-sound robustness checking procedure of any kind to certify models' entire test sets -- even those with 500,000 examples -- while retaining enough precision to be practical.
As robustness verification methods are becoming more precise, training certifiably robust neural networks is becoming ever more relevant. To this end, certified training methods compute and then optimize an upper bound on the worst-case loss over a robustness specification. Curiously, training methods based on the imprecise interval bound propagation (IBP) consistently outperform those leveraging more precise bounding methods. Still, we lack an understanding of the mechanisms making IBP so successful. In this work, we thoroughly investigate these mechanisms by leveraging a novel metric measuring the tightness of IBP bounds. We first show theoretically that, for deep linear models, tightness decreases with width and depth at initialization, but improves with IBP training, given sufficient network width. We, then, derive sufficient and necessary conditions on weight matrices for IBP bounds to become exact and demonstrate that these impose strong regularization, explaining the empirically observed trade-off between robustness and accuracy in certified training. Our extensive experimental evaluation validates our theoretical predictions for ReLU networks, including that wider networks improve performance, yielding state-of-the-art results. Interestingly, we observe that while all IBP-based training methods lead to high tightness, this is neither sufficient nor necessary to achieve high certifiable robustness. This hints at the existence of new training methods that do not induce the strong regularization required for tight IBP bounds, leading to improved robustness and standard accuracy.
Yuhao Mao, Mark Niklas Müller, Marc Fischer +1
Department of Computer Science, ETH Zürich, Swizterland
Robustness of neural networks is commonly quantified via local or global Lipschitz constants. However, Lipschitz continuity can be overly coarse or overly restrictive as global robustness measure, failing to capture nuanced, data-dependent behavior. We propose a data-driven, architecture-agnostic framework based on the discrete modulus of continuity (DMOC), a non linear generalization of Lipschitz continuity that provides a finer notion of robustness. Unlike many existing approaches, DMOC does not require access to model internals and instead evaluates regularity relative to the data distribution. This shifts the focus from the model to the data, which provide a data-driven baseline of regularity against which the network's robustness is assessed. We establish convergence results for DMOC-induced seminorms with explicit data-driven rates in terms of the separation distance, and introduce a scalable minibatch algorithm that reduces the quadratic cost of exact computation, enabling application to large-scale data sets such as ImageNet. Empirically, DMOC serves as an architecture independent diagnostic: it distinguishes trained from untrained networks, reveals underfitting and overfitting regimes, and yields, as a special case, tight Lipschitz estimates comparable to state-of-the-art method such as ECLipsE and ECLipsE-fast.
Jürgen Dölz, Michael Multerer, Michele Palma
Institute for Numerical Simulation University of Bonn Germany · Dalle Molle Institute for Artificial Intelligence USI Lugano Switzerland