The Lipschitz property of a deep neural network provides a direct measure of its sensitivity to input perturbations and, when explicitly controlled, offers a principled way to limit the propagation of errors and improve robustness. Over the past decade, Lipschitz-bounded layers have been incorporated into increasingly expressive and high-performing deep models, narrowing the gap between empirical robustness and formal, by-design guarantees of stability. This article introduces the fundamental concepts underlying Lipschitz-bounded neural networks, explaining the principles behind Lipschitz-constrained layers, the mechanisms used to enforce their bounds, and how they yield robustness certificates at the cost of a single forward pass. The tutorial concludes by discussing emerging and open directions, highlighting Lipschitz control as a general framework for offering guaranteed, by-design stability.
Figures & tables
Strategy
Robustify
Certify
Employment Time
Adversarial Training [ 13 ]
✓
✗
Train
Formal Verification [ 14 ]
✗
✓
Test
Randomized Smoothing [ 15 ]
✓
(✓)
Train/Test
Lipschitz-Bounded Networks [ 16 ]
✓
✓
Train/Test
TABLE I: Comparison of robustness-related strategies (i.e., defenses). Adversarial training increases robustness without providing guarantees, while Randomized smoothing certifies only in probabilistic terms. The Employment column indicates whether the method is applied at test time, requires training or both.
Fig. 1: Taxonomy of 1 -Lipschitz layer constructions, organized by enforcement of the Lipschitz constraints.
Fig. 2: The weight Wi lives in the ambient space, off O(d) , the set of admissible orthogonal weight matricesA plain gradient step on the classification loss ℓ takes it to W~i+1 , in a direction that has no reason to respect the constraint; the penalty term Rβ then pulls the iterate towards O(d) , producing the updated weight Wi+1 . Because the two components compete through the weight β , the iterate approaches the manifold but does not guarantee to reach it.
Fig. 3: Weights Wi lives in the manifold L , while the learnable parameters θi lives in the flat unconstrained parameter space Θ . The update of the parameter happens through a gradient step that moves θi to θi+1 . The map Φ then sends each parameter to a weight matrix that lies in O(d) by construction, so both the updated iterate Wi+1=Φ(θi+1) is admissible.
Layer
Dense: Φ(θ)
Convolution: Φ(θ)
Kernel Parameterization
AOL [ 30 ]
AOL(θ)=θ(D(θ)∑jθ⊤θij)−1/2
AOL(θ)=θ(D(θ)∑j,pθ⊤⋆θij,p)−1/2
BCOP [ 20 ]
BB(θ)=Ak , Ai+1=Ai(I+21(I−Ai⊤Ai))
BCOP(θ)=[BB(θ0),…,BB(θ2k−2)]†
SOC [ 31 ]
exp(A)=∑j≥0j!Aj , A=θ−θ⊤
exp⋆(A)=∑j≥0j!A⋆j , A=θ−θ⊤
Cholesky [ 37 ]
Chol(θ)=θL−⊤ , θ⊤θ=LL⊤
—
Spectral Parameterization
TABLE II: The parameterization map Φ for dense ( θ∈Rc×c ) and convolutional ( θ∈Rc×c×k×k ) parameters. For kernels, θ⊤ denotes the adjoint kernel and F the spatial DFT, so that Fθ is a collection of c×c complex blocks to which the dense map is applied blockwise (with ∗ in place of ⊤ ). σ denotes the ReLU, b a bias.
Fig. 4: The weights Wi starts on the manifold O(d) . The update is performed not in the ambient space (i.e., using the Euclidean gradient ∇ℓ(Wi) ), since it generally points off the manifold. However, the update is projected onto the tangent space TwiO(d) (shaded), giving the Riemannian gradient g , the steepest direction along the manifold. The step is then taken by the geodesic exponential map , which travels along the shortest path on O(d) leaving Wi with initial direction −αg , tracing the curve to Wi+1=expWi(−αg) ; the iterate has moved and is still admissible.
Fig. 5: In scale representation of the margin landscape of a naive ResNet-56 and a Lipschitz-bounded classifier [ 24 ] around a CIFAR-10 image on the 2D plane spanned by −∇M(x,y) (the boundary direction) and a random orthogonal direction. Distances are measured in ℓ2 in pixel space, where the bold curve is the decision boundary M=0 . For the ResNet-56, the margin drops from 10.2 to 0 within a surprisingly small (and unpredictable) distance from the boundary of ≈0.07 , while for the 1-Lipschitz model it decreases slowly with a bounded rate, offering a certified radius Cf(x)≈0.24 .
Fabio Brau is an Assistant Professor at the University of Cagliari, conducting research on AI security and trustworthiness. He holds a Master’s in Mathematics from the University of Pisa and a Ph.D. in Embedded Systems from the Scuola Superiore Sant’Anna, where he worked on robustness certification for deep neural networks. His research spans Adversarial Machine Learning and Trustworthy AI for Computer Vision and NLP systems, with contributions to Lipschitz-constrained architectures for certifiable robustness. He has published in top-tier venues including NeurIPS, CVPR, AAAI, and IEEE TPAMI, and is involved in European research projects on trustworthy AI.
Table 8
Giorgio Piras is a Postdoctoral Researcher at the sAIfer Lab, Department of Electrical and Electronic Engineering, University of Cagliari, Italy, working on adversarial machine learning and large language model security. He received his BSc (110/110, 2019) and MSc with honors (2021) from the University of Cagliari, and his PhD with honors (2025) from Sapienza University of Rome (National PhD-AI program), with the thesis Adversarial Pruning: Improving Evaluations and Methods , supervised by Prof. Battista Biggio. He was a visiting student at Karlsruhe Institute of Technology. He reviews for NeurIPS, ICLR, AAAI, USENIX, ACM CCS, Pattern Recognition, Neurocomputing, Machine Learning, and IEEE TIFS.
Table 9
Maura Pintor (Member, IEEE) is an Assistant Professor at the PRA Lab, Department of Electrical and Electronic Engineering, University of Cagliari, Italy, where she received her PhD in Electronic and Computer Engineering in 2022. Her research focuses on adversarial machine learning, particularly evaluating, debugging, and improving the robustness of machine learning systems to make AI more reliable, secure, and trustworthy. She is an Area Chair for NeurIPS, Associate Editor for Pattern Recognition, and Consulting Associate Editor for IEEE TIFS. She is a member of the IEEE Information Forensics and Security Technical Committee, IAPR, and ELLIS, and maintains the open-source library SecML-Torch.
Table 10
Battista Biggio (Fellow, IEEE) (MSc 2006, PhD 2010) is Full Professor at the University of Cagliari, Italy. He has provided pioneering contributions in machine learning security, playing a leading role in this field. His seminal paper on “Poisoning Attacks against Support Vector Machines” won the prestigious 2022 ICML Test of Time Award. His work on “Wild Patterns” won the 2021 Best Paper Award and Pattern Recognition Medal from Elsevier Pattern Recognition. He has managed more than 10 research projects, and serves as a PC member of ICML and USENIX Security, and as Area Chair of NeurIPS. He chaired IAPR TC1 (2016-2020), and served as Associate Editor for IEEE TNNLS, IEEE CIM, and Elsevier PRJ. He is now Associate Editor-in-Chief for PRJ. He is also Fellow of IEEE, Senior Member of ACM, and member of IAPR and ELLIS.