Neural Barriers: An Online Certifiable Learning-enhanced Adaptive High Order Safety Critical Control
Authors: Lishuo Pan, Mattia Catellani, An Cao, Lorenzo Sabattini, Nora Ayanian
Organizations: Department of Computer Science, Brown University, Providence, RI 02912 USA · Department of Sciences and Methods for Engineering, University of Modena and Reggio Emilia, 41121 Modena, Italy · Division of Physics, Mathematics and Astronomy, California Institute of Technology, Pasadena, CA 91125 USA
Control barrier functions are an effective model-based tool to formally certify the safety of a system. However, transferring their theoretical guarantees to real-world robotics systems requires high model fidelity. For example, payloads or wind disturbances can cause significant model perturbations to an aerial vehicle, leading to safety compromises. In this work, we propose a certifiable online learning-enhanced robust adaptive control barrier function, which adapts to disturbances using a Neural ODE and quantifies its adaptation uncertainty with conformal prediction. Our approach guarantees safety at all time under unknown time-varying model disturbances. It adopts a conservative strategy when the adaptation uncertainty is high; and efficiently adapts to reduce controller conservativeness as it receives more data. Our approach provides a provable safety guarantee with a probability bound under suitable Lipschitz smoothness assumptions on the underlying model and trajectory. These results demonstrate the potential of our method as a practical safety controller for robotics system operating under model perturbations.
Figures & tables
Fig. 1 : A 38 g nano quadrotor tracks a figure-eight trajectory while keeping a safe distance from two obstacles, against an 7.6\mathrm{m}\mathrm{/}\mathrm{s}$$ wind. Our Neural Barrier provides a probabilistic safety guarantee while reducing the controller conservativeness online.
Fig. 2 : System overview of our Neural Barrier controller. The safety control solves the NODE-HO-RaCBF QP (at 100\mathrm{H}\mathrm{z} ), using the most recently trained residual model (updated at $2$\mathrm{H}\mathrm{z} ) and conformal prediction bounds (updated at 10\mathrm{H}\mathrm{z}$$ ), to certify the desired control.
Fig. 3 : Once the k -th model is trained, Qcalib is cleared and the k -th CP starts to quantify its estimation quality. The data collected during its training and prior data in the queue are used for the training of the (k+1) -th model.
Fig. 4 : Typical trajectories of the baseline and our controllers under different model perturbations. single-obstacle “Attractive” residual (top left). single-obstacle “Time-varying” residual (top right) and multi-obstacle “Time-varying” residual (bottom). The Gray sphere is the obstacle. The blue dot represents the initial state, and the red dot represents the current reference. The red cross indicates the collision. Once the robot collided, the trajectory is depicted in a dashed line.
Fig. 5 : Trajectories of offline Neural Barrier controller with confidence interval computed from Thm. 2 . The model perturbations are “Time-varying” with different magnitude k .
Fig. 6 : Trajectories of the HO-RaCBF controller with different hyperparameters. The blue dot represents the initial state, and the red dot represents the current reference. The red cross indicates the collision location. Once the robot collided, the trajectory is depicted in a dashed line.
Residual Type
Attractive
Repulsive
Time-varying
Method
hmin
hneg
Avg.
Avg.
CP
hmin
hneg
Avg.
Avg.
CP
hmin
hneg
Avg.
Avg.
CP
Dist.
sDist.
Cover.
Dist.
sDist.
Cover.
Dist.
sDist.
Cover.
RCBFs No Res.
0.02
0.0%
0.04
0.00
100%
0.02
0.0%
0.04
0.00
100%
0.02
0.0%
0.04
0.00
100%
RCBFs
8.31
0.0%
1.61
1.63
100%
10.19
0.0%
2.96
3.08
100%
5.57
0.0%
1.59
1.61
100%
RaCBFs (Tuned)
0.00
0.0%
0.10
0.00
-
0.0
0.0%
0.85
0.0
-
-0.42
42.5%
0.28
0.05
-
RaCBFs Γ=1
-0.73
94.1%
0.21
0.12
-
0.87
0.0%
0.35
0.16
-
-1.44
44.3%
0.33
0.19
-
TABLE I : Quantitative results for different residuals in single-obstacle simulation: Comparison between the baseline controllers and our Neural Barrier controllers (The “HO-” prefix is dropped in all methods due to space limit). The statistics are averaged across 10 trials.
Residual Type
Attractive
Repulsive
Time-varying
Method
hmin
hneg
Avg.
Avg.
CP
hmin
hneg
Avg.
Avg.
CP
hmin
hneg
Avg.
Avg.
CP
Dist.
sDist.
Cover.
Dist.
sDist.
Cover.
Dist.
sDist.
Cover.
RCBFs No Res.
0.04
0.0%
0.34
0.30
100%
0.04
0.0%
0.34
0.30
100%
0.04
0.0%
0.34
0.30
100%
RCBFs
3.03
0.0%
2.44
2.48
100%
2.69
0.0%
3.50
3.65
100%
2.96
0.0%
1.57
1.57
100%
RaCBFs (Tuned)
-0.04
31.9%
0.35
0.32
-
0.06
0.0%
0.34
0.30
-
-0.39
41.4%
0.37
0.34
-
RaCBFs Γ=1
-1.43
42.5%
0.52
0.48
-
1.57
0.0%
0.87
0.86
-
-0.61
14.9%
0.55
0.50
-
TABLE II : Quantitative results for different residuals in multi-obstacle simulation: Comparison between the baseline controllers and our Neural Barrier controllers (The “HO-” prefix is dropped in all methods due to space limit). The statistics are averaged across 10 trials.
Fig. 7 : (1−α) quantile of nonconformity scores sampled in online Neural Barrier controller. For each trial, a polynomial of degree 8 is fitted to all quantile samples, the mean and 95% confidence interval is plotted. The dot on the error bars represents the median, and the ends depict the 25th and 75th percentiles of the quantiles in a 1.33\mathrm{s}$$ time window. The statistics are averaged across 10 trials.
Fig. 8 : The time series of CP Coverage for our onlione and offline controller in both single- and multi-obstacle instances under different residual types. The statistics are averaged across 10 trials.
Fig. 9 : The ground truth residual dynamics of the x -velocity dimension (red), the neural network prediction (gray), and the CP predictive interval (blue) plots. On the left column, the plots are online residual learning results in single-obstacle “Attractive”, “Repulsive”, “Time-varying”, multi-obstacle “Attractive”, “Repulsive”, and “Time-varying”, respectively. On the right column, the plots are the offline residual learning results. Only one trial is depicted to demonstrate the exact coverage of our Neural Barrier controller.
Fig. 10 : The top views of 150\mathrm{s}$$ trajectories using different safety controllers in the physical experiments. The red arrows indicate the wind’s location and direction (the length of each arrow represents its strength).
Method
hmin
hneg
Avg.Dist.
Avg.sDist.
RCBFs - No Wind
0.07
0.00%
0.13
0.13
RCBFs
0.02
0.06%
0.27
0.28
RaCBFs State Y
-0.11
53.14%
0.06
0.06
RaCBFs Const. Y
-0.14
53.01%
0.06
0.06
Ours Online 500
-0.23
1.43%
0.18
0.17
Ours Online 2400
-0.01
0.38%
0.16
0.16
TABLE III : Quantitative comparison between the baseline controllers and our controllers (The “HO-” prefix is dropped in all methods due to space limit) in the physical experiments. The statistics are averaged across 5 trials.
Fig. 11 : Time series of measured and predicted dynamics residual in the velocity dimension, prediction error, and (1−α) quantile used for WCP in the “Online 2400” physical experiments. The 95% confidence intervals are over 5 trials.
As the number of autonomous robots continues to grow, safety becomes increasingly important. Control barrier functions (CBFs) provide a theoretically grounded framework for ensuring safety, but existing design methods often face limitations in effectiveness, scalability, or interpretability, and may result in overly conservative safe sets. In this paper, we propose \emph{VertexCBF}, a framework for learning neural CBFs in a scalable, systematic, and explainable way. We approximate the stationary Hamilton--Jacobi value function using a neural network trained via a combination of physics-informed and sparsely supervised learning. By exploiting control-affine dynamics and a convex polytope control set, under which the Hamiltonian is maximized at the control vertices, we efficiently generate supervision points via GPU-parallel vertex-restricted tree search, while a residual architecture guarantees that the learned CBF is never larger than the specified constraint function. We evaluate the method on 15 systems and compare it against relevant baselines, showing that it reliably recovers large safe sets where the baselines are conservative or fail completely. In addition, we perform a hardware experiment in which a mobile robot safely avoids pedestrians using a neural CBF trained with our method.
Bojan Derajić, Sebastian Bernhard, Wolfgang Hönig
Technical University of Berlin, Germany · AUMOVIO, Germany · Technical University of Applied Sciences Augsburg, Germany +1
Merely pursuing performance may adversely affect safety, while a conservative policy for safe exploration will degrade the performance. How to guarantee both safety and performance in learning-based control problems is an interesting yet challenging issue. This paper aims to enhance system performance with a safety guarantee by solving reinforcement learning (RL)-based optimal control problems for nonlinear systems subject to high-relative-degree state constraints and unknown time-varying disturbance/actuator faults. A new type of control barrier functions (CBFs), termed high-order reciprocal-based control barrier function, is proposed to handle high-relative-degree constraints, which extends the design of CBFs to enforce robust safety without knowing the disturbance bound. The concept of gradient similarity is proposed to quantify the relationship between safety and performance. Finally, gradient manipulation and adaptive mechanisms are introduced in the model-based safe RL framework to enhance the performance with a safety guarantee. Two simulation examples illustrate the efficacy of the proposed algorithms.
Xinyang Wang, Hongwei Zhang, Shimin Wang +2
Shenzhen Key Lab for Advanced Motion Control and Modern Automation Equipments, School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen, Guangdong 518055, China · School of Data Science, Lingnan University, Hong Kong · School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore, with M3S, SMART, and with MIT CSAIL +1
We augment the Slotine--Li adaptive controller for Euler--Lagrange systems with three learned components: a structured-quadratic Lyapunov function Vψ whose positive-definiteness follows from a Cholesky parameterization, a residual Soft Actor--Critic policy that adds bounded torque corrections to the analytic baseline, and a physics-informed neural network that estimates unmodeled dynamics. A closed-form safety filter, derived from the single affine constraint V˙ψ+αVψ≤0, projects every policy output onto the safe set without requiring an online QP solver. We prove: global feasibility of the filter under a drift-decay condition on the control-degeneracy set; exponential stability under exact shielding, with a robust extension whose margin depends on the PINN approximation error; almost-sure convergence of the three-timescale policy--certificate--multiplier updates to a KKT point; and a PAC generalization bound for the certificate over compacts. On a 2-DOF manipulator with nonlinear friction and variable payload, the learned certificate accounts for most of the empirical gain: tracking error drops by 41% on nominal friction and 24% on aggressive friction at the centroid of the training distribution. A 7-DOF scalability study on a Franka Emika Panda confirms clean convergence of the full pipeline at industrial scale, identifies the conditions under which gains over exact model-based baselines should and should not be expected, and documents a warm-start pathology of the learned certificate that has practical implications for deployment.
Giansalvo Cirrincione, Adriano Fagiolini
Laboratoire LTI, Université de Picardie Jules Verne, Amiens, France. · MIRPALab, Department of Engineering, University of Palermo, 90128 Palermo, Italy.