Hamilton-Jacobi reachability constructs safety certificates for specified dynamics and safety constraints, tying each certificate to the deployment context for which it is synthesized. We ask whether a single certificate can instead represent a family of context-dependent safety problems and be queried across deployment conditions without re-synthesis. We learn a backward reachable tube for an eight-state vehicle model conditioned on local boundary geometry, friction coefficient, and adversarial disturbance scale. Geometry enters through an ego-frame boundary observation that defines the local containment constraint, while friction and disturbance scale enter as explicit operating-condition variables. This allows the same value function to be queried across friction coefficients from 0.4 to 2.0 and on geometries absent from synthesis. On 11 held-out evaluation geometries, the certificate maintains containment across the full tested friction range, including simultaneous geometry and grip shifts, while remaining within 1.2 percentage points in intervention rate and 0.09 m/s in speed of certificates re-synthesized with knowledge of the test geometry. We then deploy the certificate as a sampled discrete-time control barrier function filter on a full-scale vehicle near the handling limit. Lateral containment holds in every hardware session under both adversarial driving and autonomous racing, with 99th-percentile acceleration magnitude reaching 0.99 g. Across three certificates evaluated under a fixed autonomous racing controller, lap time varies by only 3.1%, demonstrating that a context-conditioned reachability certificate can transfer to deployment geometries absent from synthesis with modest performance cost.
Figures & tables
Fig. 1: Filtered closed-loop operation at the deployment site. The vehicle is drawn every 9m of driven path over one 21.7s run, with corridor boundaries, centerline, and driven path overlaid.
Family
Count
Length (m)
Half-width (m)
Rmin (m)
Small-scale circuits
23
260–553
0.47–2.14
0.7–13.7
Full-scale circuits
4
3190–8188
3.55–8.60
9.2–21.3
Skidpad figures
11
135–300
3.40 (const.)
9.1–17.0
TABLE I: Track library. Half-widths are per side; Rmin is the minimum radius of curvature. All are single-lap closed loops.
TABLE II: State, input, and context bounds. All values are read from the deployed configuration.
Fig. 2: Assumed-grip mismatch. Rows are the plant grip μtrue , columns are the queried grip μq , and boxed cells are matched. Left: violations over 30 trials. Right: intervention rate.
Configuration
Viol. (%)
Interv. (%)
Speed (m/s)
NoFilter
100.0
n/a
n/a
RobustGrip
0.0
72.3
12.66
Ours
0.0
74.2
13.07
OneTrack
0.0
74.1
13.16
TABLE III: Generalization across held-out geometry. 11 held-out tracks, 30 trials each, μq=μtrue=0.92 , sd=0.5 . NoFilter completes no trial, so its speed is not a progress measure and is omitted. OneTrack is a geometry-specialized upper bound re-synthesized for each test track.
μtrue
0.4
0.6
0.92
1.2
1.6
2.0
Violations (%)
NoFilter
100.0
100.0
100.0
100.0
100.0
100.0
NominalGrip
100.0
100.0
0.0
33.3
56.7
53.3
Intervention rate (%)
NominalGrip
n/a
n/a
72.4
76.3
78.8
88.0
RobustGrip
76.3
74.3
71.6
72.1
74.3
75.0
TABLE IV: Generalization across grip, in-distribution geometry. Columns are plant grip, with μq=μtrue and sd=0.5 . 30 trials per cell on 11 training geometries. Only NoFilter and NominalGrip incur violations; all other configurations have zero violations. OneGrip is a grip-specialized upper bound re-synthesized at each column’s grip value.
μtrue
0.4
0.6
0.92
1.2
1.6
2.0
Violations (%)
NoFilter
100.0
100.0
100.0
100.0
100.0
100.0
NominalGrip
100.0
100.0
0.0
43.3
53.3
53.3
Intervention rate (%)
NominalGrip
n/a
n/a
72.8
77.8
80.7
90.4
OneTrack
77.9
75.1
74.1
71.7
70.5
68.8
TABLE V: Simultaneous geometry and grip shift. μq=μtrue , sd=0.5 , and 30 trials are used per cell. Only NoFilter and NominalGrip incur violations; all other configurations have zero violations.
Fig. 3: Measured and model-implied acceleration envelopes for the three autonomous tracker sessions. Both channels are smoothed with a 0.1s moving average. Solid curves show the measured envelope and dashed curves the envelope predicted by the vehicle model used by the certificate; the μ0=0.92 friction circle is shown for reference.
Fig. 4: One sustained override event. Top: where it happened, with the vehicle drawn through the corner and the instant of peak override boxed; inset shows the position on the full lap. Below: the driver’s steering request against the angle actually applied, cross-track error against the corridor edges.
Geometry
Model
Lap (s)
vˉ (m/s)
vmax (m/s)
p99g
Oval
InD
18.43 ± 0.07
12.80
17.14
0.870
Oval
Finetuned
17.85 ± 0.07
13.15
18.12
0.939
Oval
OOD
18.08 ± 0.04
12.94
18.43
0.803
Bean
InD
22.38 ± 0.13
13.57
19.27
0.965
Bean
OOD
22.30 ± 0.09
13.66
19.03
0.890
TABLE VI: Autonomous racing controller on two geometries. μq=0.92 and sd=0.6 ; only the value function differs. Acceleration is the 99th percentile of the acceleration-magnitude samples after a 0.1s moving average, computed over 10 laps per configuration.
Dynamical Systems (DS) are reactive motion policies representing vector fields trained with theoretical guarantees of stability and convergence. To ensure safety during deployment in unknown environments they must be locally reshaped, either through modulation or geometric control barrier function strategies. However, depending on the geometry of the obstacles and the complexity of the DS, these local strategies can lead the system to unavoidable collisions or spurious attractors. In this work, we certify safety with a value function drawn from the notion of backward reachability tube, which measures the worst-case safety along a rollout trajectory of the nominal DS. Usually, such a value function is intractable for a controlled system due to curse of dimensionality. We show that in the DS-based learning-from-demonstration setting, the absence of a control input collapses the reachability problem to a deterministic rollout, and the presence of certain stability conditions truncates the infinite horizon to a finite one, resulting in a well-defined value function. We further show that the value function we devised is the maximal forward-invariant subset of the obstaclefree region for the nominal DS flow. The application of this certificate function is validated across five DS constructions - analytical, Neural ODE, diffeomorphic latent space, LPV-DS, SE(3)and validate it on a Franka manipulator. Modulation and geometric CBFs also suffer from saddle point in cases of headon approach towards an unsafe zone. We show that CBF-on-V avoids this pitfall entirely.
Aditya Vats, Tianyi Xia, Nadia Figueroa
The authors are with the University of Pennsylvania, Philadelphia, PA, USA.
Emergency collision avoidance under extreme driving conditions demands safety-critical control that accounts for both obstacle proximity and vehicle dynamic stability over a future time horizon, yet existing methods often rely on instantaneous or local safety evaluations. This paper proposes a safe reinforcement learning framework guided by a Hamilton-Jacobi (HJ) reachability based motion safety set that provides forward-looking safety supervision for constrained policy optimization. Specifically, a unified signed safety function is formulated by combining geometric collision margins and chassis stability limits, and is then extended through reachability analysis into a finite-horizon motion safety set that characterizes whether safety can be maintained under future vehicle state evolution. To enable practical computation, the motion safety set is approximated from offline extreme driving data, mitigating the computational burden of grid-based HJ solvers. The learned motion safety set is then embedded as a continuous safety cost into a constrained Markov decision process, and a PID-Lagrangian policy optimization scheme is employed to adaptively regulate the Lagrange multiplier for safety constraint enforcement. Simulation and real-vehicle experiments on low-adhesion obstacle-avoidance scenarios demonstrate that the proposed method achieves higher goal-reaching rates, produces smoother avoidance maneuvers, and maintains larger unified safety margins than baseline methods.
Yuhong Jiang, Shiyue Zhao, Junzhi Zhang +4
School of Vehicle and Mobility, Tsinghua University, Beijing, China · State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University, Beijing, China
Safety filters have been shown to be effective tools to ensure the safety of control systems with unsafe nominal policies. To address scalability challenges in traditional synthesis methods, learning-based approaches have been proposed for designing safety filters for systems with high-dimensional state and control spaces. However, the inevitable errors in the decisions of these models raise concerns about their reliability and the safety guarantees they offer. This paper presents Adaptive Conformal Filtering (ACoFi), a method that combines learned Hamilton-Jacobi reachability-based safety filters with adaptive conformal inference. Under ACoFi, the filter dynamically adjusts its switching criteria based on the observed errors in its predictions of the safety of actions. The range of possible safety values of the nominal policy's output is used to quantify uncertainty in safety assessment. The filter switches from the nominal policy to the learned safe one when that range suggests it might be unsafe. We show that ACoFi guarantees that the rate of incorrectly quantifying uncertainty in the predicted safety of the nominal policy is asymptotically upper bounded by a user-defined parameter. This gives a soft safety guarantee rather than a hard safety guarantee. We evaluate ACoFi in a Dubins car simulation and a Safety Gymnasium environment, empirically demonstrating that it significantly outperforms the baseline method that uses a fixed switching threshold by achieving higher learned safety values and fewer safety violations, especially in out-of-distribution scenarios.
Sacha Huriot, Ihab Tabbara, Hussein Sibai
Computer Science & Engineering Washington University in St. Louis