Adaptive Safety Filtering for Frozen ACC Policies via Conformal Residual Calibration
Authors: Zhiruo Zhou, Rigaudiere Z. Li, Chen Xiwen, Yucheng Chen, Xiaojun Zhu, Houde Liu
Organizations: Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China · Wuhan University of Technology, Wuhan, China · Shanghai Jiao Tong University, Shanghai, China
Frozen adaptive cruise control (ACC) policies can violate constraints when deployment dynamics differ from their training conditions. We propose residual-aware conformal action filtering (RACF), which calibrates residuals of a fixed nominal predictor and converts their quantile into an operating margin for finite-model action projection. Completed transitions update margins and candidate selection without retraining the policy. In a registered comparison over 2,400 controller-trial units, Adaptive RACF achieves 94.3% episode safety, improving by 19.9 percentage points over the evaluated nominal CBF-QP baseline while reducing projection frequency from 8.11% to 6.63%. A controlled study isolates a 4.54-point improvement from residual-margin injection. In a separate matched-hardware evaluation, Adaptive reduces mean amortized rollout time by 21.2% relative to Robust CBF-QP, with 161/180 versus 170/180 safe episodes. We characterize conditions linking one-step residual coverage to constraint satisfaction and quantify the observed safety-computation trade-offs.
Figures & tables
Figure 1: RACF data flow. Offline calibration uses residuals of a fixed nominal predictor to initialize the operating margin. At runtime, the policy proposes atπ from ot , while the filter uses st and selected candidate constraints to compute at . Completed transitions update the margin and candidate set for the next step. An infeasible solve or failed post-check invokes bounded maximum braking.
Table 1: Separate cohorts. Registered: N=2400 ; Proj./Corr./Jerk are mean episode projection (%), normalized correction, and P95 jerk (m/s 3 ). Matched hardware: N=180 ; time is episode wall time divided by executed steps, averaged over episodes (Supplement Sec. 3.2).
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: Additional ACC results. Panels (a–b) show episode-safe rates with 95% shared-reset cluster-bootstrap intervals for policy groups and physical conditions in the main 12-controller primary matrix. IQL/BC use seeds 0–3/0–2 ( N=800/600 ); CQL uses mixed-data seeds 0–2 ( N=600 ), CQL-low the suboptimal-data seed-0 checkpoint ( N=200 ), and IDM has N=200 . CQL plus CQL-low forms the CQL family in Table 15 . (c) Independent one-step nominal-residual coverage on 500 disjoint transitions per condition, with exact 95% binomial intervals, target 1−α=0.9 , and an 84–101.5% detail axis. (d) Adaptive-RACF episode-safe rate with 95% shared-reset cluster-bootstrap intervals for IQL, CQL, and BC under action noise in the six-policy cohort. Color denotes methods in (a–b), while the inset legend denotes policy groups in (d).
Method
Safe
Proj.
Viol.
Jerk
Base robust projection
1993
.0669
.00323
2.858
+ Static RACF
2102
.0635
.00260
3.098
Appendix
Table 1: Controlled residual-margin comparison on 2,400 controller–trials. Safe is the exact safe-episode count; Proj. and Viol. are observed-step fractions; Jerk is the mean episode P95 in m/s 3 .
Method
Episode-safe
Collisions
Raw
1001
904
Nominal QP
1437
0
Nominal CBF-QP
1785
0
Static RACF
2133
0
Adaptive RACF
2262
0
Envelope RACF
2287
0
Appendix
Table 2: Main comparison on 2,400 controller–trial units. The lower block selects Envelope–Robust costs; Table 14 includes Adaptive. Corr. is normalized action-correction magnitude, Fall. is an observed-step fraction, Jerk is episode P95 in m/s 3 , Speed is speed RMSE in m/s, and Gap is gap RMSE in m.
Set
Safe/120
Viol. steps
Proj.
Fallback
Full grid
109
119
.0701
.0135
No mass
107
130
.0776
.0137
No grade
109
119
.0701
.0135
No gain
109
119
.0713
.0127
Nominal
107
130
.0714
.0133
Appendix
Table 3: Candidate-set deletion at fixed B=0.45 . Viol. steps is a count; Proj. and Fallback are observed-step fractions; Safe/120 is the exact safe-episode count.
λ
Safe count
Δ Safe (pp)
Jerk
Speed
Gap
0
320
0.00
1.000
1.000
1.000
0.1
316
−1.11
.922
1.001
1.009
1
293
−7.50
.686
1.011
1.084
10
296
−6.67
.304
1.155
2.192
Appendix
Table 4: Smoothing screen relative to λ=0 . Safe count is out of 360, Δ Safe is in percentage points, and Jerk/Speed/Gap are ratios. All variants complete 360 rollouts with zero collisions.
Setting
q
1-step
MaxCov
Safe
Viol.
.05/500/500
.0962
.989
.000
.800
211
.10/500/500
.0920
.970
.000
.792
219
.20/500/500
.0833
.939
.000
.783
225
.10/100/500
.0882
.959
.000
.800
214
.10/250/500
.0898
.963
.000
.800
222
.10/1000/500
.0908
.964
.000
.800
221
Appendix
Table 5: Sensitivity results. Setting encodes α/n/H for conformal rows and B,H for the reference. The quantity q is a dimensionless residual threshold; 1-step, MaxCov, and Safe are fractions, and Viol. counts violation steps.
Figure 3: Predictor-alignment and component comparisons in ACC partial rollouts . (a) Safety for Static, predictor-aligned RACF, and Robust CBF-QP across nominal, mass, and gain conditions. (b) Candidate-only, margin-only, and joint updates relative to Static. (c) Episode-max qH median and range (left axis) and stopped-episode coverage in percent (right axis); the transfer series changes the controller and is descriptive.
Method
Safe/ N
Coll.
∣Δa∣
FB (%)
J95
Raw
28/180
82
0.0000
0.00
0.76
Base projection
124/180
0
0.0152
1.32
2.13
Static RACF
144/180
0
0.0151
1.32
2.20
Aligned RACF
149/180
0
0.0155
1.36
2.22
Robust CBF
170/180
0
0.0146
0.89
2.03
Envelope RACF
164/180
0
0.0158
1.36
2.28
Appendix
Table 10: Predictor-alignment methods on 180 evaluation episodes. Safe is the headway-plus-collision endpoint; FB is fallback steps divided by observed steps. J95 is the mean per-episode jerk P95.
Condition
qH range
Frozen
Transfer
Nominal
0.4016–0.4133
56/60
54/60
Mass +20%
0.4012–0.4020
54/60
55/60
Gain 0.7
0.4005–0.4005
55/60
55/60
Appendix
Table 11: Episode-max calibration range by physical condition. Values pool the six policy–scenario strata per condition; coverage counts are descriptive.
Contrast
Safe ref
Safe trt
Δ safe (pp)
Proj. Δ
FB Δ
Candidate only − static
144
144
0.00 [0,0]
−0.00049
+0.00017
Margin only − static
144
161
+9.44 [4.44,15.56]
−0.00107
+0.00048
Both − margin only
161
161
0.00 [0,0]
+0.00026
+0.00007
Both − candidate only
144
161
+9.44 [4.44,15.56]
−0.00032
+0.00038
Appendix
Table 12: 2 × 2 ablation on 180 shared evaluation units. Continuous values are treatment minus reference; intervals are shared-cluster bootstrap 95% intervals.
Figure 4: Development-matched Pareto comparisons on independent test seeds. (a) Safety versus projection frequency. (b) Safety versus episode P95 jerk. (c) Safety versus mean action correction. Crosses are Static operating points, open markers are Adaptive points, and lines connect each eligible pair. Colors identify policy families. Marker labels give the residual-window length (40 or 100) and fixed-threshold or residual-triggered updating.
Figure 5: Descriptive rolling coverage from the ACI partial rollout. The blue ACI and orange fixed- q0 curves use a 50-step rolling window and coincide where the coverage values agree. The vertical line marks step 100 and the dashed horizontal line the 0.90 target.
Figure 6: Dynamic-response trace reconstructed from existing serialized observables. (a) Residual rt , margin mt , qα , and the 2qα proxy threshold. (b) Active-candidate count and policy/applied actions; red-outlined points mark emergency braking. (c) Predicted and actual safety margins. Vertical lines mark the proxy spike and residual recovery. The sequence is descriptive and does not imply that fallback preserves safety.
Figure 7: Mass-axis development diagnostic. (a) Episode-safe rate and (b) projection rate for Raw and four RACF operating points. Masses 1500 and 1800 kg are registered; 1200 and 2100 kg are extrapolation points. This development partial rollout is not a generalization guarantee.
Figure 8: Four-axis out-of-grid development summary from Job 82049. Panels show the controllable-safe rate for Raw, Static conformal, and Contextual temporal across mass, grade, actuator gain, and lead braking. Collision-free status, projection, fallback, correction, jerk, and paired cluster-bootstrap contrasts are reported in the numerical results rather than plotted here.
Figure 9: Independent-test deltas relative to the frozen sensitivity baseline. Rows vary one hyperparameter at a time; bars show setting-minus-baseline changes in projection and fallback (pp) and mean action correction ( 10−4 normalized action units). Safety endpoint deltas and cluster-bootstrap intervals are retained in the accompanying JSON.
Figure 10: Registered supporting slices. (a) Safety on shared trial prefixes; all horizons use fixed denominators and collision-stopped trajectories remain unsafe at longer prefixes. (b) Envelope–Raw safety gain over four physical conditions and two lead scenarios. (c) Raw/Static/Adaptive policy-family safety; denominators appear below each family and exact safe counts are in Table 15 .
Method
Resid.
Adapt.
Models
Fallback
Pred. filter
set
–
model
–
Adapt. conf.
quantile
yes
–
–
RACF (ours)
margin
yes
limited
bounded
Appendix
Table 13: Operational distinction between uncertainty-aware safety interfaces. Resid. identifies the calibrated object; Adapt. denotes online updating; Models denotes explicit predictive-model constraints; Fallback records whether a bounded emergency action is part of the interface.
Figure 11: Complete nine-method comparison on the shared 2,400-unit ACC cohort. (a) Episode-safe rate. (b) Fraction of observed steps projected. (c) Mean absolute normalized action correction. (d) Mean episode P95 jerk (m/s 3 ). ACI and CPSF-style share RACF’s policies, resets, trajectories, solver, and bounded fallback.
Method
Safe/2400
SR (%)
Coll.
Proj. (%)
Fallback (%)
Correction
P95 jerk
Raw
1001
41.7
904
0.00
0.00
0.0000
0.88
Nominal QP
1437
59.9
0
7.21
0.75
0.0170
2.48
Nominal CBF-QP
1785
74.4
0
8.11
0.60
0.0177
2.64
Static RACF
2133
88.9
0
6.51
0.91
0.0172
3.13
ACI
2137
89.0
0
6.57
0.90
0.0173
3.07
CPSF-style
2130
88.8
0
6.49
0.91
0.0172
3.11
Appendix
Table 14: Safety–utility results on the shared ACC cohort. Safe is the exact safe-episode count and SR its percentage; Coll. counts collision-terminated episodes. Projection and fallback are observed-step percentages, Correction is the mean absolute normalized action change, and jerk is episode P95 in m/s 3 . Utility uses observed prefixes, so Raw collisions confound cross-method utility comparisons through survival/truncation. All filtered methods record zero collisions. Half-up rounding gives Robust fallback 0.73% from exactly 8,700/1,200,000 steps (fraction 0.00725).
Family
Policies
Trials
Raw
Static
Adaptive
BC
3
600
282
543
581
IQL
4
800
336
726
764
CQL
4
800
290
667
717
IDM
1
200
93
197
200
Appendix
Table 15: Policy-family decomposition of the registered comparison. Policies and Trials give the family denominator; Raw, Static, and Adaptive report exact safe-episode counts.
Comparison
Safe/total
Δ pp
Global fixed / pooled
170/180 / 170/180
0.00
Added f0 , qα / singleton
140/180 / 97/180
+23.89
Added f0 , Lhqα / singleton
143/180 / 126/180
+9.44
Appendix
Table 16: Constraint comparisons on 180 shared units. Safe/total gives exact paired counts; Δ pp is the first configuration minus the second. Adding f0 changes both exact- f0 inclusion and constraint cardinality.