In human-AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention after AI abstention as a one-off fallback misses the opportunity to improve future AI decisions for greater automation, yet AI adaptively learning from selectively queried human feedback breaks safety guardrails calibrated for old models. We approach this challenge with CARE---calibrated adaptive rectification and escalation---an end-to-end pipeline that combines AI models and human reviewers to guarantee safe, human-aligned decisions, while continuously learning from human feedback to achieve greater automation with fewer human queries. CARE is principled, general, modular, and works with any black-box AI model. Our novel adaptive calibration module guarantees risk control at every time step for any rectification module. We further show how CARE improves query efficiency when the AI model is well trained and the human-AI misalignment has a clear structure. Experiments on four safety-critical real-world datasets spanning driving, language, and robotics demonstrate that CARE achieves human-aligned decisions while reducing human queries by 25-81% relative to baselines.
Figures & tables
Figure 1: (left) Per-sample decision pipeline. Upon a query, CARE rectifies AI’s prediction and confidence by correcting the human–AI residual and reducing uncertainty, based on past human feedback. The rectified quantities are then passed through two checks: an active acquisition rule that escalates informative queries, and a confidence gate that ensures the risk of the final human–AI decision is below the tolerance. Once a query is escalated, the human’s feedback is used as the final decision and to update CARE ’s modules. (right) Adaptive learning from human feedback. CARE learns to adapt to the revealed human preference over time and across scenarios, continuously increasing alignment and reducing the uncertainty of the rectified AI model, and thus improving efficiency.
Figure 2: Sample queries and benchmark labels. nuPlan shows recorded motion; its label refers to the benchmark simulation result. iSafetyBench shows one video frame.
Figure 3: (a) iSafetyBench with AI-provided confidence and 10% preparation data for static methods. The x -axis is the total human query rate, and the y -axis is the empirical mean risk. Lines connect nominal risk targets of 2.5%, 5%, 7.5%, and 10%. Crosses mark risk exceeding the corresponding target. A Pareto frontier would lower-bound all methods from the lower left. (b)–(c) ASIMOV with inferred confidence, 5% risk target, and 20% preparation data for static methods. The x -axis covers all samples regardless of escalation, and the y -axis uses a rolling window of 100 samples. (d) iSafetyBench query-rate differences at the 2.5% risk target and 10% static preparation. Positive deviations indicate efficiency degradation.
Figure 4: Regional model improvement and ensemble risk at the 5% target, with 20% preparation data for static methods. Radars show successive snapshots; heatmaps summarize the final snapshot on the same held-out panel of 2,000 scenarios. The risk target is highlighted as a dashed ring in the lower row, and the red band and outlines indicate exceedance. Region definition: open-loop time to collision below 1 s for R1–R2, finite and at least 1 s for R3–R4, and infinite for R5–R6; average displacement error below 5 m for R1, R3, R5, and at least 5 m for R2, R4, R6.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
nuPlan
HateBench
iSafetyBench
ASIMOV
Stream dataset size
10,000
7,054
1,100
1,437
Reserved test panel
2,000
784
—
—
Rolling-window size
1,000
500
100
100
GP feature dimension
256
64
256
256
Particle system size
24
5
20
20
Preparation fractions (%)
[10, 20, 30, 40]
[10, 20, 30, 40]
[10, 20, 30, 40]
[10, 20, 30, 40]
Appendix
Table 1: Shared protocol and dataset-specific settings. The default operating point uses 20% preparation and a 5% risk target.
Figure 5: nuPlan risk–cost tradeoffs. (a) Targets of 1%, 3%, 5%, 7.5%, and 10%, with 20% static preparation. (b) 10%, 20%, 30%, and 40% preparation at the 5% target, shown with decreasing marker opacity.
Target
Prep.
SE
PPI+SE
static CARE
CARE
1%
10%
0.00 / 100.0
0.57 / 87.0
0.43 / 84.3
0.70 / 74.6
1%
20%
0.00 / 100.0
0.56 / 87.1
0.26 / 86.5
0.70 / 74.6
1%
30%
0.00 / 100.0
0.55 / 87.2
0.16 / 89.3
0.70 / 74.6
1%
40%
0.00 / 100.0
0.47 / 89.6
0.10 / 91.5
0.70 / 74.6
3%
10%
2.04 / 66.0
2.46 / 56.9
1.85 / 55.4
3.00 / 37.7
3%
20%
1.76 / 70.5
2.05 / 63.3
1.27 / 62.4
3.00 / 37.7
Appendix
Table 2: nuPlan risk–cost tradeoffs. Ensemble risk / oracle query rate (%) over the complete target and static-preparation grid.
Figure 6: nuPlan learning trajectories at the 5% target and 20% static preparation. Lines average trailing windows of up to 1,000 arrivals over ten trials. Gray shading marks the 2,000-arrival static preparation period. AI and PPI are shown as uncalibrated references in the model-risk panel. The dotted horizontal line is the nominal ensemble-risk target.
Region
Panel
AI risk
Final CARE (%)
size
(%)
Model risk
Ensemble risk
Queries
R1
TTC <1 , ADE <5
53
37.7
13.6±1.5
3.0±1.1
49.6±5.0
R2
TTC <1 , ADE ≥5
293
36.9
15.1±0.5
3.3±0.4
56.5±3.7
R3
Finite TTC ≥1 , ADE <5
320
6.9
5.1±0.2
3.5±0.2
13.8±1.9
R4
Finite TTC ≥1 , ADE ≥5
374
11.5
8.1±0.5
3.3±0.3
30.6±2.8
R5
Infinite TTC, ADE <5
468
6.8
6.3±0.2
3.9±0.3
11.9±1.4
Appendix
Table 3: nuPlan regional definitions and final frozen-model results at the 5% target. TTC (time to collision) is in seconds and ADE (average displacement error) is in meters.
Figure 7: HateBench risk–cost tradeoffs with different models. Targets of 1%, 3%, 5%, 7.5%, and 10%, with 10% static preparation.
Target
Prep.
SE
PPI+SE
static CARE
CARE
1%
10%
0.86 / 78.6
0.67 / 56.7
0.77 / 51.4
1.02 / 42.2
1%
20%
0.81 / 80.3
0.54 / 57.5
0.69 / 53.8
1.02 / 42.2
1%
30%
0.70 / 82.8
0.57 / 59.8
0.60 / 58.6
1.02 / 42.2
1%
40%
0.59 / 85.3
0.53 / 64.1
0.53 / 64.0
1.02 / 42.2
3%
10%
2.58 / 64.6
1.97 / 41.5
2.26 / 36.1
2.64 / 27.1
3%
20%
2.31 / 68.2
2.01 / 41.8
2.02 / 40.7
2.64 / 27.1
Appendix
Table 4: HateBench risk–cost tradeoffs with Perspective. Ensemble risk / oracle query rate (%) over the complete target and static-preparation grid.
Figure 8: HateBench learning trajectories with Perspective at the 5% target and 10% static preparation. Lines average complete trailing windows of 500 arrivals over ten trials. Gray shading marks the 705-arrival static preparation period. AI and PPI are shown as uncalibrated references in the model-risk panel. The dotted horizontal line is the nominal ensemble-risk target.
AI model
Raw AI error
Human query rate
Relative improvement
SE
PPI+SE
CARE
Query saving vs SE ↑
Query saving vs PPI+SE ↑
Model-error reduction ↑
Perspective
20.9
54.0±0.9
31.8±1.5
21.5±0.7
60.2
32.5
40.5
Moderation
15.3
35.7±1.0
30.2±1.7
15.6±1.2
56.3
48.2
15.5
Detoxify Original
34.1
64.0±0.5
39.3±1.5
27.7±1.0
56.7
29.6
58.6
Detoxify Unbiased
34.5
70.6±0.6
35.1±1.3
24.3±0.7
65.5
30.6
65.8
LFTW
18.6
52.0±1.6
36.9±1.2
24.2±1.0
53.4
34.3
27.2
Appendix
Table 5: HateBench comparison across AI models at the 5% target and 10% static preparation. All values are percentages. Raw AI error is measured on the stream; model-error reduction is measured on the reserved panel.
Figure 9: iSafetyBench risk–cost tradeoffs. (a, c) Targets of 2.5%, 5%, 7.5%, and 10%, with 10% static preparation. (b, d) 10%, 20%, 30%, and 40% preparation at the 5% target.
Figure 10: iSafetyBench learning trajectories at the 5% target and 10% static preparation. Lines average trailing windows of up to 100 arrivals over ten trials. Gray shading marks the 110-arrival static preparation period, where static model risks can include in-sample predictions. AI and PPI are shown as uncalibrated references in the model-risk panels. The dashed horizontal line is the nominal ensemble-risk target.
2.5% target
5% target
7.5% target
10% target
Method
ΔR
ΔQ
ΔR
ΔQ
ΔR
ΔQ
ΔR
ΔQ
SE
−0.5
+28.0
−1.3
+22.0
+0.1
+11.2
+0.7
+3.7
PPI+SE
−0.2
+11.8
−0.4
+17.1
+0.6
+7.1
+0.3
+9.6
GPC+SE
+0.1
−1.2
−0.2
+1.1
−0.3
+0.4
+0.1
−0.7
static CARE
−0.1
+0.3
+0.0
−0.7
−0.1
−0.6
−0.2
−0.3
CARE with PPI
+0.2
+14.4
−0.6
+5.6
−0.5
+2.1
+0.0
−3.0
Appendix
Table 6: Effect of confidence choice on iSafetyBench with 10% static preparation. Entries are paired mean differences (direct minus inferred), in percentage points, for ensemble risk ( ΔR ) and human query rate ( ΔQ ).
Figure 11: ASIMOV risk–cost tradeoffs. (a, b) Targets of 5%, 10%, 15%, and 20%, with 20% static preparation. (c, d) 10%, 20%, 30%, and 40% preparation at the 5% target with inferred confidence; the dotted horizontal line is the nominal ensemble-risk target.
Figure 12: ASIMOV learning trajectories at the 5% target and 20% static preparation with inferred confidence. Lines average trailing windows of up to 100 arrivals over ten trials. Gray shading marks the 287-arrival static preparation period, where static model risks can include in-sample predictions. AI and PPI are shown as uncalibrated references in the model-risk panel. The dotted horizontal line is the nominal ensemble-risk target. Some initial model-risk transients are outside the plotted range.