In human-AI collaborative decision making, human review can prevent unsafe AI decisions, but each human judgment is costly. Treating human intervention after AI abstention as a one-off fallback misses the opportunity to improve future AI decisions for greater automation, yet AI adaptively learning from selectively queried human feedback breaks safety guardrails calibrated for old models. We approach this challenge with CARE---calibrated adaptive rectification and escalation---an end-to-end pipeline that combines AI models and human reviewers to guarantee safe, human-aligned decisions, while continuously learning from human feedback to achieve greater automation with fewer human queries. CARE is principled, general, modular, and works with any black-box AI model. Our novel adaptive calibration module guarantees risk control at every time step for any rectification module. We further show how CARE improves query efficiency when the AI model is well trained and the human-AI misalignment has a clear structure. Experiments on four safety-critical real-world datasets spanning driving, language, and robotics demonstrate that CARE achieves human-aligned decisions while reducing human queries by 25-81% relative to baselines.
Figures & tables
Figure 1: (left) Per-sample decision pipeline. Upon a query, CARE rectifies AI’s prediction and confidence by correcting the human–AI residual and reducing uncertainty, based on past human feedback. The rectified quantities are then passed through two checks: an active acquisition rule that escalates informative queries, and a confidence gate that ensures the risk of the final human–AI decision is below the tolerance. Once a query is escalated, the human’s feedback is used as the final decision and to update CARE ’s modules. (right) Adaptive learning from human feedback. CARE learns to adapt to the revealed human preference over time and across scenarios, continuously increasing alignment and reducing the uncertainty of the rectified AI model, and thus improving efficiency.
Figure 2: Sample queries and benchmark labels. nuPlan shows recorded motion; its label refers to the benchmark simulation result. iSafetyBench shows one video frame.
Figure 3: (a) iSafetyBench with AI-provided confidence and 10% preparation data for static methods. The x -axis is the total human query rate, and the y -axis is the empirical mean risk. Lines connect nominal risk targets of 2.5%, 5%, 7.5%, and 10%. Crosses mark risk exceeding the corresponding target. A Pareto frontier would lower-bound all methods from the lower left. (b)–(c) ASIMOV with inferred confidence, 5% risk target, and 20% preparation data for static methods. The x -axis covers all samples regardless of escalation, and the y -axis uses a rolling window of 100 samples. (d) iSafetyBench query-rate differences at the 2.5% risk target and 10% static preparation. Positive deviations indicate efficiency degradation.
Figure 4: Regional model improvement and ensemble risk at the 5% target, with 20% preparation data for static methods. Radars show successive snapshots; heatmaps summarize the final snapshot on the same held-out panel of 2,000 scenarios. The risk target is highlighted as a dashed ring in the lower row, and the red band and outlines indicate exceedance. Region definition: open-loop time to collision below 1 s for R1–R2, finite and at least 1 s for R3–R4, and infinite for R5–R6; average displacement error below 5 m for R1, R3, R5, and at least 5 m for R2, R4, R6.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
nuPlan
HateBench
iSafetyBench
ASIMOV
Stream dataset size
10,000
7,054
1,100
1,437
Reserved test panel
2,000
784
—
—
Rolling-window size
1,000
500
100
100
GP feature dimension
256
64
256
256
Particle system size
24
5
20
20
Preparation fractions (%)
[10, 20, 30, 40]
[10, 20, 30, 40]
[10, 20, 30, 40]
[10, 20, 30, 40]
Appendix
Table 1: Shared protocol and dataset-specific settings. The default operating point uses 20% preparation and a 5% risk target.
Figure 5: nuPlan risk–cost tradeoffs. (a) Targets of 1%, 3%, 5%, 7.5%, and 10%, with 20% static preparation. (b) 10%, 20%, 30%, and 40% preparation at the 5% target, shown with decreasing marker opacity.
Target
Prep.
SE
PPI+SE
static CARE
CARE
1%
10%
0.00 / 100.0
0.57 / 87.0
0.43 / 84.3
0.70 / 74.6
1%
20%
0.00 / 100.0
0.56 / 87.1
0.26 / 86.5
0.70 / 74.6
1%
30%
0.00 / 100.0
0.55 / 87.2
0.16 / 89.3
0.70 / 74.6
1%
40%
0.00 / 100.0
0.47 / 89.6
0.10 / 91.5
0.70 / 74.6
3%
10%
2.04 / 66.0
2.46 / 56.9
1.85 / 55.4
3.00 / 37.7
3%
20%
1.76 / 70.5
2.05 / 63.3
1.27 / 62.4
3.00 / 37.7
Appendix
Table 2: nuPlan risk–cost tradeoffs. Ensemble risk / oracle query rate (%) over the complete target and static-preparation grid.
Figure 6: nuPlan learning trajectories at the 5% target and 20% static preparation. Lines average trailing windows of up to 1,000 arrivals over ten trials. Gray shading marks the 2,000-arrival static preparation period. AI and PPI are shown as uncalibrated references in the model-risk panel. The dotted horizontal line is the nominal ensemble-risk target.
Region
Panel
AI risk
Final CARE (%)
size
(%)
Model risk
Ensemble risk
Queries
R1
TTC <1 , ADE <5
53
37.7
13.6±1.5
3.0±1.1
49.6±5.0
R2
TTC <1 , ADE ≥5
293
36.9
15.1±0.5
3.3±0.4
56.5±3.7
R3
Finite TTC ≥1 , ADE <5
320
6.9
5.1±0.2
3.5±0.2
13.8±1.9
R4
Finite TTC ≥1 , ADE ≥5
374
11.5
8.1±0.5
3.3±0.3
30.6±2.8
R5
Infinite TTC, ADE <5
468
6.8
6.3±0.2
3.9±0.3
11.9±1.4
Appendix
Table 3: nuPlan regional definitions and final frozen-model results at the 5% target. TTC (time to collision) is in seconds and ADE (average displacement error) is in meters.
Figure 7: HateBench risk–cost tradeoffs with different models. Targets of 1%, 3%, 5%, 7.5%, and 10%, with 10% static preparation.
Target
Prep.
SE
PPI+SE
static CARE
CARE
1%
10%
0.86 / 78.6
0.67 / 56.7
0.77 / 51.4
1.02 / 42.2
1%
20%
0.81 / 80.3
0.54 / 57.5
0.69 / 53.8
1.02 / 42.2
1%
30%
0.70 / 82.8
0.57 / 59.8
0.60 / 58.6
1.02 / 42.2
1%
40%
0.59 / 85.3
0.53 / 64.1
0.53 / 64.0
1.02 / 42.2
3%
10%
2.58 / 64.6
1.97 / 41.5
2.26 / 36.1
2.64 / 27.1
3%
20%
2.31 / 68.2
2.01 / 41.8
2.02 / 40.7
2.64 / 27.1
Appendix
Table 4: HateBench risk–cost tradeoffs with Perspective. Ensemble risk / oracle query rate (%) over the complete target and static-preparation grid.
Figure 8: HateBench learning trajectories with Perspective at the 5% target and 10% static preparation. Lines average complete trailing windows of 500 arrivals over ten trials. Gray shading marks the 705-arrival static preparation period. AI and PPI are shown as uncalibrated references in the model-risk panel. The dotted horizontal line is the nominal ensemble-risk target.
AI model
Raw AI error
Human query rate
Relative improvement
SE
PPI+SE
CARE
Query saving vs SE ↑
Query saving vs PPI+SE ↑
Model-error reduction ↑
Perspective
20.9
54.0±0.9
31.8±1.5
21.5±0.7
60.2
32.5
40.5
Moderation
15.3
35.7±1.0
30.2±1.7
15.6±1.2
56.3
48.2
15.5
Detoxify Original
34.1
64.0±0.5
39.3±1.5
27.7±1.0
56.7
29.6
58.6
Detoxify Unbiased
34.5
70.6±0.6
35.1±1.3
24.3±0.7
65.5
30.6
65.8
LFTW
18.6
52.0±1.6
36.9±1.2
24.2±1.0
53.4
34.3
27.2
Appendix
Table 5: HateBench comparison across AI models at the 5% target and 10% static preparation. All values are percentages. Raw AI error is measured on the stream; model-error reduction is measured on the reserved panel.
Figure 9: iSafetyBench risk–cost tradeoffs. (a, c) Targets of 2.5%, 5%, 7.5%, and 10%, with 10% static preparation. (b, d) 10%, 20%, 30%, and 40% preparation at the 5% target.
Figure 10: iSafetyBench learning trajectories at the 5% target and 10% static preparation. Lines average trailing windows of up to 100 arrivals over ten trials. Gray shading marks the 110-arrival static preparation period, where static model risks can include in-sample predictions. AI and PPI are shown as uncalibrated references in the model-risk panels. The dashed horizontal line is the nominal ensemble-risk target.
2.5% target
5% target
7.5% target
10% target
Method
ΔR
ΔQ
ΔR
ΔQ
ΔR
ΔQ
ΔR
ΔQ
SE
−0.5
+28.0
−1.3
+22.0
+0.1
+11.2
+0.7
+3.7
PPI+SE
−0.2
+11.8
−0.4
+17.1
+0.6
+7.1
+0.3
+9.6
GPC+SE
+0.1
−1.2
−0.2
+1.1
−0.3
+0.4
+0.1
−0.7
static CARE
−0.1
+0.3
+0.0
−0.7
−0.1
−0.6
−0.2
−0.3
CARE with PPI
+0.2
+14.4
−0.6
+5.6
−0.5
+2.1
+0.0
−3.0
Appendix
Table 6: Effect of confidence choice on iSafetyBench with 10% static preparation. Entries are paired mean differences (direct minus inferred), in percentage points, for ensemble risk ( ΔR ) and human query rate ( ΔQ ).
Figure 11: ASIMOV risk–cost tradeoffs. (a, b) Targets of 5%, 10%, 15%, and 20%, with 20% static preparation. (c, d) 10%, 20%, 30%, and 40% preparation at the 5% target with inferred confidence; the dotted horizontal line is the nominal ensemble-risk target.
Figure 12: ASIMOV learning trajectories at the 5% target and 20% static preparation with inferred confidence. Lines average trailing windows of up to 100 arrivals over ten trials. Gray shading marks the 287-arrival static preparation period, where static model risks can include in-sample predictions. AI and PPI are shown as uncalibrated references in the model-risk panel. The dotted horizontal line is the nominal ensemble-risk target. Some initial model-risk transients are outside the plotted range.
The use of Large Language Models (LLMs) across diverse areas of human activity-ranging from everyday tasks to safety-critical applications-aims to enhance decision-making effectiveness with minimal human feedback. Concurrently, it seeks to align decisions with human expectations, preferences, and needs while mitigating risks associated with AI non-determinism. However, humans frequently over- or under-rely on AI recommendations, and current AI systems remain poorly calibrated to human expectations. To address these challenges, we introduce a human-AI collaborative decision-making framework designed to augment human capabilities and align AI agents with human preferences and expectations. Specifically, this paper (a) formulates the collaborative decision-making task as a stochastic game between an AI agent and a human player, and (b) proposes the Human-Centric Reflective Architecture (HCRA), which integrates human-calibrated models with reinforcement learning agents that leverage linguistic feedback in an iterative, reflective process. Evaluation results demonstrate that HCRA enhances decision-making effectiveness and delivers high-quality recommendations.
Andreas Kouridakis, Dimitrios Patiniotis Spyropoulos, George Vouros
LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a guard model that labels each proposed action as safe or unsafe, but this binary view conflates two distinct decisions: whether the action is harmful in itself, and whether it is appropriate given the user's context. It also operates at the granularity of action categories rather than individual instances, producing routine interruptions that erode autonomy and train users to wave through the most consequential alerts. We reframe the problem as a per-instance three-way routing decision over {EXECUTE, ASK, REFUSE} and instantiate it with Safety Sentry, a lightweight guard model whose inference reduces to a single decoding call. A single decoding-time threshold lets one fixed checkpoint be re-positioned across deployments of differing risk tolerance without retraining. Safety Sentry outperforms a broad set of open-weight and frontier closed-source baselines on overall accuracy and safety-related recall, while controlling both directional error rates simultaneously.
Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes increasingly challenging. This paper explores how we can leverage AI to improve the quality of human oversight. We focus on an important safety problem that is already challenging for humans: fact-verification of AI outputs. We find that combining AI ratings and human ratings based on AI rater confidence is better than relying on either alone. Giving humans an AI fact-verification assistant further improves their accuracy, but the type of assistance matters. Displaying AI explanation, confidence, and labels leads to over-reliance, but just showing search results and evidence fosters more appropriate trust. These results have implications for Amplified Oversight -- the challenge of combining humans and AI to supervise AI systems even as they surpass human expert performance.
Rishub Jain, Sophie Bridgers, Lili Janzer +3
1Google DeepMind · 2Work done while previously at Google DeepMind