ExceptionDrive: A Planning-Oriented Counterfactual Corner-Case Benchmark for Autonomous Driving
Authors: Ziyi Luo, Zhe Sun, Yehao Lu, Lei Zhou, Lisheng Wu, Xuewei Li, Zequn Qin, Xi Li
Organizations: College of Computer Science and Technology, Zhejiang University, Hangzhou, China · Yinwang Intelligent Technology Co., Ltd., Shenzhen, China
Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDrive, a counterfactual planning benchmark that uses VLM-assisted screening, localized multi-view editing, and quality auditing to insert hazards into real nuScenes scenes while preserving their context. Its 21 tasks span six safety families and define hazard or conflict regions, local safety constraints, and acceptable responses. Because hazard insertion can invalidate the recorded human trajectory, our reference-free protocol evaluates edited predictions using Unsafe Rate (UR), Hazard Clearance Compliance (HCC), Hazard Proximity Response (HPR), and Counterfactual Trajectory Shift (CTS), which measure core-region intrusion, clearance compliance, clearance relative to a prescribed margin, and counterfactual trajectory change. Seven representative planners frequently intrude into hazard regions or provide insufficient clearance. We also develop a Reminder Agent that, without sample-specific task labels, converts visual evidence and the shared taxonomy into structured records of hazard presence, type, and a recommended high-level strategy. The agent neither predicts trajectories nor controls the vehicle; its records guide a VLM-based decision agent. In zero-shot experiments, the reminders improve strategy accuracy and reduce under-warning.
Figures & tables
Figure 1: Overview of ExceptionDrive. Our framework generates localized counterfactual hazards in real driving scenes, evaluates planner responses without requiring an expert reference trajectory for the edited scene, and produces structured safety-aware reminders. The benchmark includes 21 planning tasks across six planning-safety families.
Figure 2: Counterfactual corner-case construction pipeline. An edit assistant uses the task definition and real-scene context to produce task-aware positive and negative prompts. A localized driving-scene editor inserts the target hazard while preserving viewpoint, road geometry, lighting, traffic context, and unrelated background. A review assistant then audits task correctness, physical plausibility, background preservation, realism, and action determinacy. Accepted source–edit pairs provide hazard-localized, planning-oriented samples for evaluation.
Figure 3: Illustration of the core hazard-centric evaluation metrics used in ExceptionDrive. The blue collision tube is evaluated relative to the task-defined hazard Hi and its required clearance boundary dsafe . (a) Hazard Clearance Compliance (HCC) quantifies whether the trajectory maintains the prescribed clearance margin. (b) Hazard Proximity Response (HPR) characterizes whether the avoidance response is close to the minimum sufficient clearance.
Figure 4: Representative samples from ExceptionDrive across six safety families. Each example pairs an original image with its counterfactual edit. Red dashed boxes mark the inserted local hazards. The edits preserve the original viewpoint, road layout, illumination, weather, and surrounding traffic context while modifying only the task-relevant risk factor.
Model
HCC% ↑
UR% ↓
HPR% ↑
CTS
CTS x
CTS y
Impromptu VLA Chi et al. (2026)
19.64
78.92
18.87
2.20
1.86
0.70
OmniDrive Wang et al. (2025)
9.52
83.36
9.84
1.40
1.34
0.22
AutoVLA Zhou et al. (2026)
10.37
81.93
11.26
1.97
1.67
0.64
DrivoR Kirby et al. (2026)
4.86
92.13
5.42
1.25
1.19
0.22
LightEMMA Qiao et al. (2025)
18.66
79.51
18.04
1.68
1.61
0.23
ST-P3 Hu et al. (2022)
21.03
78.11
18.59
1.57
1.55
0.24
Table 1: Main planner evaluation results on the proposed corner-case benchmark. HCC is the primary metric and measures whether the planner produces the task-defined safe response in edited hazardous scenes. UR measures unsafe response into the annotated hazard region. HPR measures the efficiency of hazard avoidance. CTS reports the magnitude of counterfactual trajectory shift between source and edited inputs, with longitudinal and lateral components denoted by CTS x and CTS y .
Planner
Emergence
Blockage
Bypass
Signal
Inter.
Surface
autoVLA
9.90
12.86
10.22
9.78
9.92
9.56
Impromptu VLA
19.63
20.28
19.98
20.14
17.86
19.95
DrivoR
4.31
4.59
2.61
5.49
3.94
8.27
ST-P3
20.02
19.86
17.48
22.92
24.75
21.17
Table 2: Family-level HCC. The table reports planner response correctness across the six safety families, revealing which types of corner cases are most challenging for each model. Higher HCC indicates a larger fraction of edited hazardous scenes in which the planner produces the task-defined safe response.
Figure 5: Qualitative planner responses across the six ExceptionDrive safety families. For each family, (a) shows the source scene and (b) the edited counterfactual scene. Red dashed boxes denote the inserted hazards, while blue, orange, and green curves show trajectories from AutoVLA, Impromptu VLA, and DrivoR, respectively.
Method
Accuracy ↑
Under-warning Rate ↓
Qwen2.5VL-7B
51.86
45.53
Qwen2.5VL-7B + Reminder
70.71
29.18
Table 3: High-level strategy evaluation for the zero-shot VLM driving agent.
Mercedes-Benz AG, Research & Development, Sindelfingen, Germany · University of Tübingen, Tübingen, Germany · Karlsruhe Institute of Technology, ITIV, Karlsruhe, Germany