Towards Reliable Vision-Language Models for Autonomous Driving
Authors: Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner
Organizations: Technical University of Munich, Heilbronn, Germany · Mercedes Benz AG, Sindelfingen, Germany · University of Massachusetts, Amherst, USA
Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions and their associated confidence. Such degradation is especially concerning in autonomous driving, where safety-critical decisions require models to make accurate predictions and recognize when their predictions may be unreliable. In this work, we evaluate five VLMs (Qwen3.5-9B, Gemma4-E4B, LLaVA-OneVision-7B, DriveFusion/DriveFusionQA-4B, and NVIDIA Alpamayo-1.5-10B) across four driving-related QA datasets with different visual input settings, including single-frame, multi-view, multi-frame, and monocular inputs. Our results show that the effects of visual corruption vary across models, datasets, and input settings, with changes in accuracy and confidence reliability and also differing across conditions. We then apply Visual Evidence Augmentation (VEA), a recent inference-time method to examine whether it can improve model reliability under degraded visual conditions. We find that VEA improves performance for some models and datasets, although the gains are not consistent across all settings.
Figures & tables
Figure 1: Gemma4-E4B responses to clean and perturbed inputs.
Qwen3.5-9B
DriveFusionQA-4B
Gemma4-E4B
Condition
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Clean
55.8
53.6
21.5
29.3
50.4
49.0
13.5
27.2
55.2
55.1
19.5
28.4
Glare
56.4
54.6
19.8
28.2
50.9
48.7
12.3
26.9
54.7
52.3
20.2
28.8
Fog
55.3
52.2
21.9
29.5
52.2
49.4
11.2
26.6
53.6
52.8
21.3
29.5
Motion Blur
55.3
53.5
21.3
29.3
50.7
51.4
10.9
26.5
53.3
49.6
20.8
29.8
Lens Occl.
53.8
52.3
22.6
30.1
50.9
51.0
11.2
26.7
53.5
52.3
21.9
29.7
Table 1: DrivingVQA results under clean and corrupted conditions. We report Accuracy, AUROC VAUQ , ECE, and Brier score. Green and red show improvement and degradation relative to clean performance.
Qwen3.5-9B
DriveFusionQA-4B
Gemma4-E4B
Condition
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Clean
42.1
67.0
49.5
47.5
35.3
61.6
23.8
27.2
38.0
63.5
54.7
51.7
Glare
29.5
60.0
63.4
60.0
36.4
54.9
22.0
28.5
29.8
55.1
63.3
60.1
Fog
40.6
61.9
52.2
50.5
37.7
59.7
21.6
28.1
35.2
59.3
58.0
55.0
Motion Blur
41.5
61.8
51.9
50.4
36.3
58.3
23.1
28.3
34.6
57.8
59.5
56.8
Lens Occl.
40.9
65.4
51.2
49.4
37.2
59.3
23.3
28.1
36.3
62.4
57.1
54.0
Table 2: NuScenes-QA-mini results under clean and corrupted conditions. We report Accuracy, AUROC VAUQ , ECE, and Brier score. Green and red show improvement and degradation relative to clean performance.
Qwen3.5-9B
DriveFusionQA-4B
Gemma4-E4B
Condition
DVQA
NQA
OSR
STRIDE
DVQA
NQA
OSR
STRIDE
DVQA
NQA
OSR
STRIDE
Clean
+3.20
-0.18
+8.00
+6.00
+1.30
+1.25
+0.00
-0.90
+0.80
-0.81
-12.00
-1.10
Glare
+0.70
+5.37
-10.00
+8.60
+0.10
+0.72
+2.00
-1.60
-0.60
-0.54
-14.00
+0.80
Fog
+2.20
+0.27
-2.00
+6.50
-0.40
-1.43
-2.00
+1.30
+1.70
-0.09
-8.00
+1.60
Motion Blur
+1.50
-0.27
+4.00
+5.50
+0.20
+0.90
+0.00
+1.40
+3.00
-0.54
+2.00
-2.40
Lens Occl.
+5.00
+0.54
+8.00
+6.70
-0.10
+0.54
+0.00
+0.30
+2.80
+0.09
+8.00
+0.50
Table 3: Change in accuracy after applying VEA across all four datasets. Green and red indicate improvement and degradation, respectively.
Figure 2: Effect of VEA on Accuracy, AUROC VAUQ , ECE, and Brier score on DrivingVQA under clean and corrupted conditions.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Qwen3.5-9B
DriveFusionQA-4B
Gemma4-E4B
Condition
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Clean
18.1
51.1
53.1
43.3
8.9
54.4
31.2
18.1
13.1
53.1
61.7
49.8
Glare
15.9
49.3
55.4
44.3
9.4
51.5
30.1
17.8
12.1
55.1
62.3
49.7
Fog
17.7
43.9
54.1
44.3
7.6
50.8
31.7
17.2
14.6
49.5
58.9
47.6
Motion Blur
19.9
46.2
51.8
43.1
7.7
46.8
31.4
17.5
16.7
55.0
58.3
48.2
Lens Occl.
16.8
49.3
54.4
43.7
8.5
50.1
30.9
17.6
12.8
54.8
61.5
49.2
Appendix
Table 4: STRIDE-QA Bench results under clean and corrupted conditions. We report Accuracy, AUROC VAUQ , ECE, and Brier score. Green and red show improvement and degradation relative to clean performance.
Qwen3.5-9B
DriveFusionQA-4B
Gemma4-E4B
Condition
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Acc ↑
AUROC ↑
ECE ↓
Brier ↓
Clean
42.0
53.4
27.4
32.1
20.0
44.5
54.6
46.9
42.0
56.7
34.4
36.7
Glare
50.0
63.0
19.3
27.8
22.0
45.0
52.7
46.2
46.0
42.5
31.2
35.9
Fog
48.0
57.4
21.8
29.1
28.0
42.9
45.8
43.1
42.0
41.2
34.7
38.0
Motion Blur
36.0
63.2
33.5
33.1
22.0
36.1
50.9
44.2
36.0
38.2
40.9
40.9
Lens Occl.
38.0
48.6
30.6
33.1
24.0
42.1
49.5
43.9
40.0
37.8
35.7
38.8
Appendix
Table 5: OSR results under clean and corrupted conditions. We report Accuracy, AUROC VAUQ , ECE, and Brier score. Green and red show improvement and degradation relative to clean performance.
Figure 3: Effect of VEA on Accuracy, AUROC VAUQ , ECE, and Brier score on NuScenes-QA-mini under clean and corrupted conditions.
Figure 4: Effect of VEA on Accuracy, AUROC VAUQ , ECE, and Brier score on OSR under clean and corrupted conditions.
Figure 5: Effect of VEA on Accuracy, AUROC VAUQ , ECE, and Brier score on STRIDE-QA under clean and corrupted conditions.
Figure 6: Additional evidence extraction time introduced by VEA across models and datasets. Values show the average added time per sample.
top- k image token positions by mean attention weight
Appendix
Table 6: Hardware, generation, VAUQ hyperparameters, and evaluation sample sizes.
Model
Total layers
# Selected
Selected layer indices
Qwen3.5-9B
32
1
23
DriveFusionQA-4B
36
4
25, 27, 29, 31
Gemma4-E4B
42
5
29, 32, 35, 38, 41
LLaVA-OV-7B
28
3
25, 26, 27
Alpamayo
36
4
0, 2, 4, 13
Appendix
Table 7: VEA steering layers selected per model (from profiling, top 10% of layers by score, minimum 1). Alpamayo-1.5-10B uses a separate VEA implementation not covered by this profiling step.