cs.AIOct 1, 2026

Towards Reliable Vision-Language Models for Autonomous Driving

Authors: Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas, Stefan Wagner

Organizations: Technical University of Munich, Heilbronn, Germany · Mercedes Benz AG, Sindelfingen, Germany · University of Massachusetts, Amherst, USA

Abstract

Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions and their associated confidence. Such degradation is especially concerning in autonomous driving, where safety-critical decisions require models to make accurate predictions and recognize when their predictions may be unreliable. In this work, we evaluate five VLMs (Qwen3.5-9B, Gemma4-E4B, LLaVA-OneVision-7B, DriveFusion/DriveFusionQA-4B, and NVIDIA Alpamayo-1.5-10B) across four driving-related QA datasets with different visual input settings, including single-frame, multi-view, multi-frame, and monocular inputs. Our results show that the effects of visual corruption vary across models, datasets, and input settings, with changes in accuracy and confidence reliability and also differing across conditions. We then apply Visual Evidence Augmentation (VEA\mathrm{V}{\scriptstyle \mathrm{EA}}), a recent inference-time method to examine whether it can improve model reliability under degraded visual conditions. We find that VEA\mathrm{V}{\scriptstyle \mathrm{EA}} improves performance for some models and datasets, although the gains are not consistent across all settings.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Task-Aligned Stability Analysis of Vision-Language Models for Autonomous Driving Hazard Detection

    Jun 10, 2026Everett RichardsAutonomous DrivingInstability

  2. Understanding Adversarial Transferability in Vision-Language Models for Autonomous Driving: A Cross-Architecture Analysis

    Apr 30, 2026David Fernandez, Pedram MohajerAnsari, Amir Salarpour +1Autonomous DrivingAdversarial Training