cs.CVSep 28, 2026

CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving

Authors: Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li, Haotian Chen, Fan Shi, Mengyang Zhao, Wenjuan Meng, +8 more

Organizations: Fudan University · Yinwang Intelligent Technology Co., Ltd · Fuzhou University · Sun Yat-Sen University · Huawei Technology

Abstract

Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth, dynamic risk is equally critical for deciding how to reason in time-critical situations. We therefore propose CAR-VLA, a unified driving VLA model that jointly considers scene complexity and dynamic risk to guide reasoning depth, urgency, and focus. CAR-VLA maps four complexity--risk categories to three reasoning modes: \textit{Fast Intuition} for direct trajectory generation in simple low-risk scenes, \textit{Slow Thinking} for deliberate reasoning in complex low-risk scenes, and \textit{Reflex Response} for compact, hazard-focused reasoning in high-risk scenes regardless of complexity. Rather than merely shortening deliberation, Reflex Response centers reasoning on the most critical hazard and the immediate safe response. We train CAR-VLA through progressive supervised learning that links scene assessment, reasoning-mode selection, and trajectory generation, followed by reasoning-augmented reinforcement learning to improve driving quality and reasoning behavior. Experiments on NAVSIM v1(91.1 PDMS), NAVSIM v2(90.3 EPDMS), and Navhard(35.0 EPDMS) demonstrate competitive driving performance. Qualitative comparisons on navtest and in-house high-risk scenarios further illustrate risk-aware reasoning and hazard-responsive trajectory generation. The code for this paper will be released publicly at: https://github.com/chenxl124578/CAR-VLA.git

Figures & tables

Explore similar work

CardsList
  1. ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

    May 27, 2026Mohammadreza Teymoorianfard, Jean-Philippe Monteuuis, Jonathan Petit +1Autonomous DrivingBreak

  2. Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving

    Jun 30, 2026Anh Dung Dinh, Simon Khan, Flora SalimDiffusion-Based Vision-Language-ActionsAutonomous Driving

  3. SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model

    Apr 21, 2026Zewei Zhou, Ruining Yang, Xuewei +8Diffusion-Based Vision-Language-ActionsBridging Language