Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth, dynamic risk is equally critical for deciding how to reason in time-critical situations. We therefore propose CAR-VLA, a unified driving VLA model that jointly considers scene complexity and dynamic risk to guide reasoning depth, urgency, and focus. CAR-VLA maps four complexity--risk categories to three reasoning modes: \textit{Fast Intuition} for direct trajectory generation in simple low-risk scenes, \textit{Slow Thinking} for deliberate reasoning in complex low-risk scenes, and \textit{Reflex Response} for compact, hazard-focused reasoning in high-risk scenes regardless of complexity. Rather than merely shortening deliberation, Reflex Response centers reasoning on the most critical hazard and the immediate safe response. We train CAR-VLA through progressive supervised learning that links scene assessment, reasoning-mode selection, and trajectory generation, followed by reasoning-augmented reinforcement learning to improve driving quality and reasoning behavior. Experiments on NAVSIM v1(91.1 PDMS), NAVSIM v2(90.3 EPDMS), and Navhard(35.0 EPDMS) demonstrate competitive driving performance. Qualitative comparisons on navtest and in-house high-risk scenarios further illustrate risk-aware reasoning and hazard-responsive trajectory generation. The code for this paper will be released publicly at: https://github.com/chenxl124578/CAR-VLA.git
Figures & tables
Fig. 1: Comparison of Reasoning-based VLA Paradigms. (a) Prior adaptive reasoning methods adaptively choose whether to perform reasoning, without recognizing that different scenes require different reasoning styles. (b) Our CAR-VLA analyzes scene complexity and risk, and routes different scene categories to different reasoning modes. These choices represent different reasoning focuses and depths, satisfying the reasoning requirements of diverse scenarios. The results below demonstrate CAR-VLA’s superior performance on the NAVSIM leaderboard.
Fig. 2: Overview of CAR-VLA. (a) CAR-VLA separately assesses scene complexity and risk, mapping four scene categories to three driving modes: Fast Intuition for simple, low-risk scenes, Slow Thinking for complex, low-risk scenes, and Reflex Response for high-risk scenes regardless of complexity. The selected mode guides subsequent trajectory generation. (b) Teacher-assisted data construction produces two complementary datasets: scene assessment data containing assessment reasoning and scene labels, and mode-adaptive driving data containing scene labels, driving reasoning, and trajectories. (c) Progressive supervised fine-tuning jointly learns general driving knowledge and scene assessment in Stage I, then learns scene classification, mode-specific reasoning, and trajectory generation in Stage II. (d) Reasoning-augmented reinforcement learning updates the policy through GSPO using driving, geometry, format, and reasoning rewards. The reasoning reward combines scene-assessment correctness with alignment to reference reasoning.
Method
NAVSIM v1
NAVSIM v2
NC ↑
DAC ↑
EP ↑
TTC ↑
C. ↑
PDMS ↑
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Traditional End-to-End Methods
HydraMDP++ [ 22 ]
97.6
96.0
80.4
93.1
100.0
86.6
97.2
97.5
99.4
99.6
83.1
96.5
94.4
98.2
70.9
81.4
DriveSuprim [ 23 ]
97.8
97.3
86.7
93.6
100.0
89.9
97.5
96.5
99.4
99.6
88.4
96.6
95.5
98.3
77.0
83.1
DiffusionDrive [ 24 ]
98.2
96.2
82.2
94.7
100.0
88.1
98.2
95.9
99.4
99.8
87.5
97.3
96.8
98.3
87.7
84.5
ResAD [ 25 ]
98.0
97.5
83.3
94.1
100.0
88.8
97.8
97.2
99.5
99.8
88.2
96.9
97.0
98.4
88.2
85.5
TABLE I: Performance Comparison on Navtest split in NAVSIM v1 and v2 using Closed-Loop Metrics.
Method
Stage
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
S. ↑
EPDMS ↑
Traditional End-to-End Methods
LTF
S1
96.2
79.6
99.1
99.6
84.1
95.1
94.2
97.6
79.1
–
25.1
S2
77.8
70.2
84.3
98.1
85.1
85.1
45.4
95.7
76.0
–
GuideFlow
S1
96.6
80.5
96.3
99.3
82.3
94.9
91.5
97.7
67.8
–
27.1
S2
87.3
76.7
88.8
99.2
84.3
85.1
49.7
93.1
44.5
–
VLA-based Methods
TABLE II: Performance Comparison on Navhard split in NAVSIM v2 using Closed-Loop Metrics.
Fig. 4: Qualitative comparison of AutoVLA and CAR-VLA. CAR-VLA adapts its reasoning to simple, high-risk (top) and complex, low-risk (bottom) scenes through Reflex Response and Slow Thinking , respectively, achieving higher PDMS in both cases and correctly identifying the risk factors.
Fig. 5: Qualitative comparison in a pedestrian-crossing scenario. CAR-VLA recognizes the high risk despite low scene complexity and invokes Reflex Response , producing a trajectory closer to the ground truth than AutoVLA (ADE: 1.289 m vs. 9.519 m).
Reward
NC
DAC
DDC
TLC
EP
TTC
LK
HC
EC
EPDMS
SFT(baseline)
98.6
96.6
99.5
99.9
86.3
98.0
97.4
98.3
86.0
88.5
Rd.+Rf.
97.1
95.5
98.9
99.8
93.5
96.7
96.2
97.9
84.7
87.2 (-1.3)
+Rg.
98.4
97.8
99.3
99.8
88.2
97.7
96.9
98.2
86.4
89.7 (+1.2)
+Rr.
98.5
98.1
99.4
99.9
89.0
97.9
97.4
98.3
85.8
90.3 (+1.8)
TABLE III: Ablation study of RL reward components. Values in parentheses denote EPDMS changes relative to the baseline.
Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework capturing both semantic and structural aspects of reasoning, along with safety-centric measures. We also introduce a benchmark for evaluating attacks and defenses on reasoning-trajectory interactions in autonomous driving. Our results highlight the need for rigorous evaluation and improved defenses to ensure the safety of reasoning-enabled VLA systems in autonomous driving.
Mohammadreza Teymoorianfard, Jean-Philippe Monteuuis, Jonathan Petit +1
Modern Vision-Language-Action (VLA) planners for autonomous driving emit a chain-of-causation (CoC) reasoning step \emph{before} producing a trajectory. The reasoning is autoregressive and dominates inference latency, while the trajectory head is parallel and cheap. Latency is an operational constraint in autonomous driving, so accelerating the reasoning step is the central problem we address. We observe that CoC reasoning has two qualitatively different needs: most tokens continue routine setup that follows naturally from the ego-trajectory history, and a small fraction encode commitments that require fresh visual evidence about an unexpected situation. We split this reasoning into two specialized paths: a \emph{routine reasoner} that handles the predictable continuation by attending to trajectory history, and a \emph{deliberative reasoner} (the unmodified VLA target) that handles novel cases by attending to current visual evidence, using the speculative decoding framework as the architectural template for how the two paths cooperate. Unlike standard speculative decoding, our routine reasoner is not a smaller replica of the target; the two reasoners are deliberately specialized to read different parts of the prompt. We propose two techniques to realize this. First, we introduce \textbf{FlatRoPE}, a 1D rotary positional embedding in the draft that breaks the rotational symmetry of the target's 3D M-RoPE, redirecting attention away from visual tokens and onto trajectory-history tokens. Second, we introduce \textbf{Action-aware RL (AARL)}, a post-training stage that uses an action-quality reward together with a static-reference KL anchor. Together, our two-reasoner system reduces the reasoning-step running time by approximately 4× relative to the original Alpamayo planner.
Anh Dung Dinh, Simon Khan, Flora Salim
School of Computer Science and Engineering UNSW Sydney · Air Force Research Laboratory
Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially in long-tail scenarios. However, existing VLA models often struggle with the high latency in action generation using an autoregressive generation framework and exhibit limited robustness. In this paper, we propose SpanVLA, a novel end-to-end autonomous driving framework, integrating an autoregressive reasoning and a flow-matching action expert. First, SpanVLA introduces an efficient bridge to leverage the vision and reasoning guidance of VLM to efficiently plan future trajectories using a flow-matching policy conditioned on historical trajectory initialization, which significantly reduces inference time. Second, to further improve the performance and robustness of the SpanVLA model, we propose a GRPO-based post-training method to enable the VLA model not only to learn from positive driving samples but also to learn how to avoid the typical negative behaviors and learn recovery behaviors. We further introduce mReasoning, a new real-world driving reasoning dataset, focusing on complex, reasoning-demanding scenarios and negative-recovery samples. Extensive experiments on the NAVSIM (v1 and v2) demonstrate the competitive performance of the SpanVLA model. Additionally, the qualitative results across diverse scenarios highlight the planning performance and robustness of our model.
Zewei Zhou, Ruining Yang, Xuewei +8
Tony · University of California, Los Angeles, USA · Motional, USA +1