Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth, dynamic risk is equally critical for deciding how to reason in time-critical situations. We therefore propose CAR-VLA, a unified driving VLA model that jointly considers scene complexity and dynamic risk to guide reasoning depth, urgency, and focus. CAR-VLA maps four complexity--risk categories to three reasoning modes: \textit{Fast Intuition} for direct trajectory generation in simple low-risk scenes, \textit{Slow Thinking} for deliberate reasoning in complex low-risk scenes, and \textit{Reflex Response} for compact, hazard-focused reasoning in high-risk scenes regardless of complexity. Rather than merely shortening deliberation, Reflex Response centers reasoning on the most critical hazard and the immediate safe response. We train CAR-VLA through progressive supervised learning that links scene assessment, reasoning-mode selection, and trajectory generation, followed by reasoning-augmented reinforcement learning to improve driving quality and reasoning behavior. Experiments on NAVSIM v1(91.1 PDMS), NAVSIM v2(90.3 EPDMS), and Navhard(35.0 EPDMS) demonstrate competitive driving performance. Qualitative comparisons on navtest and in-house high-risk scenarios further illustrate risk-aware reasoning and hazard-responsive trajectory generation. The code for this paper will be released publicly at: https://github.com/chenxl124578/CAR-VLA.git
Figures & tables
Fig. 1: Comparison of Reasoning-based VLA Paradigms. (a) Prior adaptive reasoning methods adaptively choose whether to perform reasoning, without recognizing that different scenes require different reasoning styles. (b) Our CAR-VLA analyzes scene complexity and risk, and routes different scene categories to different reasoning modes. These choices represent different reasoning focuses and depths, satisfying the reasoning requirements of diverse scenarios. The results below demonstrate CAR-VLA’s superior performance on the NAVSIM leaderboard.
Fig. 2: Overview of CAR-VLA. (a) CAR-VLA separately assesses scene complexity and risk, mapping four scene categories to three driving modes: Fast Intuition for simple, low-risk scenes, Slow Thinking for complex, low-risk scenes, and Reflex Response for high-risk scenes regardless of complexity. The selected mode guides subsequent trajectory generation. (b) Teacher-assisted data construction produces two complementary datasets: scene assessment data containing assessment reasoning and scene labels, and mode-adaptive driving data containing scene labels, driving reasoning, and trajectories. (c) Progressive supervised fine-tuning jointly learns general driving knowledge and scene assessment in Stage I, then learns scene classification, mode-specific reasoning, and trajectory generation in Stage II. (d) Reasoning-augmented reinforcement learning updates the policy through GSPO using driving, geometry, format, and reasoning rewards. The reasoning reward combines scene-assessment correctness with alignment to reference reasoning.
Method
NAVSIM v1
NAVSIM v2
NC ↑
DAC ↑
EP ↑
TTC ↑
C. ↑
PDMS ↑
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
EPDMS ↑
Traditional End-to-End Methods
HydraMDP++ [ 22 ]
97.6
96.0
80.4
93.1
100.0
86.6
97.2
97.5
99.4
99.6
83.1
96.5
94.4
98.2
70.9
81.4
DriveSuprim [ 23 ]
97.8
97.3
86.7
93.6
100.0
89.9
97.5
96.5
99.4
99.6
88.4
96.6
95.5
98.3
77.0
83.1
DiffusionDrive [ 24 ]
98.2
96.2
82.2
94.7
100.0
88.1
98.2
95.9
99.4
99.8
87.5
97.3
96.8
98.3
87.7
84.5
ResAD [ 25 ]
98.0
97.5
83.3
94.1
100.0
88.8
97.8
97.2
99.5
99.8
88.2
96.9
97.0
98.4
88.2
85.5
TABLE I: Performance Comparison on Navtest split in NAVSIM v1 and v2 using Closed-Loop Metrics.
Method
Stage
NC ↑
DAC ↑
DDC ↑
TLC ↑
EP ↑
TTC ↑
LK ↑
HC ↑
EC ↑
S. ↑
EPDMS ↑
Traditional End-to-End Methods
LTF
S1
96.2
79.6
99.1
99.6
84.1
95.1
94.2
97.6
79.1
–
25.1
S2
77.8
70.2
84.3
98.1
85.1
85.1
45.4
95.7
76.0
–
GuideFlow
S1
96.6
80.5
96.3
99.3
82.3
94.9
91.5
97.7
67.8
–
27.1
S2
87.3
76.7
88.8
99.2
84.3
85.1
49.7
93.1
44.5
–
VLA-based Methods
TABLE II: Performance Comparison on Navhard split in NAVSIM v2 using Closed-Loop Metrics.
Fig. 4: Qualitative comparison of AutoVLA and CAR-VLA. CAR-VLA adapts its reasoning to simple, high-risk (top) and complex, low-risk (bottom) scenes through Reflex Response and Slow Thinking , respectively, achieving higher PDMS in both cases and correctly identifying the risk factors.
Fig. 5: Qualitative comparison in a pedestrian-crossing scenario. CAR-VLA recognizes the high risk despite low scene complexity and invokes Reflex Response , producing a trajectory closer to the ground truth than AutoVLA (ADE: 1.289 m vs. 9.519 m).
Reward
NC
DAC
DDC
TLC
EP
TTC
LK
HC
EC
EPDMS
SFT(baseline)
98.6
96.6
99.5
99.9
86.3
98.0
97.4
98.3
86.0
88.5
Rd.+Rf.
97.1
95.5
98.9
99.8
93.5
96.7
96.2
97.9
84.7
87.2 (-1.3)
+Rg.
98.4
97.8
99.3
99.8
88.2
97.7
96.9
98.2
86.4
89.7 (+1.2)
+Rr.
98.5
98.1
99.4
99.9
89.0
97.9
97.4
98.3
85.8
90.3 (+1.8)
TABLE III: Ablation study of RL reward components. Values in parentheses denote EPDMS changes relative to the baseline.