Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, and verify relations before deciding how to act, leaving the support for risk conclusions implicit. We address this relevance--applicability gap with a Driving-Risk Knowledge Graph (DRKG) and Semantic Web Rule Language (SWRL) reasoning stage before VLM decision-making. Structured perception instantiates scene facts, from which SWRL rules derive events and directed risk relations when their antecedents are jointly satisfied. Recognized events, bound risk relations, and semantic descriptions of activated rules form compact evidence that conditions the VLM and diffusion planner. In matched comparisons on nuReasoning, our method improved the nuReasoning planning score (NPS) by 1.30 points and the non-at-fault collision score (NC) by 2.76 points over the relevance retrieval-based baseline. These gains indicate that scene-applicable risk evidence improves safety-weighted planning relative to semantically retrieved risk knowledge.
Figures & tables
Fig. 1: Typical RAG-based driving methods provide semantically relevant knowledge while leaving scene grounding and applicability assessment to the VLM, whereas our DRKG-supported framework explicitly derives scene-grounded risk evidence before VLM reasoning and decision making.
Fig. 2: Overview of the scene-grounded risk-evidence vision language-driving pipeline. Structured perception results and ego state instantiate scene facts using the ontology of the driving-risk knowledge graph (DRKG). Its Semantic Web Rule Language (SWRL) rules derive scene events and risk relations between scene entities. Recognized events, risk relations and associated rule descriptions are combined with ego and scene context in a prompt, while camera observations provide the visual input to the vision–language model (VLM). The VLM decision representation conditions a diffusion planner to generate the ego trajectory.
Fig. 3: Closure-based SWRL risk entailment from scene facts. The initial fact set Fs is expanded by applying the ontology T and SWRL rules RG until no new assertion is derived. Intermediate events and relations lead to risk assertions in the closure Fscl . The right panel shows the general rule form and an illustrative vehicle-interaction rule.
Method
Risk rules
Risk chains
Recall ↑
Appl. ↑
Recall ↑
Appl. ↑
KnowVal top-5
18.32
11.11
6.52
1.09
DriveReg top-5
29.30
12.58
27.05
1.55
KnowVal top-16
65.57
8.05
50.73
1.11
DriveReg top-16
33.88
4.04
29.51
0.58
SWRL-activated
100.00
100.00
100.00
100.00
TABLE I: Rule recall and applicability (Appl.) against SWRL-derived scene references (%).
Risk information
NC ↑
DA ↑
EP ↑
CF ↑
HL ↑
NPS ↑
ADE ↓
KnowVal top-5
86.88
92.91
89.94
89.76
61.00
62.01
1.675
DriveReg top-5
87.27
93.70
89.91
89.76
60.76
62.53
1.676
KnowVal top-16
87.53
93.18
89.95
90.29
61.59
62.88
1.644
DriveReg top-16
87.14
93.18
89.59
89.24
61.11
62.49
1.681
Scene-grounded risk evidence ( ours )
90.29
92.91
89.79
90.29
60.17
64.18
1.728
TABLE II: Planning results with different risk-information sources.
Risk-evidence input
NC ↑
DA ↑
EP ↑
CF ↑
HL ↑
NPS ↑
ADE ↓
None
86.22
92.91
90.19
89.24
60.78
61.55
1.700
Inferred results
89.24
92.65
89.62
90.55
61.11
63.22
1.707
Activated rule descriptions
89.11
92.39
89.63
90.29
60.65
62.68
1.718
Inferred results and rule descriptions ( ours )
90.29
92.91
89.79
90.29
60.17
64.18
1.728
TABLE III: Planning ablation of scene-grounded risk-evidence representation.
Method
NC ↑
DA ↑
EP ↑
CF ↑
HL ↑
NPS ↑
ADE ↓
UniAD [ 26 ]
88.87
87.62
89.62
92.62
48.80
55.65
2.054
DiffusionDrive [ 27 ]
90.22
88.25
90.46
94.96
51.96
57.86
1.930
AutoVLA [ 28 ]
90.92
86.48
89.33
99.90
49.89
59.05
2.063
SpanVLA [ 29 ]
93.78
88.35
85.72
99.80
49.13
60.59
1.890
Alpamayo-1.5 (zero-shot) [ 30 ]
90.26
86.13
86.51
97.93
33.79
50.45
2.925
nuVLA (planning only) [ 9 ]
94.87
92.10
87.38
99.70
55.22
64.98
1.937
TABLE IV: Selected planning baselines reported on nuReasoning [ 9 ] and our method.
Fig. 4: Qualitative comparison in an occluded pedestrian-crossing scene. The front camera shows a vehicle braking suddenly, whereas the top view reveals a pedestrian hidden from the ego view. Our method identifies the applicable risk rule and selects strong deceleration. The KnowVal-style baseline retrieves semantically related pedestrian rules (2 of top-5 are shown) and selects gentle deceleration.
Risk category
Types
Rules
Vars.
Atoms
Vehicle interaction
8
13
6.38 (4–11)
11.08 (6–18)
Vulnerable road users
4
6
4.33 (4–6)
7.67 (7–10)
Road geometry and traffic conditions
11
14
6.00 (5–9)
10.64 (7–20)
Occlusion and blind-spot risk
4
6
4.67 (4–6)
9.00 (7–11)
Total
27
39
5.67 (4–11)
10.08 (6–20)
TABLE V: Coverage and complexity of risk-specific SWRL rules.
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric knowledge limits generalization in long-tail scenarios. While Retrieval-Augmented Generation (RAG) offers a solution by accessing external expert priors, standard visual retrieval suffers from high latency and semantic ambiguity. To address these challenges, we propose \textbf{VLADriver-RAG}, a framework that grounds planning in explicit, structure-aware historical knowledge. Specifically, we abstract sensory inputs into spatiotemporal semantic graphs via a \textit{Visual-to-Scenario} mechanism, effectively filtering visual noise. To ensure retrieval relevance, we employ a \textit{Scenario-Aligned Embedding Model} that utilizes Graph-DTW metric alignment to prioritize intrinsic topological consistency over superficial visual similarity. These retrieved priors are then fused within a query-based VLA backbone to synthesize precise, disentangled trajectories. Extensive experiments on the Bench2Drive benchmark establish a new state-of-the-art, achieving a Driving Score of 89.12.
Rui Zhao, Haofeng Hu, Zhenhai Gao +2
College of Automotive Engineering, Jilin University · The National Key Laboratory of Automotive Chassis Integration and Bionics, Jilin University · ReeFocus AI Technology
Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between reasoning and trajectory generation. However, the robustness of such systems under realistic input perturbations remains largely unexplored. We show that these models are highly vulnerable to realistic input perturbations, achieving up to 89% attack success rate (ASR) on reasoning and up to 72% on trajectory manipulation in closed-loop simulation, leading to increased collision rates and degraded safety metrics. Using NVIDIA's recent Alpamayo models as representative industry-developed VLAs, we conduct the first systematic black-box study of reasoning-enabled VLA models under realistic textual input corruptions, evaluating their impact on reasoning and driving behavior. We introduce a reasoning-aware evaluation framework capturing both semantic and structural aspects of reasoning, along with safety-centric measures. We also introduce a benchmark for evaluating attacks and defenses on reasoning-trajectory interactions in autonomous driving. Our results highlight the need for rigorous evaluation and improved defenses to ensure the safety of reasoning-enabled VLA systems in autonomous driving.
Mohammadreza Teymoorianfard, Jean-Philippe Monteuuis, Jonathan Petit +1
Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and expose intermediate decisions in natural language, yet current rationales often lack the step-by-step decision semantics needed to keep the rationale causally connected to the planned motion. We introduce Neuro-Symbolic Drive, a neuro-symbolic driving framework that supervises a driving VLA with rule-grounded reasoning traces extracted directly from classical rule-based planners. Our key observation is that rule-based planners are symbolic AI systems that already function as executable reasoning engines: they reason about active safety constraints, search over candidate maneuvers, and select a final trajectory. We instrument these planners in simulation to capture both the executed trajectory and the internal decision trace at each rule-evaluation step. Each trace is serialized into structured rule-grounded reasoning and paired with the trajectory to fine-tune Qwen3.5-4B as a driving VLA. Because these traces are derived directly from the planner states that determine the action, they ensure reasoning is structurally coupled to motion generation by construction, rather than by post-hoc alignment. On our simulator-generated benchmark, detailed rule-grounded reasoning reduces ADE@3s from 0.47 to 0.26 and miss rate from 8.30% to 6.40% under three-camera perception, and from 0.54 to 0.26 and 10.13% to 5.99% under eight-camera perception. Neuro-Symbolic Drive thus converts neuro-symbolic planning logic into structured supervision. Code base: https://github.com/XiangboGaoBarry/Neural-Symbolic-Drive.
Xiangbo Gao, Xiukun Huang, Boyu Lu +5
1Texas A&M University · 2Carnegie Mellon University · University of Maryland +2