Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, and verify relations before deciding how to act, leaving the support for risk conclusions implicit. We address this relevance--applicability gap with a Driving-Risk Knowledge Graph (DRKG) and Semantic Web Rule Language (SWRL) reasoning stage before VLM decision-making. Structured perception instantiates scene facts, from which SWRL rules derive events and directed risk relations when their antecedents are jointly satisfied. Recognized events, bound risk relations, and semantic descriptions of activated rules form compact evidence that conditions the VLM and diffusion planner. In matched comparisons on nuReasoning, our method improved the nuReasoning planning score (NPS) by 1.30 points and the non-at-fault collision score (NC) by 2.76 points over the relevance retrieval-based baseline. These gains indicate that scene-applicable risk evidence improves safety-weighted planning relative to semantically retrieved risk knowledge.
Figures & tables
Fig. 1: Typical RAG-based driving methods provide semantically relevant knowledge while leaving scene grounding and applicability assessment to the VLM, whereas our DRKG-supported framework explicitly derives scene-grounded risk evidence before VLM reasoning and decision making.
Fig. 2: Overview of the scene-grounded risk-evidence vision language-driving pipeline. Structured perception results and ego state instantiate scene facts using the ontology of the driving-risk knowledge graph (DRKG). Its Semantic Web Rule Language (SWRL) rules derive scene events and risk relations between scene entities. Recognized events, risk relations and associated rule descriptions are combined with ego and scene context in a prompt, while camera observations provide the visual input to the vision–language model (VLM). The VLM decision representation conditions a diffusion planner to generate the ego trajectory.
Fig. 3: Closure-based SWRL risk entailment from scene facts. The initial fact set Fs is expanded by applying the ontology T and SWRL rules RG until no new assertion is derived. Intermediate events and relations lead to risk assertions in the closure Fscl . The right panel shows the general rule form and an illustrative vehicle-interaction rule.
Method
Risk rules
Risk chains
Recall ↑
Appl. ↑
Recall ↑
Appl. ↑
KnowVal top-5
18.32
11.11
6.52
1.09
DriveReg top-5
29.30
12.58
27.05
1.55
KnowVal top-16
65.57
8.05
50.73
1.11
DriveReg top-16
33.88
4.04
29.51
0.58
SWRL-activated
100.00
100.00
100.00
100.00
TABLE I: Rule recall and applicability (Appl.) against SWRL-derived scene references (%).
Risk information
NC ↑
DA ↑
EP ↑
CF ↑
HL ↑
NPS ↑
ADE ↓
KnowVal top-5
86.88
92.91
89.94
89.76
61.00
62.01
1.675
DriveReg top-5
87.27
93.70
89.91
89.76
60.76
62.53
1.676
KnowVal top-16
87.53
93.18
89.95
90.29
61.59
62.88
1.644
DriveReg top-16
87.14
93.18
89.59
89.24
61.11
62.49
1.681
Scene-grounded risk evidence ( ours )
90.29
92.91
89.79
90.29
60.17
64.18
1.728
TABLE II: Planning results with different risk-information sources.
Risk-evidence input
NC ↑
DA ↑
EP ↑
CF ↑
HL ↑
NPS ↑
ADE ↓
None
86.22
92.91
90.19
89.24
60.78
61.55
1.700
Inferred results
89.24
92.65
89.62
90.55
61.11
63.22
1.707
Activated rule descriptions
89.11
92.39
89.63
90.29
60.65
62.68
1.718
Inferred results and rule descriptions ( ours )
90.29
92.91
89.79
90.29
60.17
64.18
1.728
TABLE III: Planning ablation of scene-grounded risk-evidence representation.
Method
NC ↑
DA ↑
EP ↑
CF ↑
HL ↑
NPS ↑
ADE ↓
UniAD [ 26 ]
88.87
87.62
89.62
92.62
48.80
55.65
2.054
DiffusionDrive [ 27 ]
90.22
88.25
90.46
94.96
51.96
57.86
1.930
AutoVLA [ 28 ]
90.92
86.48
89.33
99.90
49.89
59.05
2.063
SpanVLA [ 29 ]
93.78
88.35
85.72
99.80
49.13
60.59
1.890
Alpamayo-1.5 (zero-shot) [ 30 ]
90.26
86.13
86.51
97.93
33.79
50.45
2.925
nuVLA (planning only) [ 9 ]
94.87
92.10
87.38
99.70
55.22
64.98
1.937
TABLE IV: Selected planning baselines reported on nuReasoning [ 9 ] and our method.
Fig. 4: Qualitative comparison in an occluded pedestrian-crossing scene. The front camera shows a vehicle braking suddenly, whereas the top view reveals a pedestrian hidden from the ego view. Our method identifies the applicable risk rule and selects strong deceleration. The KnowVal-style baseline retrieves semantically related pedestrian rules (2 of top-5 are shown) and selects gentle deceleration.
Risk category
Types
Rules
Vars.
Atoms
Vehicle interaction
8
13
6.38 (4–11)
11.08 (6–18)
Vulnerable road users
4
6
4.33 (4–6)
7.67 (7–10)
Road geometry and traffic conditions
11
14
6.00 (5–9)
10.64 (7–20)
Occlusion and blind-spot risk
4
6
4.67 (4–6)
9.00 (7–11)
Total
27
39
5.67 (4–11)
10.08 (6–20)
TABLE V: Coverage and complexity of risk-specific SWRL rules.
College of Automotive Engineering, Jilin University · The National Key Laboratory of Automotive Chassis Integration and Bionics, Jilin University · ReeFocus AI Technology