Vision-language-action~(VLA) models have emerged as a promising paradigm for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a think with 2D first, drive with dedicated 3D priors paradigm. It first grounds 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support the trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.
Figures & tables
Figure 1: Comparison of geometry integration paradigms for VLA . (a) Structure-perception fusion injects agent and map tokens. (b) Geometric fusion incorporates global geometric tokens. (c) GeoCoTDrive introduces an explicit geometric chain-of-thought pipeline: VLM first grounds planning-relevant regions, then retrieves localized 3D geometric priors, and finally interleaves the grounding text with geometry tokens to condition trajectory generation.
Figure 2 : Overview of GeoCoTDrive. The model first grounds planning-critical 2D regions, retrieves localized geometry tokens from a geometric foundation model, and then interleaves them into the autoregressive context for trajectory generation.
Figure 3 : Overview of the PlanningGrounding data construction pipeline. The pipeline generates planning-oriented region annotations from driving scenes by identifying decision-critical visual cues. An annotation example and the overall category distribution of the dataset are provided.
Method
Ego Status
L2 (m) ↓
Collision (%) ↓
Intersection (%) ↓
BEV
Planner
1s
2s
3s
Avg.
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Traditional / modular paradigm
ST-P3 Hu et al. (2022)
–
–
1.59
2.64
3.73
2.65
0.69
3.62
8.39
4.23
2.53
8.17
14.40
8.37
UniAD Hu et al. (2023)
✓
✓
0.20
0.42
0.75
0.46
0.02
0.25
0.84
0.37
0.20
1.33
3.24
1.59
VAD-Base Jiang et al. (2023)
✓
✓
0.17
0.34
0.60
0.37
0.04
0.27
0.67
0.33
0.21
2.13
5.06
2.47
AD-MLP Li et al. (2024c)
–
✓
0.15
0.32
0.59
0.35
0.00
0.27
0.85
0.37
0.27
2.52
6.60
2.93
Table 1 : Planning results on nuScenes valset. Best and second-best results among VLM/VLA-based methods are highlighted. The result follows the evaluation protocol of OmniDrive Wang et al. (2025b) .
Method
Base Model
Sensor
Closed-loop Metrics ( ↑ )
NC
DAC
EP
TTC
Comf.
PDMS
Traditional / modular paradigm
TransFuser Chitta et al. (2022)
-
Image+LiDAR
97.8
92.6
78.9
92.0
99.9
83.8
PARA-Drive Weng et al. (2024)
-
Image
97.9
92.4
79.3
93.0
99.8
84.0
Hydra-MDP Li et al. (2024a)
-
Image+LiDAR
98.3
96.0
78.7
94.6
100.0
86.5
DiffusionDrive Liao et al. (2025)
-
Image+LiDAR
98.2
96.2
82.2
94.7
100.0
88.1
Table 2 : Planning results on NAVSIM navtest split . The result follows the evaluation protocol of official NAVSIMv1 Dauner et al. (2024) with non-reactive simulation.
Table 6
Method
NAVSIM
nuScenes
NC
DAC
TTC
EP
PDMS
L2
CR
Intersection
Baseline
98.0
94.0
94.4
80.4
85.9
0.16 / 0.32 / 0.58
0.04 / 0.11 / 0.52
0.64 / 2.46 / 5.31
Global
97.9
94.1
94.2
80.2
85.8
0.16 / 0.31 / 0.57
0.02 / 0.17 / 0.56
0.48 / 2.34 / 4.90
GeoCoT
98.4
95.6
95.0
81.6
87.6
0.14 / 0.29 / 0.53
0.00 / 0.07 / 0.27
0.37 / 1.44 / 4.25
Table 5 : Ablation study on geometry integration strategies. We compare the vanilla baseline, global geometry fusion, and the proposed GeoCoT.
Figure 4: Visualization of geometry token response with different fusion strategies. (A) Global geometry tokens interleaved (B) GeoCoT retrieved geometric tokens interleaved.
Method
L2
CR
Intersection
GeoCoTDrive w/ Perception Label
0.17 / 0.30 / 0.55
0.02 / 0.12 / 0.39
0.50 / 2.20 / 4.98
GeoCoTDrive w/ PlanningGrounding
0.14 / 0.29 / 0.53
0.00 / 0.07 / 0.27
0.37 / 1.44 / 4.25
Table 6 : Comparison of different grounding annotation sources for GeoCoTDrive. “Perception label” uses the human-annotated perception boxes.
Method
Geometric Model
NC
DAC
TTC
EP
PDMS
Baseline
–
98.0
94.0
94.4
80.4
85.9
GeoCoTDrive
VGGT
98.3
95.3
94.6
81.7
87.2
GeoCoTDrive
DA3-LARGE
98.4
95.6
95.0
81.6
87.6
Table 7 : Effect of geometric foundation model. Models are loaded from official checkpoints.
Figure 5 : Effect of grounding quality on downstream driving behavior. Lower values indicate better safety performance. denotes GeoCoTDrive with perturbed grounding quality, and denotes OmniDrive.
Figure 6: Qualitative planning results of GeoCoTDrive on nuScenes, NAVSIM and Bench2Drive.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Main supervision
Grounding granularity
NuScenes-QA [ 46 ]
Template-based VQA from structured scene graphs.
Object-level attributes and relations derived from 3D detection annotations.
DriveLM [ 49 ]
Graph Visual Question Answering for perception, prediction, planning, behavior, and motion.
Object-level graph nodes and interaction-level reasoning.
LingoQA [ 44 ]
Video-based driving QA with natural-language answers and explanations.
Mainly language-level supervision without explicit box-level grounding.
OmniDrive [ 56 ]
Counterfactual QA using simulated trajectories, expert trajectories, 3D objects, and map elements.
3D object/map-centric detection and trajectory-level consequence reasoning.
PlanningGrounding
Command-conditioned grounding of planning-relevant visual regions.
Table 8: Comparison with existing driving VQA datasets on grounding supervision.
Source Dataset
# Samples
Category Distribution (%)
Critical Object
Road Boundary
Conflict
Occluded Unknown
Dense Object
nuScenes
28K
30.68
43.65
7.05
6.14
12.45
NAVSIM
102K
34.47
34.05
6.86
3.99
20.60
Bench2Drive
16K
38.31
47.16
2.99
3.48
8.05
Total
146K
34.16
37.33
6.47
4.35
17.66
Appendix
Table 9: Dataset composition of PlanningGrounding. The table summarizes the number of grounding QA samples from each source dataset and the distribution of five planning-oriented grounding categories. “K” denotes thousands of samples.
Stage
Trainable Modules
Objective
Epochs
Peak LR
Stage-I
Vision + LLM
Full QA training
3
4×10−5
Stage-II
LLM + Geometric Aligner
GeoCoT
3
2×10−5
Appendix
Table 10 : Two-stage training structure of GeoCoTDrive. Stage-I adapts the vision-language backbone to autonomous-driving QA and planning knowledge, while Stage-II trains GeoCoT planning.
Grid Size
Geo. Tokens
NC
DAC
TTC
PDMS
1×1
1
98.2
94.6
94.5
86.4
2×2
4
98.3
95.1
94.6
87.0
4×4
16
98.5
95.4
95.0
87.3
5×5
25
98.5
95.4
95.0
87.3
Appendix
Table 11 : Ablation on the sampling grid size for localized geometry retrieval. The grid size denotes the number of sampled geometric features within each grounded region.
Stage-I QA
Stage-II GeoCoT
NC
DAC
TTC
PDMS
✓
×
98.0
94.0
94.4
85.9
×
✓
98.3
94.7
94.5
86.6
✓
✓
98.5
95.4
95.0
87.3
Appendix
Table 12 : Ablation on training stages. Stage-I denotes full QA training, and Stage-II denotes geometry-interleaved planning training.
Scenario
Method
PDMS
NC
DAC
EP
Left turn
Baseline
83.96
98.34
91.04
78.63
GeoCoTDrive
86.66 (+2.70)
99.04 (+0.70)
93.80 (+2.76)
80.78 (+2.15)
Right turn
Baseline
80.10
97.30
89.84
73.23
GeoCoTDrive
82.32 (+2.22)
97.61 (+0.31)
91.80 (+1.96)
74.35 (+1.12)
Appendix
Table 13 : Performance on turning scenarios in NAVSIM. Values in parentheses indicate absolute gains over the baseline.
Figure 7: Qualitative planning results on nuScenes and Bench2Drive. GeoCoTDrive is compared with the corresponding baseline methods.
Figure 8: PlanningGrounding-nuScenes dataset. Human annotations are shown for comparison.
Figure 9: PlanningGrounding-NAVSIM and PlanningGrounding-Bench2Drive dataset examples.
School of Electronic Information Engineering, Beihang University · Institute for AI Industry Research (AIR), Tsinghua University · National College for Excellent Engineers, Beihang University +5