Vision-language-action~(VLA) models have emerged as a promising paradigm for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a think with 2D first, drive with dedicated 3D priors paradigm. It first grounds 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support the trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.
Figures & tables
Figure 1: Comparison of geometry integration paradigms for VLA . (a) Structure-perception fusion injects agent and map tokens. (b) Geometric fusion incorporates global geometric tokens. (c) GeoCoTDrive introduces an explicit geometric chain-of-thought pipeline: VLM first grounds planning-relevant regions, then retrieves localized 3D geometric priors, and finally interleaves the grounding text with geometry tokens to condition trajectory generation.
Figure 2 : Overview of GeoCoTDrive. The model first grounds planning-critical 2D regions, retrieves localized geometry tokens from a geometric foundation model, and then interleaves them into the autoregressive context for trajectory generation.
Figure 3 : Overview of the PlanningGrounding data construction pipeline. The pipeline generates planning-oriented region annotations from driving scenes by identifying decision-critical visual cues. An annotation example and the overall category distribution of the dataset are provided.
Method
Ego Status
L2 (m) ↓
Collision (%) ↓
Intersection (%) ↓
BEV
Planner
1s
2s
3s
Avg.
1s
2s
3s
Avg.
1s
2s
3s
Avg.
Traditional / modular paradigm
ST-P3 Hu et al. (2022)
–
–
1.59
2.64
3.73
2.65
0.69
3.62
8.39
4.23
2.53
8.17
14.40
8.37
UniAD Hu et al. (2023)
✓
✓
0.20
0.42
0.75
0.46
0.02
0.25
0.84
0.37
0.20
1.33
3.24
1.59
VAD-Base Jiang et al. (2023)
✓
✓
0.17
0.34
0.60
0.37
0.04
0.27
0.67
0.33
0.21
2.13
5.06
2.47
AD-MLP Li et al. (2024c)
–
✓
0.15
0.32
0.59
0.35
0.00
0.27
0.85
0.37
0.27
2.52
6.60
2.93
Table 1 : Planning results on nuScenes valset. Best and second-best results among VLM/VLA-based methods are highlighted. The result follows the evaluation protocol of OmniDrive Wang et al. (2025b) .
Method
Base Model
Sensor
Closed-loop Metrics ( ↑ )
NC
DAC
EP
TTC
Comf.
PDMS
Traditional / modular paradigm
TransFuser Chitta et al. (2022)
-
Image+LiDAR
97.8
92.6
78.9
92.0
99.9
83.8
PARA-Drive Weng et al. (2024)
-
Image
97.9
92.4
79.3
93.0
99.8
84.0
Hydra-MDP Li et al. (2024a)
-
Image+LiDAR
98.3
96.0
78.7
94.6
100.0
86.5
DiffusionDrive Liao et al. (2025)
-
Image+LiDAR
98.2
96.2
82.2
94.7
100.0
88.1
Table 2 : Planning results on NAVSIM navtest split . The result follows the evaluation protocol of official NAVSIMv1 Dauner et al. (2024) with non-reactive simulation.
Table 6
Method
NAVSIM
nuScenes
NC
DAC
TTC
EP
PDMS
L2
CR
Intersection
Baseline
98.0
94.0
94.4
80.4
85.9
0.16 / 0.32 / 0.58
0.04 / 0.11 / 0.52
0.64 / 2.46 / 5.31
Global
97.9
94.1
94.2
80.2
85.8
0.16 / 0.31 / 0.57
0.02 / 0.17 / 0.56
0.48 / 2.34 / 4.90
GeoCoT
98.4
95.6
95.0
81.6
87.6
0.14 / 0.29 / 0.53
0.00 / 0.07 / 0.27
0.37 / 1.44 / 4.25
Table 5 : Ablation study on geometry integration strategies. We compare the vanilla baseline, global geometry fusion, and the proposed GeoCoT.
Figure 4: Visualization of geometry token response with different fusion strategies. (A) Global geometry tokens interleaved (B) GeoCoT retrieved geometric tokens interleaved.
Method
L2
CR
Intersection
GeoCoTDrive w/ Perception Label
0.17 / 0.30 / 0.55
0.02 / 0.12 / 0.39
0.50 / 2.20 / 4.98
GeoCoTDrive w/ PlanningGrounding
0.14 / 0.29 / 0.53
0.00 / 0.07 / 0.27
0.37 / 1.44 / 4.25
Table 6 : Comparison of different grounding annotation sources for GeoCoTDrive. “Perception label” uses the human-annotated perception boxes.
Method
Geometric Model
NC
DAC
TTC
EP
PDMS
Baseline
–
98.0
94.0
94.4
80.4
85.9
GeoCoTDrive
VGGT
98.3
95.3
94.6
81.7
87.2
GeoCoTDrive
DA3-LARGE
98.4
95.6
95.0
81.6
87.6
Table 7 : Effect of geometric foundation model. Models are loaded from official checkpoints.
Figure 5 : Effect of grounding quality on downstream driving behavior. Lower values indicate better safety performance. denotes GeoCoTDrive with perturbed grounding quality, and denotes OmniDrive.
Figure 6: Qualitative planning results of GeoCoTDrive on nuScenes, NAVSIM and Bench2Drive.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Main supervision
Grounding granularity
NuScenes-QA [ 46 ]
Template-based VQA from structured scene graphs.
Object-level attributes and relations derived from 3D detection annotations.
DriveLM [ 49 ]
Graph Visual Question Answering for perception, prediction, planning, behavior, and motion.
Object-level graph nodes and interaction-level reasoning.
LingoQA [ 44 ]
Video-based driving QA with natural-language answers and explanations.
Mainly language-level supervision without explicit box-level grounding.
OmniDrive [ 56 ]
Counterfactual QA using simulated trajectories, expert trajectories, 3D objects, and map elements.
3D object/map-centric detection and trajectory-level consequence reasoning.
PlanningGrounding
Command-conditioned grounding of planning-relevant visual regions.
Table 8: Comparison with existing driving VQA datasets on grounding supervision.
Source Dataset
# Samples
Category Distribution (%)
Critical Object
Road Boundary
Conflict
Occluded Unknown
Dense Object
nuScenes
28K
30.68
43.65
7.05
6.14
12.45
NAVSIM
102K
34.47
34.05
6.86
3.99
20.60
Bench2Drive
16K
38.31
47.16
2.99
3.48
8.05
Total
146K
34.16
37.33
6.47
4.35
17.66
Appendix
Table 9: Dataset composition of PlanningGrounding. The table summarizes the number of grounding QA samples from each source dataset and the distribution of five planning-oriented grounding categories. “K” denotes thousands of samples.
Stage
Trainable Modules
Objective
Epochs
Peak LR
Stage-I
Vision + LLM
Full QA training
3
4×10−5
Stage-II
LLM + Geometric Aligner
GeoCoT
3
2×10−5
Appendix
Table 10 : Two-stage training structure of GeoCoTDrive. Stage-I adapts the vision-language backbone to autonomous-driving QA and planning knowledge, while Stage-II trains GeoCoT planning.
Grid Size
Geo. Tokens
NC
DAC
TTC
PDMS
1×1
1
98.2
94.6
94.5
86.4
2×2
4
98.3
95.1
94.6
87.0
4×4
16
98.5
95.4
95.0
87.3
5×5
25
98.5
95.4
95.0
87.3
Appendix
Table 11 : Ablation on the sampling grid size for localized geometry retrieval. The grid size denotes the number of sampled geometric features within each grounded region.
Stage-I QA
Stage-II GeoCoT
NC
DAC
TTC
PDMS
✓
×
98.0
94.0
94.4
85.9
×
✓
98.3
94.7
94.5
86.6
✓
✓
98.5
95.4
95.0
87.3
Appendix
Table 12 : Ablation on training stages. Stage-I denotes full QA training, and Stage-II denotes geometry-interleaved planning training.
Scenario
Method
PDMS
NC
DAC
EP
Left turn
Baseline
83.96
98.34
91.04
78.63
GeoCoTDrive
86.66 (+2.70)
99.04 (+0.70)
93.80 (+2.76)
80.78 (+2.15)
Right turn
Baseline
80.10
97.30
89.84
73.23
GeoCoTDrive
82.32 (+2.22)
97.61 (+0.31)
91.80 (+1.96)
74.35 (+1.12)
Appendix
Table 13 : Performance on turning scenarios in NAVSIM. Values in parentheses indicate absolute gains over the baseline.
Figure 7: Qualitative planning results on nuScenes and Bench2Drive. GeoCoTDrive is compared with the corresponding baseline methods.
Figure 8: PlanningGrounding-nuScenes dataset. Human annotations are shown for comparison.
Figure 9: PlanningGrounding-NAVSIM and PlanningGrounding-Bench2Drive dataset examples.
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing VLAs' training relies heavily on text-centric visual question answering and chain-of-thought reasoning data, which emphasizes linguistic reasoning rather than action-grounded planning. As a result, the learned representations capture semantic knowledge but lack spatial dependencies crucial for reliable trajectory prediction. We propose DriveTeach-VLA, a framework that explicitly teaches VLAs what to see and where to look. Driving-aware Vision Distillation (DVD) injects driving-specific perceptual priors into the vision encoder, while 2D Trajectory-Guided Prompts (2D-TGP) provide spatial conditioning aligned with feasible driving trajectories. Together, they form a vision-guided learning pipeline: what to see (DVD pretraining) - where to look (TGP-guided SFT) - how to act (TGP-guided GRPO). DriveTeach-VLA achieves the state-of-the-art performance on NAVSIM and nuScenes. Our code is available at: https://github.com/ShivaTeam/DriveTeach-VLA.
Yuguang Yang, Canyu Chen, Zhewen Tan +10
School of Electronic Information Engineering, Beihang University · Institute for AI Industry Research (AIR), Tsinghua University · National College for Excellent Engineers, Beihang University +5
Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D world around them. Existing approaches either inject features from a frozen 3D foundation model without an objective that ensures the policy uses them, or constrain geometry with sparse box and map losses that provide no dense spatial signal. We introduce VLGA, the first vision-language-action model supervised to reconstruct the dense 3D world it drives through. VLGA introduces geometry as a fourth modality alongside vision, language, and action through a dedicated expert supervised by a per-pixel pointmap regression loss against LiDAR. Extensive experiments conducted on challenging nuScenes and Bench2Drive datasets for open-loop and closed-loop evaluations, respectively, show the superiority of VLGA over counterpart VLA methods. In particular, on open-loop nuScenes, VLGA sets a new state of the art among VLA methods without ego status, with the lowest L2 (0.50,m average) and 3-second collision rate (0.18%). On closed-loop Bench2Drive, VLGA attains the state-of-the-art driving score of 79.08, +0.71 over the strongest prior VLA, at comparable efficiency and comfort.
Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verbose natural-language Chain-of-Thought (CoT) is poorly suited to real-time control because it is open-ended, costly to decode, and difficult to optimize as an action-facing representation. We propose XCoT-VLA, which replaces descriptive rationales with compact executable CoT tokens learned from automatically constructed Reason-Action supervision. Logged trajectories provide action evidence, while scene context supplies causal semantics. The predicted XCoT sequence remains in context and conditions fixed trajectory queries through shared multimodal self-attention. Deterministic token-function routing applies the Reason FFN to XCoT tokens and the Control FFN to trajectory queries for flow-matching trajectory generation. We further introduce XCoT Policy Optimization (XCPO) as an optional refinement extension in the same executable token space. XCoT-VLA reduces longitudinal ADE from 1.645 to 1.323 on a general-distribution set and lateral FDE from 1.616 to 0.648 in lane-change scenarios. By representing driving-oriented reasoning with only 2-6 executable XCoT tokens, our method substantially reduces autoregressive reasoning overhead and remains within the real-time planning budget. These results demonstrate that driving-oriented reasoning can be compact, executable, and directly connected to trajectory generation.