doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving
Authors: Parthib Roy, Yash Tandon, Marcus Blennemann, Giovanni Tapia Lopez, Angel Martinez-Sanchez, Mohan M. Trivedi, Ross Greer
Organizations: Machine Intelligence, Interaction, and Imagination (Mi3) Laboratory, University of California, Merced, Merced, CA, USA. · Laboratory for Intelligent & Safe Automobiles (LISA), University of California, San Diego, La Jolla, CA, USA.
Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at https://github.com/Mi3-Lab/doPlan. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.
Figures & tables
Fig. 1: Representative doPlan annotation illustrating temporally extended, multi-stage passenger intent. A single human-written instruction is aligned with successive route segments, with front-camera frames showing the corresponding driving context.
Property
doScenes
doPlan
Passenger instructions
2,450
5,154
Window duration
12 s
30.0–508.8 s
Cumulative instruction context
∼ 8.17 h
169.1 h
Unique annotated context
∼ 2.4 h
50.9 h
Temporal horizon
Fixed
Variable
TABLE I: Comparison of doScenes and doPlan.
Fig. 2: Distribution of instruction-bearing window durations ( n=5,154 ). The median is 80.4 s; 64.6% of windows exceed 60 s and 32.3% exceed 120 s.
Annotation Window (s)
Mean Words
Multi-Sentence
Multi-Step
30.0–45.2 ( n=1,031 )
9.8
24.2%
28.8%
45.2–65.0 ( n=1,031 )
11.2
27.8%
35.1%
65.2–97.4 ( n=1,031 )
13.6
33.3%
48.1%
97.4–169.8 ( n=1,031 )
17.7
46.4%
53.2%
169.8–508.8 ( n=1,030 )
27.8
52.3%
62.6%
TABLE II: Instruction structure across annotation-window duration groups in doPlan.
Label
Description
Count
%
Non-Referential
No specific scene reference.
782
15.17
Static
Fixed scene element.
2,749
53.34
Dynamic
Moving agent.
629
12.20
Both
Static and dynamic elements.
954
18.51
Ambiguous
Reference type is unclear.
39
0.76
Other/Malformed
Other or malformed label.
1
0.02
TABLE III: Referential labels and their distribution in doPlan.
Video observed
R@1
R@5
R@10
5 s
13.9%
38.2%
52.2%
15 s
14.9%
40.4%
55.5%
30 s
16.5%
43.3%
57.5%
60 s
17.8%
44.7%
59.9%
Full window
23.1%
54.2%
68.6%
TABLE IV: Retrieval of the correct passenger instruction from 100 candidates as more of the driving sequence is observed.
Model
ADE ↓ (m)
No Instr.
Correct
Unrelated
Counterfactual
OpenEMMA
3.342
3.307
3.359
3.340
AutoVLA
2.663
2.784
2.846
2.872
Alpamayo 1.5 ( w=1.0 )
1.979
2.100
2.106
2.081
Alpamayo 2.0 ( w=3.0 )
1.451
1.834
1.816
1.858
TABLE V: Trajectory alignment across language conditions on 4,315 instructions. ADE (average displacement error) is measured over a common 5 s horizon.
w
1.0
1.5
2.0
2.5
3.0
3.5
4.0
S (m)
0.219
0.297
0.375
0.455
0.533
0.607
0.679
TABLE VI: Effect of Alpamayo 2.0 navigation-guidance weight on trajectory separation.
Fig. 3: Increasing navigation guidance makes Alpamayo 2.0 more responsive to passenger language. For the same driving scene, replacing the original “right lane” instruction (blue) with its left–right counterfactual, “left lane” (red), yields a trajectory separation of S=0.95 m at w=1.0 and S=3.84 m at w=4.0 . At higher guidance, the counterfactual trajectory shifts toward the left lane while the original trajectory bends more sharply to the right.
Time to turn
OpenEMMA
Alp. 1.5
AutoVLA
Alp. 2.0
≤5 s
-0.001
-0.022
+0.818
+0.006
5–10 s
+0.003
+0.037
+0.546
-0.080
10–30 s
-0.004
-0.001
+0.282
+0.051
30–60 s
+0.000
+0.045
+0.267
+0.025
>60 s
+0.003
-0.008
+0.199
+0.057
TABLE VII: Mean directional response R versus time to the recorded turn. Positive values indicate a shift toward the requested direction.
Driving Vision-Language-Action Models (Driving VLAs) aim to use language to improve end-to-end planning, but the language-action gap limits this promise. We propose DriveMA, a Driving VLA framework built on verifiable meta-actions, which summarize future ego motion into compact language-domain intentions and can be constructed from expert trajectories with a trajectory-grounded annotation pipeline and can be verified against generated trajectories through rule-based projection. DriveMA exploits this verifiability with action-centric supervised training and a data-efficient turn-level credit assignment reinforcement learning framework, explicitly aligning high-level decisions with low-level trajectory planning through dense rewards and precise credit assignment. DriveMA sets a new state of the art on the Waymo Open Dataset Vision-based E2E Driving, achieving a Rater Feedback Score of 8.060 with a 2B model and further improving it to 8.079 with a 4B model; it also obtains competitive closed-loop planning performance on NAVSIM. These results show that even a simple meta-action interface can achieve state-of-the-art planning when made verifiable and optimized for language-action alignment. Code, data, and models are available at https://tsinghua-mars-lab.github.io/DriveMA.
Weicheng Zheng, Yixin Huang, Qiao Sun +2
1Shanghai Qi Zhi Institute · 3Tongji University · 2IIIS Tsinghua University
Reasoning is essential for autonomous driving (AD) in long-tail scenarios, where vehicles must apply commonsense knowledge, understand spatial relations, infer agent interactions, and make safe decisions. However, existing AD datasets and benchmarks mainly target perception, prediction, or planning, and provide limited supervision for reasoning over realistic long-tail driving scenes. We introduce nuReasoning, a large-scale real-world dataset and benchmark for reasoning-centric AD. Following the lineage of nuScenes and nuPlan, nuReasoning advances real-world AD datasets and benchmarks toward reasoning in long-tail driving scenarios. The dataset contains 20,000 clips, each 20 seconds long, collected across multiple cities, with synchronized multi-camera images, LiDAR data, HD maps, object annotations, and human-verified reasoning annotations spanning Spatial Reasoning, Decision Reasoning, and Counterfactual Reasoning. Unlike prior datasets that focus primarily on visual question answering, nuReasoning supports both reasoning evaluation and planning evaluation, enabling a direct study of how reasoning supervision affects driving performance. Experiments show that fine-tuning VLMs on nuReasoning substantially improves driving-specific question answering, while incorporating reasoning supervision into VLA training improves planning performance even when textual reasoning outputs are disabled at inference time. These results establish nuReasoning as a foundation for evaluating and improving robust, interpretable, reasoning-driven AD systems in realistic long-tail settings.
Zhiyu Huang, Johnson Liu, Rui Song +13
Xuewei · University of California, Los Angeles · 2Motional
Vision-language driving models increasingly use reasoning supervision to bridge perception, prediction, and planning, but existing driving rationales are often free-form and expensive to generate with frontier models. We present VeriDrive, a framework for constructing planning-oriented, verifiable counterfactual supervision. VeriDrive converts driving reasoning into a structured Perception-Evaluation-Revision chain that grounds key objects in future motion, evaluates alternative ego trajectories with rule-checkable evidence, revises risky intent toward expert behavior, and produces final planning targets. To scale data construction, VeriDrive combines local generation with validator-guided selective correction, escalating only invalid or difficult samples. We build the VeriDrive dataset on nuScenes and train under the Omni-Q protocol. Controlled open-loop experiments show that VeriDrive improves L2, Collision, and Intersection over OmniDrive while reducing logged token usage, generation time, and actual paid LLM/VLM cost. These results show that auditable intermediate fields and structured revision targets can improve vision-language planning supervision under realistic annotation budgets. Code, prompts, and validator scripts are coming soon and will be released after the review process.
Zikai Zhang, Hubert P. H. Shum, Toby P. Breckon
Department of Computer Science Durham University Stockton Road, Durham, DH1 3LE, UK