doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving
Authors: Parthib Roy, Yash Tandon, Marcus Blennemann, Giovanni Tapia Lopez, Angel Martinez-Sanchez, Mohan M. Trivedi, Ross Greer
Organizations: Machine Intelligence, Interaction, and Imagination (Mi3) Laboratory, University of California, Merced, Merced, CA, USA. · Laboratory for Intelligent & Safe Automobiles (LISA), University of California, San Diego, La Jolla, CA, USA.
Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at https://github.com/Mi3-Lab/doPlan. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.
Figures & tables
Fig. 1: Representative doPlan annotation illustrating temporally extended, multi-stage passenger intent. A single human-written instruction is aligned with successive route segments, with front-camera frames showing the corresponding driving context.
Property
doScenes
doPlan
Passenger instructions
2,450
5,154
Window duration
12 s
30.0–508.8 s
Cumulative instruction context
∼ 8.17 h
169.1 h
Unique annotated context
∼ 2.4 h
50.9 h
Temporal horizon
Fixed
Variable
TABLE I: Comparison of doScenes and doPlan.
Fig. 2: Distribution of instruction-bearing window durations ( n=5,154 ). The median is 80.4 s; 64.6% of windows exceed 60 s and 32.3% exceed 120 s.
Annotation Window (s)
Mean Words
Multi-Sentence
Multi-Step
30.0–45.2 ( n=1,031 )
9.8
24.2%
28.8%
45.2–65.0 ( n=1,031 )
11.2
27.8%
35.1%
65.2–97.4 ( n=1,031 )
13.6
33.3%
48.1%
97.4–169.8 ( n=1,031 )
17.7
46.4%
53.2%
169.8–508.8 ( n=1,030 )
27.8
52.3%
62.6%
TABLE II: Instruction structure across annotation-window duration groups in doPlan.
Label
Description
Count
%
Non-Referential
No specific scene reference.
782
15.17
Static
Fixed scene element.
2,749
53.34
Dynamic
Moving agent.
629
12.20
Both
Static and dynamic elements.
954
18.51
Ambiguous
Reference type is unclear.
39
0.76
Other/Malformed
Other or malformed label.
1
0.02
TABLE III: Referential labels and their distribution in doPlan.
Video observed
R@1
R@5
R@10
5 s
13.9%
38.2%
52.2%
15 s
14.9%
40.4%
55.5%
30 s
16.5%
43.3%
57.5%
60 s
17.8%
44.7%
59.9%
Full window
23.1%
54.2%
68.6%
TABLE IV: Retrieval of the correct passenger instruction from 100 candidates as more of the driving sequence is observed.
Model
ADE ↓ (m)
No Instr.
Correct
Unrelated
Counterfactual
OpenEMMA
3.342
3.307
3.359
3.340
AutoVLA
2.663
2.784
2.846
2.872
Alpamayo 1.5 ( w=1.0 )
1.979
2.100
2.106
2.081
Alpamayo 2.0 ( w=3.0 )
1.451
1.834
1.816
1.858
TABLE V: Trajectory alignment across language conditions on 4,315 instructions. ADE (average displacement error) is measured over a common 5 s horizon.
w
1.0
1.5
2.0
2.5
3.0
3.5
4.0
S (m)
0.219
0.297
0.375
0.455
0.533
0.607
0.679
TABLE VI: Effect of Alpamayo 2.0 navigation-guidance weight on trajectory separation.
Fig. 3: Increasing navigation guidance makes Alpamayo 2.0 more responsive to passenger language. For the same driving scene, replacing the original “right lane” instruction (blue) with its left–right counterfactual, “left lane” (red), yields a trajectory separation of S=0.95 m at w=1.0 and S=3.84 m at w=4.0 . At higher guidance, the counterfactual trajectory shifts toward the left lane while the original trajectory bends more sharply to the right.
Time to turn
OpenEMMA
Alp. 1.5
AutoVLA
Alp. 2.0
≤5 s
-0.001
-0.022
+0.818
+0.006
5–10 s
+0.003
+0.037
+0.546
-0.080
10–30 s
-0.004
-0.001
+0.282
+0.051
30–60 s
+0.000
+0.045
+0.267
+0.025
>60 s
+0.003
-0.008
+0.199
+0.057
TABLE VII: Mean directional response R versus time to the recorded turn. Positive values indicate a shift toward the requested direction.