cs.ROSep 29, 2026

doPlan: A Variable-Horizon Dataset for Multi-Stage Language-Conditioned Planning in Autonomous Driving

Authors: Parthib Roy, Yash Tandon, Marcus Blennemann, Giovanni Tapia Lopez, Angel Martinez-Sanchez, Mohan M. Trivedi, Ross Greer

Organizations: Machine Intelligence, Interaction, and Imagination (Mi3) Laboratory, University of California, Merced, Merced, CA, USA. · Laboratory for Intelligent & Safe Automobiles (LISA), University of California, San Diego, La Jolla, CA, USA.

Abstract

Autonomous vehicles interacting with passengers through natural language must reason beyond immediate commands. Passenger intent may span multiple stages of behavior, depend on future events, refer to surrounding agents or landmarks, and remain relevant as driving conditions evolve. Existing language-enabled driving datasets largely focus on short, localized interactions, leaving these longer-horizon forms of passenger intent comparatively underexplored. We introduce doPlan, to our knowledge the first publicly available, human-annotated real-world dataset designed to study passenger language as persistent task context. Built on nuPlan, doPlan contains 5,154 human-written passenger instructions spanning 169.1 hours of cumulative instruction-aligned context over 50.9 hours of unique driving, with annotation windows ranging from 30.0 to 508.8 s. The annotations capture immediate, deferred, event-conditioned, persistent, and multi-stage passenger intent. The dataset, annotation interface, and supporting resources are publicly available at https://github.com/Mi3-Lab/doPlan. We evaluate four language-conditioned driving models and find that sensitivity to passenger language does not reliably translate into behavior consistent with the requested direction. More broadly, among 2,161 examples with a matched future maneuver, the first associated maneuver occurs a median of 24.6 s after the evaluation point, and only 9.8% occur within the models' common 5 s prediction horizon. These findings highlight the need to connect persistent passenger intent with successive planning decisions. doPlan provides a setting for studying how unresolved goals can be retained, grounded in evolving scenes, and tracked across multiple stages, including how a planner determines when a future goal becomes relevant to the current plan.

Figures & tables

Explore similar work

CardsList
  1. DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions

    May 29, 2026Weicheng Zheng, Yixin Huang, Qiao Sun +2Diffusion-Based Vision-Language-ActionsDrives

  2. nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving

    May 29, 2026Zhiyu Huang, Johnson Liu, Rui Song +13Naturalistic Driving DataAutonomous Driving

  3. VeriDrive: Verifiable Counterfactual Supervision for Cost-Efficient Vision-Language Planning

    Jun 5, 2026Zikai Zhang, Hubert P. H. Shum, Toby P. BreckonPhotometric SupervisionHuman Driving