cs.AISep 28, 2026

STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts

Authors: Wanchun Ni, Tao Qi, Leonel Aguilar, Jiugeng Sun, Marlene Wagner, Verena Zimmermann, Mennatallah El-Assady

Organizations: ETH Zurich · Beijing University of Posts and Telecommunications

Abstract

Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks are not transferable to this domain. These challenges make scalable, reliable evaluation difficult. We introduce STRIDE, the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. STRIDE addresses these challenges through three design choices. First, we derive our VRDST evaluation protocol from sociological theories to define a complete evaluation space. Second, it decomposes high-level context into scenario-adaptive behavioral questions. Third, every question is resolved against a deterministic measurement tool library that yields reproducible answers. Together, STRIDE enables complete, verifiable, automated, and scalable evaluation across diverse contexts without requiring human trajectory data. We instantiate STRIDE in the crowd domain as STRIDE-Bench, comprising 1K scenarios, 6K behavioral questions, and 11K measurements with calibrated expected answers across 30 real-world maps. Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment, achieving 80% human agreement. We further evaluate several text-to-trajectory models, finding limited context-alignment capability and persistent challenges in fine-grained context conditioning. We believe that the STRIDE framework provides a first step toward principled evaluation of context-aligned pedestrian trajectory generation.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TrajPrism: A Multi-Task Benchmark for Language-Grounded Urban Trajectory Understanding

    May 11, 2026Lihuan Li, Wilson Wongso, Baiyu Chen +6Urban MobilityJourney

  2. CityTrajBench: A Unified Benchmark for City-Scale Vehicle Trajectory Generation

    Jun 1, 2026Shibo Zhu, Xiaodan Shi, Dayin Chen +4Urban MobilityStudent-Generated Trajectories

  3. TrajGenAgent: A Hierarchical LLM Agent for Human Mobility Trajectory Generation

    Jun 10, 2026Siyu Li, Toan Tran, Lingyi Zhao +2MobilityTraccia