cs.CVAug 27, 2026

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

Authors: Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, +10 more

Organizations: Shanghai Jiao Tong University · National University of Singapore · Meituan · The Chinese University of Hong Kong · University of Oxford · Shanghai University

Abstract

Multimodal large language models (MLLMs) can interpret a street view, but reliable urban action depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a real-scale city. We propose UrbanGround, an urban sandbox built from Hong Kong's territory-wide 3D geospatial data. It combines the city's geographic structure with continuous, collision-constrained control through a shared evaluation interface. Agents use first-person observations and an interactive map to select actions across tasks ranging from local question answering to long-horizon navigation. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can gather and interpret local visual evidence to answer spatial questions. Then we ask whether these abilities support navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far MLLM agents can explore reliably in open-ended urban environments.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

    Aug 9, 2026Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa +2Urban EnvironmentsObject Goal Navigation

  2. SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

    May 29, 2026Tianhui Liu, Jie Feng, Zhiheng Zheng +6Spatial ReasoningRecent Vision-Language Models

  3. SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

    Jun 8, 2026Hongcheng Gao, Hailong Qu, Jingyi Tang +18Spatial ReasoningMultimodal Agents