Organizations: Shanghai Jiao Tong University · National University of Singapore · Meituan · The Chinese University of Hong Kong · University of Oxford · Shanghai University
Multimodal large language models (MLLMs) can interpret a street view, but reliable urban action depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a real-scale city. We propose UrbanGround, an urban sandbox built from Hong Kong's territory-wide 3D geospatial data. It combines the city's geographic structure with continuous, collision-constrained control through a shared evaluation interface. Agents use first-person observations and an interactive map to select actions across tasks ranging from local question answering to long-horizon navigation. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can gather and interpret local visual evidence to answer spatial questions. Then we ask whether these abilities support navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far MLLM agents can explore reliably in open-ended urban environments.
Figures & tables
Figure 2: Dynamic simulation components of U r b a n Ground . The same urban environment can be rendered under different times of day and weather conditions, while a georeferenced pedestrian network supports animated pedestrian populations.
Figure 3: The spatial agency evaluation ladder increases the state that must remain usable across action. Level 1 supports RQ1. Levels 2–4 support RQ2. Level 5 and matched visual interventions support RQ3.
Figure 3
Model
Visual Recognition
Orientation
Active Exploration
Overall
Acc.
PNA
Acc.
PNA
Acc.
PNA
Acc.
PNA
GPT-5.6-Luna
78.8
64.7
48.3
71.6
58.8
92.2
63.2
76.6
GPT-5.5
82.5
63.8
40.0
74.8
62.5
91.4
63.6
76.8
GPT-5.4
75.0
66.2
31.7
75.0
57.5
91.1
56.8
77.7
GPT-5.2
77.5
68.0
33.3
74.7
48.8
90.9
55.0
78.2
Claude-Opus-5
91.3
69.4
58.3
77.9
82.5
92.8
79.1
80.2
Table 1: Answer accuracy and pedestrian-network adherence compare recognition, orientation, and active exploration question answering across the evaluated MLLM agents. Overall is the instance-count-weighted average across the three task types.
Figure 6: Example of active exploration by GPT-5.5. The agent is asked to answer which bank is next to Beijing Tong Ren Tang.
Model
Provided
Inferred
Multiple
Overall
SN
LN
IN
CN
PS
II
TW
MS
GPT-5.6-Luna
81.3
5.0
16.0
3.3
26.7
10.0
1.7
1.7
21.3
GPT-5.5
75.0
0.0
20.0
0.0
35.0
11.7
3.3
0.0
20.8
GPT-5.4
15.0
1.3
20.0
0.0
15.0
11.7
1.7
0.0
8.3
GPT-5.2
25.0
0.0
28.0
0.0
8.3
15.0
1.7
3.3
10.6
Claude-Opus-5
48.8
2.5
30.0
3.3
30.0
11.7
3.3
3.3
17.9
Table 2: Navigation success across tasks grouped by whether the destination is provided, inferred, or multiple, where SN, LN, IN, CN, PS, II, TW, and MS denote short navigation, long navigation, instructional navigation, constrained navigation, place search, intent inference, time-window navigation, and multi-stop navigation, respectively, and Overall is the instance-count-weighted average across all eight subtasks.
Figure 7: Navigation success as a function of interaction horizon. The panels report success across four equal-count bins for (a) ShortNav and (b) InstructionNav.
Figure 8: Long-range navigation progress and multi-stop completion. The left panel reports the proportion of LongNav episodes that end closer to the goal than at initialization. The right panel reports the mean proportion of required destinations reached in multi-stop planning. Error bars show 95% confidence intervals.
Figure 9: Example of GPT-5.5 traversing a complex interchange using the official footbridge and continuing steadily toward the goal.
Figure 10: Example of GPT-5.5 crossing the road toward the goal direction before becoming blocked by a central obstacle and being unable to continue.
Model
Local QA Accuracy
Short Navigation Success
Clear
Dusk
Night
Cloudy
Rain
Clear
Dusk
Night
Cloudy
Rain
GPT-5.6-Luna
63.2
56.4
60.9
62.3
57.7
81.3
73.8
71.3
77.5
76.3
GPT-5.5
63.6
59.1
61.4
62.7
61.4
75.0
72.5
75.0
76.3
72.5
GPT-5.4
56.8
44.5
48.6
58.2
53.6
15.0
13.8
18.8
11.3
16.3
GPT-5.2
55.0
44.1
48.6
55.9
53.6
25.0
26.3
17.5
27.5
21.3
Claude-Opus-5
79.1
66.4
69.5
75.5
70.5
48.8
46.3
41.3
51.3
46.3
Table 3: Local question-answering accuracy and short-navigation success compare robustness across different daytime and weather conditions.
Model
Road Closure
Pedestrians
SR ↑
PNA ↑
CCR ↑
SR ↑
PNA ↑
PCR ↓
GPT-5.6-Luna
0.0
98.4
20.0
0.0
98.6
82.5
GPT-5.5
0.0
94.7
13.3
0.0
93.8
83.8
GPT-5.4
0.0
96.4
16.7
0.0
98.5
80.0
GPT-5.2
0.0
94.2
16.7
1.3
95.4
78.8
Claude-Opus-5
0.0
93.6
46.7
2.5
94.4
87.5
Table 4: Dynamic-environment results, where SR is the goal-reaching rate, PNA is the fraction of action time spent on the pedestrian network, CCR is the closure compliance rate that represents the fraction of road-closure episodes that respect the closure, and PCR is the pedestrian-collision rate.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Environment / benchmark
Real-city geography
Continuous motion
Collision response
3D Pedestrian network
Locomotion primitives
V-IRL ( Yang et al., 2024 )
✓
✗
✗
✗
✗
CitySeeker ( Wang et al., 2026 )
✓
✗
✗
✗
✗
CityNav ( Dalal et al., 2026 )
✓
✗
✗
✗
✗
UrbanWorld ( Shang et al., 2024 )
✓
✓
✗
✗
✗
MetaUrban ( Wu et al., 2025 )
✗
✓
✓
✗
✗
SimWorld ( Ren et al., 2025 )
✗
✓
✓
✗
✗
Appendix
Table 5: Comparison of urban environments and navigation benchmarks in Section 2 . ✓ documented support, ✗ absent from the described setting,
Figure 11: GPT-5.5 LongNav runs grouped by their run-ending trajectory pattern. Substantial progress denotes a reduction of at least 20% in horizontal goal distance. The final recorded checkpoint defines the endpoint of each run.
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.
Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent vision-language models (VLMs) have shown promising performance on observation-conditioned spatial perception and reasoning tasks, it remains unclear whether they can build coherent spatial understanding, act upon it, and refine their actions through multi-turn feedback. To study this problem, we introduce \textbf{SpatialAct}, a simulator-grounded benchmark for probing \textit{action-conditioned spatial reasoning} in 3D scenes. Starting from the most challenging setting, Multi-turn Interactive Refinement, we further design its decomposed counterpart, Single-step Error Detection and Fix, together with five fundamental spatial ability tasks to diagnose the underlying causes of model failures. Experiments reveal a clear reasoning-to-action gap: current VLMs can perform well on isolated spatial reasoning tasks, but struggle to maintain coherent spatial beliefs and produce reliable actions during multi-turn feedback, substantially underperforming humans. These results suggest that current VLM agents still lack robust spatial state tracking under action-induced environment changes, even when low-level control is abstracted away.
Tianhui Liu, Jie Feng, Zhiheng Zheng +6
1The Hong Kong University of Science and Technology (Guangzhou) · 2Zhongguancun Academy · 3Tsinghua University +1
Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predominantly rely on passive evaluation (e.g., static VQA) or simulator-specific pipelines, failing to assess general interactive spatial understanding. We introduce SpatialWorld, a unified benchmark designed specifically for evaluating the interactive spatial understanding of multimodal agents in complex real-world tasks. Integrating eight heterogeneous simulation backends under a shared, simulator-agnostic protocol, SpatialWorld features 760 human-annotated tasks across diverse domains (e.g., household routines, travel, social collaboration). Agents must solve tasks under vision-only partial observability, actively gathering egocentric visual evidence and expressing decisions via a unified, text-based action interface native to MLLMs. For reliable evaluation, each task includes a human-validated initial state, a reference trajectory, and a terminal-state verifier. Evaluating 15 advanced agents reveals that robust spatial task solving remains challenging: the strongest model, GPT-5, achieves an average task success rate (TSR) of only 17.4%, while the leading open-source model, Qwen-3.5, reaches 14.1%. Further analysis exposes a clear mismatch between task success and execution efficiency, alongside substantial domain-specific performance variations. These bottlenecks in active exploration and long-horizon planning position SpatialWorld as a rigorous testbed for future spatial agents.
Hongcheng Gao, Hailong Qu, Jingyi Tang +18
Tsinghua University · Chongqing University · Peking University +6