Organizations: Shanghai Jiao Tong University · National University of Singapore · Meituan · The Chinese University of Hong Kong · University of Oxford · Shanghai University
Multimodal large language models (MLLMs) can interpret a street view, but reliable urban action depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a real-scale city. We propose UrbanGround, an urban sandbox built from Hong Kong's territory-wide 3D geospatial data. It combines the city's geographic structure with continuous, collision-constrained control through a shared evaluation interface. Agents use first-person observations and an interactive map to select actions across tasks ranging from local question answering to long-horizon navigation. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can gather and interpret local visual evidence to answer spatial questions. Then we ask whether these abilities support navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far MLLM agents can explore reliably in open-ended urban environments.
Figures & tables
Figure 2: Dynamic simulation components of U r b a n Ground . The same urban environment can be rendered under different times of day and weather conditions, while a georeferenced pedestrian network supports animated pedestrian populations.
Figure 3: The spatial agency evaluation ladder increases the state that must remain usable across action. Level 1 supports RQ1. Levels 2–4 support RQ2. Level 5 and matched visual interventions support RQ3.
Figure 3
Model
Visual Recognition
Orientation
Active Exploration
Overall
Acc.
PNA
Acc.
PNA
Acc.
PNA
Acc.
PNA
GPT-5.6-Luna
78.8
64.7
48.3
71.6
58.8
92.2
63.2
76.6
GPT-5.5
82.5
63.8
40.0
74.8
62.5
91.4
63.6
76.8
GPT-5.4
75.0
66.2
31.7
75.0
57.5
91.1
56.8
77.7
GPT-5.2
77.5
68.0
33.3
74.7
48.8
90.9
55.0
78.2
Claude-Opus-5
91.3
69.4
58.3
77.9
82.5
92.8
79.1
80.2
Table 1: Answer accuracy and pedestrian-network adherence compare recognition, orientation, and active exploration question answering across the evaluated MLLM agents. Overall is the instance-count-weighted average across the three task types.
Figure 6: Example of active exploration by GPT-5.5. The agent is asked to answer which bank is next to Beijing Tong Ren Tang.
Model
Provided
Inferred
Multiple
Overall
SN
LN
IN
CN
PS
II
TW
MS
GPT-5.6-Luna
81.3
5.0
16.0
3.3
26.7
10.0
1.7
1.7
21.3
GPT-5.5
75.0
0.0
20.0
0.0
35.0
11.7
3.3
0.0
20.8
GPT-5.4
15.0
1.3
20.0
0.0
15.0
11.7
1.7
0.0
8.3
GPT-5.2
25.0
0.0
28.0
0.0
8.3
15.0
1.7
3.3
10.6
Claude-Opus-5
48.8
2.5
30.0
3.3
30.0
11.7
3.3
3.3
17.9
Table 2: Navigation success across tasks grouped by whether the destination is provided, inferred, or multiple, where SN, LN, IN, CN, PS, II, TW, and MS denote short navigation, long navigation, instructional navigation, constrained navigation, place search, intent inference, time-window navigation, and multi-stop navigation, respectively, and Overall is the instance-count-weighted average across all eight subtasks.
Figure 7: Navigation success as a function of interaction horizon. The panels report success across four equal-count bins for (a) ShortNav and (b) InstructionNav.
Figure 8: Long-range navigation progress and multi-stop completion. The left panel reports the proportion of LongNav episodes that end closer to the goal than at initialization. The right panel reports the mean proportion of required destinations reached in multi-stop planning. Error bars show 95% confidence intervals.
Figure 9: Example of GPT-5.5 traversing a complex interchange using the official footbridge and continuing steadily toward the goal.
Figure 10: Example of GPT-5.5 crossing the road toward the goal direction before becoming blocked by a central obstacle and being unable to continue.
Model
Local QA Accuracy
Short Navigation Success
Clear
Dusk
Night
Cloudy
Rain
Clear
Dusk
Night
Cloudy
Rain
GPT-5.6-Luna
63.2
56.4
60.9
62.3
57.7
81.3
73.8
71.3
77.5
76.3
GPT-5.5
63.6
59.1
61.4
62.7
61.4
75.0
72.5
75.0
76.3
72.5
GPT-5.4
56.8
44.5
48.6
58.2
53.6
15.0
13.8
18.8
11.3
16.3
GPT-5.2
55.0
44.1
48.6
55.9
53.6
25.0
26.3
17.5
27.5
21.3
Claude-Opus-5
79.1
66.4
69.5
75.5
70.5
48.8
46.3
41.3
51.3
46.3
Table 3: Local question-answering accuracy and short-navigation success compare robustness across different daytime and weather conditions.
Model
Road Closure
Pedestrians
SR ↑
PNA ↑
CCR ↑
SR ↑
PNA ↑
PCR ↓
GPT-5.6-Luna
0.0
98.4
20.0
0.0
98.6
82.5
GPT-5.5
0.0
94.7
13.3
0.0
93.8
83.8
GPT-5.4
0.0
96.4
16.7
0.0
98.5
80.0
GPT-5.2
0.0
94.2
16.7
1.3
95.4
78.8
Claude-Opus-5
0.0
93.6
46.7
2.5
94.4
87.5
Table 4: Dynamic-environment results, where SR is the goal-reaching rate, PNA is the fraction of action time spent on the pedestrian network, CCR is the closure compliance rate that represents the fraction of road-closure episodes that respect the closure, and PCR is the pedestrian-collision rate.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Environment / benchmark
Real-city geography
Continuous motion
Collision response
3D Pedestrian network
Locomotion primitives
V-IRL ( Yang et al., 2024 )
✓
✗
✗
✗
✗
CitySeeker ( Wang et al., 2026 )
✓
✗
✗
✗
✗
CityNav ( Dalal et al., 2026 )
✓
✗
✗
✗
✗
UrbanWorld ( Shang et al., 2024 )
✓
✓
✗
✗
✗
MetaUrban ( Wu et al., 2025 )
✗
✓
✓
✗
✗
SimWorld ( Ren et al., 2025 )
✗
✓
✓
✗
✗
Appendix
Table 5: Comparison of urban environments and navigation benchmarks in Section 2 . ✓ documented support, ✗ absent from the described setting,
Figure 11: GPT-5.5 LongNav runs grouped by their run-ending trajectory pattern. Substantial progress denotes a reduction of at least 20% in horizontal goal distance. The final recorded checkpoint defines the endpoint of each run.