DUDA-Bench: Benchmarking LLM Agents on Multimodal Data-Driven Urban Diagnosis
Organizations: The Hong Kong University of Science and Technology (Guangzhou) Guangzhou, China
Abstract
Urban diagnosis integrates heterogeneous observations to identify urban problems, localize affected areas, and investigate contributing factors, informing evidence-based urban planning and management. However, its reliance on labor-intensive, case-specific expert workflows limits scalability and reuse, motivating the exploration of agent-based execution. To evaluate this capability, we introduce DUDA-Bench, a hierarchical and interactive benchmark that formalizes data-driven urban diagnosis as a multi-stage agent workflow. It comprises 86 atomic and 22 workflow tasks spanning four analytical stages, grounded in multimodal data from 12 cities covering five urban problem types. Evaluations of seven backbone models and five agent systems reveal a substantial gap between isolated analytical competence and end-to-end diagnosis, with system benefits varying across backbones. Trajectory analysis shows that unresolved evidence gaps propagate across stages, while successful recovery involves revising assumptions and actions using feedback. These findings highlight limitations in coordinating analytical capabilities across stages, particularly adaptive planning, evidence integration, and verification. More broadly, DUDA-Bench provides a framework for translating expert analytical workflows into hierarchical agent tasks and process-aware evaluation, supporting systematic assessment of end-to-end analytical capabilities.
Figures & tables
| Category | Benchmark | Data Coverage | Agentic Workflow Tasks | Evaluation | ||||||||
| Map Data | Urban Indicators | Remote Sensing | Street View | Temporal Data | Cities | Interactive Execution | Multi-stage Workflow | Planning | Outcome | Process- aware | ||
| Urban-agent benchmarks | CityBench Feng et al. (2025b) | ✓ | ✓ | ✓ | ✓ | ✓ | 13 | ✓ | ✓ | ✗ | ✓ | ✗ |
| CityEQA Zhao et al. (2025) | ✗ | ✗ | ✗ | ✓ | ✗ | – | ✓ | ✓ | ✓ | ✓ | ✗ | |
| USTBench Lai et al. (2026) | ✓ | ✓ | ✗ | ✗ | ✓ | – | ✓ | ✓ | ✗ | ✓ | ✓ | |
| Urban-multimodal benchmarks | UrBench Zhou et al. (2025) | ✗ | ✗ | ✓ | ✓ | ✗ | 11 | ✗ | ✗ | ✗ | ✓ | ✗ |
| DynamicVL Xuan et al. (2026) | ✗ | ✓ | ✓ | ✗ | ✓ | 42 | ✗ | ✗ | ✗ | ✓ | ✗ | |
| Model | Atomic PR | Workflow Outcome | Efficiency | ||||
| Recognition | Localization | Indicator F1 | Workflow Score | Avg. Turns | Tokens (K) | ||
| GPT-5.6-Terra | 0.419 | 0.250 | 0.080 | 0.063 | 0.071 | 9.9 | 184.7 |
| Gemini-3.7-Flash | 0.558 | 0.432 | 0.169 | 0.195 | 0.182 | 15.9 | 239.1 |
| Claude-Sonnet-5 | 0.558 | 0.386 | 0.159 | 0.111 | 0.135 | 18.2 | 228.5 |
| GPT-5.6-Sol | 0.500 | 0.318 | 0.080 | 0.103 | 0.091 | 12.1 | 177.9 |
| Kimi-K3 | 0.465 | 0.591 | 0.080 | 0.189 | 0.134 | 13.4 | 196.8 |
| Agent System | Recognition | Localization | Indicator F1 | Workflow Score |
| Qwen3.8-Max | ||||
| Direct | 0.659 | 0.216 | 0.126 | 0.171 |
| Plan-and-Execute | 0.591 | 0.125 | 0.128 | 0.126 |
| ReAct | 0.727 | 0.274 | 0.281 | 0.277 |
| MAF-Magentic | 0.295 | 0.057 | 0.081 | 0.069 |
| LAMBDA | 0.227 | 0.125 | 0.070 | 0.097 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Workflow Tasks ( ) | Atomic Tasks ( ) | ||
| Problem Complexity (3) | N | Capability Types (5) | N |
| Single-issue | 5 | Indicator quantification | 32 |
| Dual-issue | 7 | Temporal change analysis | 20 |
| Multi-issue | 10 | Spatial problem localization | 22 |
| Visual change interpretation | 6 | ||
| Evidence-to-metric grounding | 6 | ||
| Aspect | Count | Coverage |
| Spatial and temporal coverage | Count | Range |
| Cities / Regions | 12 / 22 | City / study-region cases |
| Observation Years | 10 | 2015–2024 |
| Data composition | No. of types | Included data |
| Source Families | 8 | OSM; WorldPop; GHSL; VIIRS; Landsat; Sentinel-2; Mapillary; WorldCover |
| Map | 2 | Roads; POIs/AOIs |
| Source | Data Content and Benchmark Use | Analysis Scale | Resolution | Access |
| OpenStreetMap | Administrative boundaries and vector features, including roads, buildings, land use, and amenities, for spatial delineation and accessibility analysis. | City/Region | – | OSM |
| WorldPop | Annual gridded population estimates for population change, density, and population-weighted accessibility analysis. | City/Region | 1 km | WorldPop |
| VIIRS Annual VNL | Annual nighttime-light composites used as proxies for urban activity and its temporal changes. | City | 15 arcsec | EOG |
| Landsat Collection 2 Level-2 | Surface-temperature rasters for thermal analysis, with Landsat 8 RGB imagery supplementing early periods without available Sentinel-2 imagery. | City/Region | 30 m grid | PC catalog |
| Sentinel-2 L2A | True-color satellite imagery for visual analysis of urban expansion, greenery, and surface change. | City/Region | 10 m (RGB) | PC catalog |
| GHSL GHS-BUILT-S | Multi-epoch built-up-surface estimates for analyzing urban expansion, sprawl, and stagnation. | City | 100 m | GHSL |
| ID | Checkpoint | Evaluation Criterion |
| C1 | Issue Context | Identifies the urban issue and establishes an appropriate analytical context. |
| C2 | Spatiotemporal Representation | Establishes the relevant spatial scope, temporal scope, and analysis units. |
| C3 | Multimodal Evidence Planning | Selects or plans evidence sources that are relevant to the analytical objective. |
| C4 | Indicator Acquisition | Obtains or derives the indicators needed to characterize the urban condition. |
| C5 | Comparative Analysis | Performs appropriate temporal, spatial, or reference-based comparisons using the obtained evidence. |
| C6 | Multimodal Evidence Integration | Relates evidence across relevant modalities rather than treating observations in isolation. |
| Evaluation | Raw Agreement | Accuracy vs. Gold |
| Annotator A vs. Annotator B | 93.6% | – |
| Annotator A vs. GLM-5.3 | 89.8% | – |
| Human annotator (A) | – | 97.8% |
| GLM-5.3 | – | 90.4% |
| Accuracy gap | – | 7.4 pp |
| Capability | Domain Skill | Representative Operations | Function and Output |
| Data discovery | data-card | aoi_data_inventory ; data_access_guide | Identifies available data sources, years, files, spatial scopes, and access methods for the current case. |
| Numerical analysis | urban-metrics | annual_mean_metric ; annual_total_metric ; temporal_change_rate ; temporal_series ; landcover_class_share ; population_weighted_ facility_distance ; raster_percentile | Computes annual statistics, temporal changes, land-cover shares, accessibility measures, and raster percentiles. |
| Map and facility queries | osm-map-query | poi_query ; building_query ; road_query ; landuse_query | Retrieves OSM roads, buildings, land-use polygons, and facilities such as hospitals, schools, and parks. |
| Satellite imagery access | raster-imagery | temporal_satellite_ retrieval ; regional_satellite_ retrieval | Selects Sentinel-2 imagery by year and region, using Landsat imagery when the required early observation is unavailable. |
| Raster visualization | raster-imagery | raw_raster_visualization ; normalized_raster_ visualization ; thematic_raster_ visualization | Converts LST, population, built-up, nighttime-light, and land-cover rasters into interpretable maps and summary statistics. |
| Spatial localization | spatial- localization | aoi_grid_generation ; region_bbox_resolution | Divides a city or study region into analysis units and resolves benchmark region identifiers to predefined spatial bounds. |