cs.CVSep 2, 2026

Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM-Based Urban Sensing

Authors: Kaizhen Tan

Organizations: New York University

Abstract

Street-view imagery is increasingly analysed with vision-language models (VLMs) to infer urban attributes, but predictive accuracy alone does not show how much a photograph contributes beyond data already available for the same place. Using three VLMs, we compare image-based predictions with existing urban data for seven attributes drawn from five public resources. Each urban unit is evaluated with imagery, with location or text context, and against non-image predictions from nearby observations or public records. We also replace images and add conflicting records to test which source the predictions follow. Nearby observations or public records matched or exceeded image-only predictions for road damage, curb ramps, population, and house price. Images were more informative for building type, building function, and the floor count of low-rise buildings. For floor count, the advantage of images over nearby OpenStreetMap labels increased by 5.7 percentage points per doubling of distance to the nearest labelled building and declined for tall buildings whose rooflines often fell outside the frame. When images and records disagreed, predictions usually moved towards the supplied record. Released OpenFACADES floor annotations, generated with OpenStreetMap floor values as input, showed the same dependence: their agreement with the reference increased with building height, whereas that of image-only reruns decreased. The value of street-view imagery therefore depends on whether an attribute is visible and how well the place is already covered by existing data. Because machine-derived labels are often reused as references, these comparisons also bear on how urban datasets are documented and evaluated.

Figures & tables

Appendix figures & tables17 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery

    Jul 16, 2026Zahratu Shabrina, Muhammad Asa, Jin Rui +2VLM EvaluationVLM Reasoning

  2. Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated

    May 30, 2026Rashid MushkaniVLM EvaluationInter-Rater Reliability