Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM-Based Urban Sensing
Organizations: New York University
Abstract
Street-view imagery is increasingly analysed with vision-language models (VLMs) to infer urban attributes, but predictive accuracy alone does not show how much a photograph contributes beyond data already available for the same place. Using three VLMs, we compare image-based predictions with existing urban data for seven attributes drawn from five public resources. Each urban unit is evaluated with imagery, with location or text context, and against non-image predictions from nearby observations or public records. We also replace images and add conflicting records to test which source the predictions follow. Nearby observations or public records matched or exceeded image-only predictions for road damage, curb ramps, population, and house price. Images were more informative for building type, building function, and the floor count of low-rise buildings. For floor count, the advantage of images over nearby OpenStreetMap labels increased by 5.7 percentage points per doubling of distance to the nearest labelled building and declined for tall buildings whose rooflines often fell outside the frame. When images and records disagreed, predictions usually moved towards the supplied record. Released OpenFACADES floor annotations, generated with OpenStreetMap floor values as input, showed the same dependence: their agreement with the reference increased with building height, whereas that of image-only reruns decreased. The value of street-view imagery therefore depends on whether an attribute is visible and how well the place is already covered by existing data. Because machine-derived labels are often reused as references, these comparisons also bear on how urban datasets are documented and evaluated.
Figures & tables
| Attribute | Resource | Potential image cue | Existing urban data |
|---|---|---|---|
| Road damage | SVRDD ( Ren et al.,, 2024 ) | surface condition | nearby road observations |
| Curb ramp | Project Sidewalk ( Saha et al.,, 2019 ) | corner geometry | municipal asset register |
| Floor count | OpenFACADES ( Liang et al.,, 2025 ) | visible elevation and roofline | nearby OSM floor tags |
| Building type | OpenFACADES ( Liang et al.,, 2025 ) | facade form and use cues | local class prevalence |
| Building function | BuildingSense ( Su et al.,, 2026 ) | signage and entrances | nearby labels and place text |
| Population | CityLens ( Liu et al.,, 2026 ) | indirect neighbourhood cues | neighbouring statistics |
| Source | Information supplied | Role in the study |
|---|---|---|
| Task context | city, coordinates, or released text | VLM without the image |
| Nearby observations | labels from other urban units | spatial alternative |
| Public record | register or official statistics | administrative alternative |
| Attribute | Metric | Non-image source: score | Image only | Image gain range | ||
|---|---|---|---|---|---|---|
| GPT | Gemini | Qwen | ||||
| Road damage | Balanced accuracy | Nearby frames: 72.3 | 50.4 | 54.7 | 52.6 | to |
| Curb ramp | Accuracy | SDOT register: 78.0 | 76.0 | 60.7 | 52.7 | to |
| Floors, 1–4 | Within 1 floor | Nearby tags: 59.3 | 92.0 | 88.7 | 86.0 | to |
| Floors, 5–14 | Within 1 floor | Nearby tags: 43.2–43.9 | 40.3 | 64.2 | 54.7 | to |
| Floors, 15+ | Within 1 floor | Nearby tags: 17.4–17.7 | 3.4 | 21.5 | 18.4 | to |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Attribute | Sources considered | Main comparison source | Urban setting |
|---|---|---|---|
| Road damage | prevalence, nearby frames | nearby frames | gaps in a dense survey |
| Curb ramp | prevalence, nearby labels, SDOT | SDOT register | municipal inventory |
| Floor count | prevalence, nearby OSM | nearby OSM tags | incomplete building inventory |
| Building type | prevalence, nearby OSM | modal OSM class | local class distribution |
| Building function | prevalence, nearby labels | nearby labels | sampled local labels |
| Population | prevalence, official neighbours | official neighbours | area statistics |
| Exclusion radius | 50 m | 100 m | 200 m | 400 m |
|---|---|---|---|---|
| Balanced accuracy | 72.3 [69.1, 75.3] | 69.4 [66.3, 72.5] | 66.3 [63.0, 69.5] | 64.3 [61.2, 67.4] |
| Macro accuracy | 82.6 [80.6, 84.6] | 80.3 [78.3, 82.3] | 77.8 [75.7, 80.0] | 76.5 [74.4, 78.5] |
| Predictor | Macro acc. | Macro-F1 | Balanced acc. | Positive recall |
|---|---|---|---|---|
| Prevalence benchmark | 75.3 | 0.0 | 50.0 | 0.0 |
| Buffered spatial | 82.6 | 59.6 | 72.3 | 54.9 |
| GPT image only | 74.0 | 7.9 | 50.4 | 4.8 |
| Gemini image only | 73.9 | 27.5 | 54.7 | 23.7 |
| Qwen image only | 73.3 | 23.6 | 52.6 | 19.7 |
| Model | Current | Source-era | Difference | Changed |
|---|---|---|---|---|
| GPT-4o-mini | 80.0% | 74.0% | [ , ] | 30.0% |
| Gemini 2.5 Flash Lite | 63.0% | 68.0% | [ , ] | 31.0% |
| Qwen3.5 Flash | 50.0% | 54.0% | [ , ] | 4.0% |
| Task | Donor separation |
|---|---|
| Road damage | another district, with at least one different damage label |
| Curb ramp | opposite reference class |
| Floor count | another building, usually differing by more than one floor |
| Building type | different reference class |
| Building function | independently sampled building |
| Population | area in another city |
| Predictor | Accuracy | Macro-F1 | Balanced accuracy |
|---|---|---|---|
| Modal class | 58.8 | 12.3 | 16.7 |
| Spatial predictor | 27.8 | 18.7 | 25.2 |
| GPT image only | 72.2 | 64.1 | 70.8 |
| Gemini image only | 77.3 | 74.2 | 76.1 |
| Qwen image only | 76.3 | 64.3 | 63.1 |
| Term | Estimate | Bootstrap SE | 95% interval |
|---|---|---|---|
| Intercept | |||
| nearest-label distance (m) | |||
| camera-to-building distance (m) | |||
| Capture year minus 2022 | |||
| Height: 1–4 floors | |||
| Height: 5–14 floors |
| Attribute | Image | Coordinates | Difference | |
|---|---|---|---|---|
| Floors, within 1 | 450 | 50.7% | 17.3% | |
| Building type | 291 | 76.6% | 46.4% | |
| Construction decade | 93 | 71.0% | 34.4% | |
| Distance to hospital | 450 | 20.4% | 22.4% |
| Predictor | Accuracy | Macro-F1 | Balanced accuracy |
|---|---|---|---|
| Spatial (50 m) | 56.0% | 6.5% | 8.4% |
| GPT image only | 63.0% | 23.9% | 27.6% |
| Gemini image only | 59.0% | 24.2% | 24.0% |
| Qwen image only | 63.0% | 25.4% | 31.1% |
| Model | Nearby place added | Address shuffle | Image added, no nearby place |
|---|---|---|---|
| GPT | |||
| Gemini | |||
| Qwen |
| Task | GPT | Gemini | Qwen |
|---|---|---|---|
| Road damage | |||
| Curb ramp | |||
| Floor count | |||
| Building type | |||
| Building function | |||
| Population |
| Task | Model | Image prompt | Statistics input | Difference |
|---|---|---|---|---|
| Population | GPT | 2.76 | 1.88 | |
| Gemini | 2.18 | 1.90 | ||
| Qwen | 1.97 | 1.98 | ||
| House price | GPT | 2.94 | 0.25 | |
| Gemini | 4.38 | 0.41 | ||
| Qwen | 1.85 | 0.24 |