DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response
Organizations: Computer Vision and Learning System, Linköping University, Sweden · Vantor, Linköping, Sweden
Abstract
Vision--language models (VLMs) show promise for disaster-response remote sensing, but existing benchmarks mainly emphasize scene-level or damage-centric assessment. To study this building-centric gap, we introduce \method{}, a diagnostic benchmark built on xBD, a pre/post-disaster satellite dataset with building-level damage labels. \method{} enriches building instances with OpenStreetMap-derived functional labels and contains 134{,}108 task-specific instruction records across 15 task types, spanning instance-level assessment, scene-level counting, multi-instance reasoning, and structured report generation. The benchmark supports RGB pre/post-disaster imagery, single- and multi-view instance formulations, and scene-level RGB/SAR diagnostic inputs. Experiments with general-domain and remote-sensing VLMs show that models perform better on visible damage cues than on building-function understanding, multi-instance reasoning, counting, and grounded reporting. Instruction tuning improves performance on several tasks but does not close this building-centric gap.
Figures & tables
| Dataset | Bldg. Inst. | Damage | Modality | Report | Function |
|---|---|---|---|---|---|
| RSICD [ 15 ] | ✗ | ✗ | Optical + caption | Caption | ✗ |
| BRIGHT [ 5 ] | ✗ | ✓ | Optical/SAR | ✗ | ✗ |
| FloodNet [ 20 ] | ✗ | ✓ | UAV RGB + QA | VQA | ✗ |
| xBD [ 8 ] | ✓ | ✓ | Pre/post optical | ✗ | ✗ |
| CrisisMMD [ 1 ] | ✗ | Partial | Image + text | Summ. | ✗ |
| CrisisFACTS [ 16 ] | ✗ | ✗ | Text | Summ. | ✗ |
| Family | Task types | Input | Metrics |
|---|---|---|---|
| Instance | Building-function classification; damage-level estimation; temporal-change detection; joint function–damage assessment | Target building views with pre/post imagery | Acc., macro-F1, MAE, exact match |
| Scene | Total building count; damage-level count; function count | Full pre/post scene imagery | MAE, RMSE, key-wise MAE |
| Report | Short structured report generation; long structured report generation | Full scene imagery; RGB–SAR scene setting | BLEU-4, ROUGE-L, BERTScore |
| Multi-instance | Selection by damage; selection by function; joint function–damage selection; damage comparison; spatial reasoning; response-priority reasoning | Marked buildings in pre/post scene imagery | Acc., exact match, Set-F1 |
| Instance-level | Multi-instance | ||||||||
| Func. | Dmg. | Temp. | F-Sel. | Spatial | Priority | ||||
| Model | Acc. | mF1 | Acc. | MAE | Acc. | mF1 | Set-F1 | Acc. | Acc. |
| Open-source models | |||||||||
| Qwen2.5-VL-7B | 74.9 | 21.0 | 81.2 | 0.32 | 44.8 | 16.8 | 66.5 | 47.4 | 33.3 |
| Qwen3-VL-8B | 86.6 | 25.9 | 72.7 | 0.46 | 44.1 | 27.9 | 66.4 | 53.4 | 31.0 |
| Qwen3-VL-30B | 91.0 | 30.1 | 64.2 | 0.69 | 22.7 | 13.3 | 70.3 | 39.8 | 31.0 |
| Counting | Reports | |||||
| Model | Tot. | Func. | Dmg. | B4 | R-L | BS |
| Open-source models | ||||||
| Qwen2.5-VL-7B | 66.0 | 6.5 | 19.2 | 1.8 | 16.3 | 84.3 |
| Qwen3-VL-8B | 175.6 | 9.8 | 20.5 | 3.1 | 17.7 | 84.8 |
| Qwen3-VL-30B | 69.7 | 15.2 | 18.3 | 2.2 | 16.6 | 84.2 |
| LLaVA-OneVision | 78.6 | 5.7 | 20.2 | 3.4 | 19.4 | 85.4 |
| OSM Labels | Benchmark Category |
|---|---|
| house , apartments , residential , detached , bungalow , static_caravan | residential |
| hospital , clinic , pharmacy , doctors , dentist , veterinary | medical_facilities |
| school , college , university , kindergarten , childcare | educational_facilities |
| fire_station , police , courthouse , shelter , prison , townhall , government | government_emergency |
| supermarket , convenience , mall , restaurant , fast_food , bank , retail , commercial , doityourself , computer , car_wash , pet | commercial |
| warehouse , factory , storage_rental , industrial , fuel , power , works | industrial_utilities |