We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing KGQA benchmarks for Web knowledge, GeoOutageBench considers a spatiotemporal KG that integrates visual, textual, and structured data from outage records, remote sensing, weather observations, storm and power events, geographic entities, and domain ontologies. It provides a competency query taxonomy at different difficulty levels from spatiotemporal containment and proximity, spatiotemporal co-occurrence analysis, multimodal evidence, to hypothetical evaluation. Over multimodal KG and query classes, GeoOutageBench provides user-configurable evaluation of three important, highly coherent yet less studied tasks: (1) LLMs' understanding for ambiguous geospatiotemporal questions in terms of NL to SPARQL interpretation, (2) query-driven assessment of ontology utility, and (3) answer accuracy of multimodal KGQA retrieval. GeoOutageBench provides a design principle and foundation for assessing LLM-KG systems that support real-world infrastructure resilience analysis. Our benchmark, source code, data, results, and other documentation are available at https://github.com/UCF-SAGE/GeoOutageBench.
Figures & tables
Figure 1. Abstract view of how one question can support multiple KGQA interpretations. The same term can bind the state, year, SVI threshold, and event-count semantics differently, while alternative ontology-compatible traversals can reach the same answer through different paths in KG. A one-column abstract branching diagram for a GeoOutageBench natural-language question. The question points to four underspecified slots labeled as spatial, temporal, semantic, and semantic ambiguity: state, year, SVI threshold, and the meaning of dozen or so storm events. The first expanded interpretation box shows a query graph bound to Florida, year 2022, SVI at least 0.75, Thunderstorm Wind events, and count at least 12. The second interpretation keeps the Louisiana 2018 SVI at least 0.85, Wind-event count between 10 and 14 semantics, but grounds the county path in a nighttime-light image, acquisition year, and VIIRS sensor before joining with SVI and storm-event records.
Figure 2. GeoOutageBench framework. Dataset assets are linked through GeoOutageOnto and instantiated in GeoOutageKG; natural-language questions are mapped to ambiguity-aware SPARQL interpretations; and LLMs, GraphDB ( Ontotext, 2026 ) , and OntoCheck ( Kundu et al., 2025 ) support the benchmark tasks and custom metrics P . Top-down sketch of the GeoOutageBench framework. An orange GeoOutageBench box containing the KGQA profile tuple is centered above three aligned process panels connected by arrows. The Dataset panel contains a Multimodal Data group with text-only boxes for Tabular Data, Remote Sensing Data, and Outage Maps, followed by GeoOutageOnto and GeoOutageKG boxes. The Query Construction panel contains Natural Language Questions and ambiguity-aware SPARQL Interpretations grounded in the ontology, including semantic, spatial, temporal, and spatiotemporal readings with one-hop, two-hop, and multi-hop graph traversals. The Tools and Benchmark Tasks panel contains a Tools group with LLMs, GraphDB, and OntoCheck, a raised Benchmark Tasks group with NL2SPARQL, spatiotemporal KGQA, and ontology-utility assessment, and a custom evaluation metrics box for P covering ambiguity, schema-linking, spatiotemporal, answer-quality, and ontology-utility metrics.
Figure 3. A fraction of GeoOutageOnto. Data classes are displayed in blue, geospatial classes are in yellow, temporal properties are in orange, and miscellaneous instrument/superclass classes are in gray. goo stands for GeoOutageOnto, mds stands for MDS-Onto ( Rajamohan et al., 2025 ) , and xsd stands for XML Schema Definition ( World Wide Web Consortium, 2004 ) . Ontology diagram showing the GeoOutageOnto subset. Data classes are displayed in blue, geospatial classes are in yellow, temporal properties are in orange, and miscellaneous instrument/superclass classes are in gray. \texttt{goo} stands for GeoOutageOnto, \texttt{mds} stands for MDS-Onto, and \texttt{xsd} stands for XML Schema Definition. Edges represent semantic relationships among outage, storm, remote sensing, administrative region, and observation entities.
Class
Instances
Triples
CustomerOutageRecord
11733064
117330640
NTLImage
335469
7907444
OutageMap
191661
2108267
StormEventRecord
2012856
36211668
HurricaneEventRecord
3266
128831
SVIRecord
21997
395940
Table 1. GeoOutageKG instance and triple counts by class.
Figure 4. Multimodal data construction for GeoOutageBench. County geography provides the shared spatial key used to segment nighttime-light imagery, derive radiance-loss outage maps, and ground outage, storm, hurricane, and vulnerability records as county-level GeoOutageKG evidence. One-column multimodal data workflow for GeoOutageBench. Remote-sensing, geographic, and tabular sources feed a shared county key. The raster branch clips nighttime-light imagery by county, places the NTL image stack and outage-map stack side by side with a rightward arrow between them, and shows radiance loss below the stacks. The tabular branch uses one combined normalization and county-date alignment block for outage, storm, hurricane, and vulnerability records, then passes through overlapping cylindrical Counties boxes. Both branches instantiate county-level GeoOutageKG evidence bundles.
Coverage by hop depth
Category
1-hop
2-hop
Multi-hop
Total Q / I
Example question
Underlying operation
Spatial containment
12 / 13
7 / 8
7 / 8
12 / 13
Which U.S. county [ jurisdiction type ] is depicted in the power outage severity map [ map type ] corresponding to a specific customer outage record [ outage record ]?
Region containment, jurisdiction lookup, or spatial aggregation
Spatial proximity
4 / 9
4 / 9
4 / 9
4 / 9
On Ian’s Florida landfall day [ event date ], what was the peak outage count in the southwest Florida county centered on Fort Myers [ place descriptor ], and were any hurricanes active within the general area of the county [ distance/radius abstraction ] that day?
Buffer, distance filter, spatial join
Temporal interval
11 / 14
3 / 3
2 / 2
11 / 14
How many and which hurricanes [ cyclone type ] made landfall in Louisiana [ jurisdiction ] in August [ month ] across all years?
Temporal filtering, aggregation, and interval comparison
Event sequence
3 / 7
3 / 7
3 / 7
3 / 7
Which storm events [ event type ] in Florida [ state ] started before same-county EAGLE-I records [ outage record source ] that later exceeded 50,000 outages [ outage threshold ] on the same day [ temporal alignment ]?
Temporal precedence over event graph
Spatiotemporal co-occurrence
3 / 7
3 / 7
2 / 4
3 / 7
Which counties [ jurisdiction type ] in the hurricane-prone state [ state descriptor ] had high socioeconomic vulnerability [ vulnerability descriptor ] in a recent year [ relative year descriptor ] and a dozen or so wind-related storm events [ event-count condition ] recorded that same year [ temporal alignment ]?
Spatial–temporal join
Table 2. Query classification by primary category. Coverage entries are template instances / SPARQL interpretations and are inclusive: an interpretation at a given hop depth also satisfies shallower depths. Ambiguous spans in the example questions are shown as surface form [ binding type ].
Model
Syntax
Compat.
Exact
Align.
Class F1
Prop. F1
Schema F1
Spatial F1
Temporal F1
ST F1
GPT-5.5
1.000
1.000
0.000
0.745
0.716
0.679
0.678
0.870
0.715
0.662
Gemini 3.1 Pro
1.000
0.979
0.021
0.749
0.834
0.703
0.737
0.783
0.675
0.578
Claude Opus 4.7
1.000
1.000
0.000
0.805
0.807
0.774
0.772
0.886
0.754
0.711
Table 4. Task 1 macro results for NL2SPARQL over 48 GeoOutageBench questions.
Model
Axis
Align.
S F1
T F1
ST F1
GPT-5.5
Non-ST
0.585
1.000
0.250
0.250
GPT-5.5
S
0.699
0.969
0.625
0.607
GPT-5.5
T
0.751
1.000
0.626
0.676
GPT-5.5
ST
0.776
0.812
0.807
0.725
Gemini 3.1 Pro
Non-ST
0.688
1.000
0.250
0.250
Gemini 3.1 Pro
S
0.697
0.875
0.500
0.375
Table 5. Assessment of LLMs in Task 1, with finer-grained analysis over spatiotemporal axis. S abbreviates Spatial, T abbreviates Temporal, and ST abbreviates Spatiotemporal.
Figure 5. Single- versus multimodal querying for the outage-and-storm case study. The multimodal variant supplements county-level tabular answers with date-aligned NTL images and outage maps, enabling evidence-aware visual inspection and richer answer-level assessment. Two-panel case-study figure comparing single-modal and multimodal KGQA. The single-modal panel asks for Florida counties in 2022 with high socioeconomic vulnerability and at least twelve wind-related storm events, returning SVI records, storm-event counts, and county rows. The multimodal panel requests nighttime-light images and derived outage maps around Hurricane Ian for those counties, returning county bindings along with image assets from September 29 and September 30, 2022.
Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logistics, yet large language models (LLMs) struggle with the required geometric and topological computation despite storing considerable geographic knowledge. Existing benchmarks localize these failures only partially: they are synthetic or smallscale, largely monolingual, and offer limited control over geographic coverage. We introduce MultiGlobeQA, a multilingual benchmark of 46,060 question-answer pairs spanning 14 spatial-function families and 15 answer formats, with execution-based ground truth over three knowledge graphs. It covers 201 countries and territories via income- and density-stratified sampling, with parallel questions in English and 16 additional high- and low-resource languages. Across parametric, reasoning, and agentic settings, LLMs collapse on tasks requiring grid indexing and shape computation, while topological relations and directions fare best. Retrieval and tool use yield considerable gains, yet performance plateaus below two thirds even when gold facts are supplied, indicating that computation, not access to knowledge, is the bottleneck. Models also underperform on low-income regions, a gap that gold facts widen rather than close.
Martin Böckling, Elizaveta Nosova, Heiko Paulheim +1
Data and Web Science Group, University of Mannheim, Germany
Remote-sensing vision-language models (RS-VLMs) have advanced Earth-observation analysis toward visual interpretation and instruction-following, yet fall short of operational geo-intelligence, which demands tool-grounded spatial reasoning and structured, evidence-backed decisions. We introduce GeoDisaster, an operational geospatial disaster reasoning benchmark with 2,921 verified instances across 43 question types and five task families: deforestation monitoring, multi-hazard analysis, building-damage assessment, flood-safe routing, and Sentinel-1 SAR flood monitoring. Instances integrate heterogeneous EO/GIS evidence-optical and SAR imagery, raster masks, vector geometries, road networks, and exposure layers-spanning hazard detection, damage assessment, exposure estimation, and diagnostic report generation. Ground-truth answers are grounded in executable geospatial workflows and deterministic consistency checks, removing the need for language-model annotation. We further propose an orchestrated multi-agent framework with 18 disaster-oriented tools, where role-specialized agents coordinate through explicit execution contracts, aligned via Role-Contract Expectation Alignment (RCEA): failure-aware supervised fine-tuning combined with contract-grounded reinforcement learning over dense step-level signals. Experiments show that GeoDisaster challenges existing RS-VLMs and agentic systems, while RCEA improves tool use, evidence grounding, state consistency, and decision generation.
Maram Hasan, Aman Verma, Savitra Roy +5
Indian Institute of Technology Bombay · Mohamed bin Zayed University of Artificial Intelligence
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into their generalization capabilities. In this paper, we present \benchName, a comprehensive benchmark for probing LLMs on geo-related tasks. We leverage a careful selection of twelve publicly available datasets from diverse geo-related tasks and domains, and evaluate a set of LLMs on geo-spatial and temporal understanding using our benchmark. Our results show that reasoning and size have a strong impact on overall performance. GeoBenchLLM is publicly available at https://github.com/Rfr2003/GeoBenchLLM.
Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno +1