cs.LGSep 30, 2026

Network-based Spatial Context Retrieval for Open-weight LLMs: A Faithfulness Benchmark for Grounded Geographic Reasoning

Authors: Joan Perez

Organizations: Urban Geo Analytics

Abstract

Large language models (LLMs) encode substantial latent geographic knowledge, yet they reason poorly over space and are unreliable when queried from coordinates alone. Useful behaviour emerges only when structured spatial context is supplied in the prompt. This raises a question geographic evaluation has left unexamined: once the right context is supplied, does the model reason from it, or override it with its own parametric recall? We take up this question with an open pipeline for network-based spatial context retriev-al. In it, the surroundings of a selected point are defined by the pedestrian street network, the area actually reachable on foot. Using only open data and open-weight models, the pipeline retrieves features from OpenStreetMap and the GHS-POP population grid, computes indicators over the network catchment in code, and injects them as a compact spatial brief. On this basis we build a faithfulness benchmark. It labels every claim a model makes by its source (grounded in the brief, or drawn from training knowledge) and its correctness, and it probes each case with a planted false premise that the brief refutes. We evaluate sixteen open-weight model configurations across three families (Qwen, Gemma and Llama, with Gemma in two generations), four size classes and, where available, both thinking and non-thinking modes, on three con-trasting cities, resampling every case over ten seeds. The results show that resistance to the planted premise varies more strongly by model family and generation than by scale, while brief-reading competence forms a partly separate dimension. These behaviours are not captured by conventional world-correctness scores or single-shot evaluation. We release the implementation, spatial briefs, model outputs, and claim-level labels as a reproducible workflow at github.com/perezjoan/NSCR-LLM.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

    Aug 4, 2026Martin Böckling, Elizaveta Nosova, Heiko Paulheim +1Geospatial ReasoningMultilingual Benchmark

  2. GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

    Aug 7, 2026Rodrigo Ferreira Rodrigues, Karim Radouane, Jose G Moreno +1Geogs-SlamTemporal Reasoning

  3. SpatialChain: A Benchmark for Auditing Spatial Reasoning Faithfulness in VLMs

    Oct 5, 2026Rafael Teixeira Sousa, Vinícius Paulo Lopes de Oliveira, Elisa Ayumi Masasi de Oliveira +5Visual Question Answering Benchmarks