cs.CVSep 30, 2026

KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs

Authors: Aravindh Mahendran, Michael King, Matthew Koichi Grimes, Antoine Yang, Tyler Zhu, Joseph Heyward, Tengda Han, Shiry Ginosar, +7 more

Organizations: Google DeepMind, Berlin, Germany · Google DeepMind, London, UK · Princeton University, Princeton, USA · Toyota Technological Institute at Chicago, Chicago, USA · Google DeepMind, USA · Google, Zurich, Switzerland

Abstract

We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark is publicly available at https://perception-test-challenge.github.io/kilometervision.html.

Figures & tables

Explore similar work

CardsList
  1. GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

    Aug 6, 2026Qifeng Zhang, Kaixiang Huang, Heng Dong +6Stable Spatial UnderstandingVisual Question Answering Benchmarks

  2. SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

    Aug 3, 2026Jing Wu, Jianhua Wu, Jiayi Guan +5Recent Vision-Language ModelsStable Spatial Understanding

  3. Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

    May 28, 2026Cheolhong Min, Jaeyun Jung, Daeun Lee +5Entanglement