KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs
Authors: Aravindh Mahendran, Michael King, Matthew Koichi Grimes, Antoine Yang, Tyler Zhu, Joseph Heyward, Tengda Han, Shiry Ginosar, +7 more
Organizations: Google DeepMind, Berlin, Germany · Google DeepMind, London, UK · Princeton University, Princeton, USA · Toyota Technological Institute at Chicago, Chicago, USA · Google DeepMind, USA · Google, Zurich, Switzerland
We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark is publicly available at https://perception-test-challenge.github.io/kilometervision.html.
Figures & tables
Figure 1 : KilometerVision – the first benchmark probing city-scale spatial understanding from real-world videos spanning 1km distances. Left: Real-world hour-long video of a walking tour in Istanbul and relevant questions one would ask when visiting an unfamiliar place. Centre: Grounding the video on the map instantly reveals loop closures, distances between landmarks, or sequences of streets visited. Right: Using the landmark-route-map paradigm from cognitive science, we define tasks to comprehensively evaluate visual city-scale spatial intelligence in video-language models.
Figure 2 : Left: Distribution of question types in the KilometerVision benchmark. Right: Histogram over video lengths across question types.
Figure 4 : Left: Distribution of line-of-sight distances (in metres) to landmarks used in our dataset. We also include questions with landmarks that haven’t been visited on the tour, for which the correct answer is ‘not seen’. Centre-left: Distribution of orientations at the end of the walking tour for the compass task. Centre-right: Distribution of Euclidean distances between the start and the end of each tour. Right: Distribution of loop lengths in seconds; maximum video length is 600 seconds (10 minutes). We also include 11 question-answer pairs where there is no loop.
Figure 5 : (a) Example of a VPS trace with a loop. (b) Example of a VPS trace without a loop. (c) Given the VPS coordinates for a video segment (shown in red), we select high-confidence intermediate VPS points to query the Google Maps Compute Routes API and get a closely-matching trace (shown in blue). Black segments indicate the optimal pairings found by DTW (Dynamic Time Warping) metric. In green, we show the maximum distance (error).
Figure 6 : Top : Map trace images used as options for the Map trace task. Bottom : A more challenging map trace task variation where the route traces are in-painted over maps without text labels or landmark icons.
Figure 7 : Performance of different VLMs on KilometerVision, compared to human performance and a blind baseline.
Figure 8 : Impact of frame rate (left) and of spatial resolution (right) on performance for different models.
Figure 9 : Left: Gemini 2.5 Flash performance on Map Trace questions with additional camera and/or VPS information. Right: Performance on Landmark and Loop Closure questions separated into Visual Recognition ( i.e . determining the presence or absence of a landmark or loop) and selecting the correct distance / time option, compared to overall (combined) performance.
Figure 10 : Example of thinking traces from Gemini 2.5 Flash and Claude Opus on loop closure questions. Both models rely purely on visual recognition to solve the task. Gemini correctly identifies the location at the end of the video 1 , and correctly eliminates another option 2 . The model incorrectly believes that a different statue that it can see in the distance is the same as the one it sees at the end of the video 3 . The model correctly determines that the camera is in the same location at time 9:17 but incorrectly decides that this was not visited earlier 4 . Claude Opus ’s reasoning is generally correct and it correctly identifies option ID 2 as the right answer.
Table 11
Figure 11 : The optimal true positive rate achievable across all settings as the allowed false positive rate is increased.
Figure 12 : 3 of the 5 options from an easier version of the route summary and map trace questions in which the options are in roughly the same location but paths are not overlapping.
Figure 13 : Accuracy of Gemini 2.5 Flash on the route summary and map trace questions where options are non-overlapping (Easy) vs overlapping (Hard). We include the Hard variant in the official benchmark.
Figure 14 : Sample answer from a compass question for Gemini 2.5 Flash (left) and Claude Opus (right). We highlight correct statements in green and mistakes in red. The number labels point to different sections in the video whose sample frames are shown at the top and bottom. We explain these for Gemini 2.5 Flash next. [1] At 00:00-00:42 a 45 degree turn alongside camera translation, which is incorrectly labelled as a ‘Stationary pan’. [2] The model correctly recognises the landmark and the overall turn direction. [3] At 00:42 it has correctly understood that a 45 degree turn took place. ‘South’ is correct. [4] 00:42 - 1:15 is a 90 degree right turn. The model mistakes this for a 5-10 degree turn. Claude Opus, on the other hand, is able to answer the question correctly but is incorrect in the middle where Southwest should in fact be south and the overall rotation is not 90 degree but closer to 135 degrees.
Figure 15 : In this ablation experiment for the Route summary task, the model is shown a video and asked to pick a route summary that matches the video. As helper, the model also receives a map with the correct map trace followed by the person in the video. Intuitively, this should make the task much easier. The top part of the text insert shows the response from Gemini 2.5 Flash with parts of the thinking trace. The model relies largely on street signs in the video and street names on the map, and puts together the route summary in its own words. Orientation / direction are recognised relative to streets such as ‘southwest on Zelená’, ’turns right from Zelená’. However, it gets the final answer wrong because it mistakes ‘Head northwest on Hlavné námestie’ with ‘southwest from Hlavné námestie’. The model failed to accurately describe the motion on Hlavné námestie itself and thus got confused. Below the line is the response from Claude Opus. It is very general in it’s analysis and doesn’t try to estimate turn directions. This leads it to rely mostly on matching street names and thus gets the answer right.
Figure 16 : Sample answer from a landmark recognition and distance-to-landmark estimation question. We highlight correct statements in green and mistakes in red. The numbered labels for the Gemini2.5 flash answer, in this case, simply tag relevant sections of the answer which we explain next: [1] Landmark (shown in the inset bottom right large image) has been recognized correctly as “La Vecchia Scuola” and the model detects it at the correct time stamp. The thinking trace, not shown for brevity, does this instantly by saying "I’ve pinpointed the landmark within the video at 05:15.000, confirming the image’s presence." [2] “La Vecchia Scuola” is not at at 8-9 low Petergate and the coordinates are in-fact a third location. [3] While “Bartle garth” is correct, it incorrectly guesses the video end location as Bedern hall which in-fact lies on the other end of Bartle garth. In the thinking trace, not shown for brevity, we detect “I’ve determined the path: Low Petergate to Goodramgate to Bartle Garth.” which shows that the model is only recognising familiar streets. Note that all references to tool use in the answer are hallucinations. Tool use was not enabled in any of our experiments. Over-reliance on memorised landmarks and their locations makes the model latch on to the wrong address and wrong coordinates. Course street name level path integration is insufficient to answer this question. Claude also identifies the landmark correctly and comes to a similar but incorrect conclusion of 140m as the right answer. The Correct answer should in fact be 178m in this case.