Continuing the Perception Test challenge series, we organised the fourth edition as a workshop at the European Conference on Computer Vision (ECCV) 2026 in Malmö, Sweden. This edition focused on spatial intelligence and featured four different tracks: unified multiple-choice videoQA and grounded videoQA from the original Perception Test benchmark, alongside two new tracks based on city-scale walking-tour videos (KilometerAudio and KilometerVision). In this report, we describe the new benchmarks used for the city-scale tracks and summarise the winning solutions across all tracks, including a generalist model that competed across all tracks with satisfactory performance. The winning solutions in the newly added city-scale tracks demonstrated that complex spatial and multimodal reasoning can be solved by expensive agentic pipelines, but remains difficult for multimodal models used standalone.
Figures & tables
Figure 1: Example of map trace tasks in the KilometerVision benchmark. The model receives a video and a text question ( “Which of these maps correctly represent the path of the person in this video?” ) together with options given as raster images, with or without text labels inpainted on the maps.
Figure 2: Example of landmark task in the KilometerVision benchmark, where the question contains a reference image in addition to the text prompt ( “Does the image show a location in the video, and if so, what is the straight-line distance between the camera location of the image and the camera location of the final video frame?” ).
Figure 3: Example KilometerAudio questions from the validation set (correct answer in green). Each requires reasoning over the audio track of an hour-long walking-tour video.
Figure 4: Multiple-choice audio-video benchmarks. The proposed KilometerAudio benchmark has the longest videos (more than 70 minutes on average) and the questions focus exclusively on audio events; Video-MME (Long) has videos up to 1h long, but only a fraction of the questions focus on audio events.
Split
Num videos
Num questions
Train
–
–
Validation
6
8
Test
218
500
Table 1: Splits of the KilometerAudio track.
Figure 6
Figure 7: Best multiple-choice video QA test results per edition. The 11,528 original questions are common to all editions, and the 1,842 unified questions were added in 2025; 2023, 2024, and the frequency baseline (striped) are therefore on the original questions only, while for 2025 and 2026 we show the accuracy on all 13,370 questions and on each subset.
Rank
Team name
top-1
Baseline
Random
0.313
Runner-up
F423 (zero-shot)
0.910
Best
Njust-KMG (fine-tuned)
0.914
Table 2: Multiple-choice video QA results on the test set.
Figure 8: Left: accuracy of the best 2026 submission (Njust-KMG) on each of the 141 unique questions. Right: distribution across areas of the questions that are quasi-solved (accuracy >95% , top) and of the 10 hardest questions (accuracy ≤80% , bottom).
Rank
Team name
HOTA
DetA
AssA
Baseline
MDETR+static
0.052
0.033
0.091
Runner-up
Huawei SpaceMind
0.654
0.651
0.661
Best
Banana Bats
0.664
0.654
0.677
Table 3: Grounded video QA results on the test set.
Figure 9: Baseline vs best results over time in terms of overall HOTA, detection, and association accuracy for the grounded video QA task.
Split
Num videos
Num questions
Train
–
–
Validation
13
17
Test
54
986
Table 4: Splits of the KilometerVision track.
Rank
Team name
top-1
Baseline
Random
0.196
Runner-up
Huawei SpaceMind
0.723
Best
ohmyyuan
0.753
Best-performing †
CloudAI-evolve
0.905
Table 5: KilometerVision results. † Best-performing entry, not eligible for a prize; see text.
Figure 10: Top-5 public test results (top-1 accuracy) in KilometerAudio (top) and KilometerVision (bottom), including the best-performing entry, CloudAI-evolve.
Rank
Team name
top-1
Baseline
Random
0.174
Runner-up
WW11
0.752
Best
IUCV
0.782
Best-performing †
CloudAI-evolve
0.864
Table 6: KilometerAudio results. † Best-performing entry, not eligible for a prize; see text.
We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark is publicly available at https://perception-test-challenge.github.io/kilometervision.html.
Aravindh Mahendran, Michael King, Matthew Koichi Grimes +12
Google DeepMind, Berlin, Germany · Google DeepMind, London, UK · Princeton University, Princeton, USA +3
Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation of the output (e.g. the answer to a question), without evaluation of intermediate reasoning steps, and most provide answers only in the text domain. We introduce Minerva-Ego, a benchmark for evaluating complex egocentric visual reasoning. We extend recent high-quality video data sources recorded from egocentric / embodied settings with a set of challenging, multi-step multimodal questions and spatiotemporally-dense human-annotated reasoning traces. Benchmarking experiments show that state-of-the-art models still have a large gap to human performance. To investigate this gap in detail, we annotate each reasoning trace in the dataset with the objects of interest required to solve the question, as spatiotemporal mask annotations. Through extensive evaluations, we identify that prompting frontier models with hints of 'where' and 'when' to look yields substantial improvements in performance. Minerva-Ego can be downloaded at https://github.com/google-deepmind/neptune.
Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams, whereas existing benchmarks largely fail to evaluate this capability. They restrict videos to short clips, isolate modalities, or reduce questions to one-hop perception. We introduce TraceAV-Bench, the first benchmark to jointly evaluate multi-hop reasoning over long audio-visual trajectories and multimodal hallucination robustness. TraceAV-Bench comprises 2,200 rigorously validated multiple-choice questions over 578 long videos, totaling 339.5 hours, spanning 4 evaluation dimensions and 15 sub-tasks. Each question is grounded in an explicit reasoning chain that averages 3.68 hops across a 15.1-minute temporal span. The dataset is built by a three-step semi-automated pipeline followed by a strict quality assurance process. Evaluation of multiple representative OmniLLMs on TraceAV-Bench reveals that the benchmark poses a persistent challenge across all models, with the strongest closed-source model (Gemini 3.1 Pro) reaching only 68.29% on general tasks, and the best open-source model (Ming-Flash-Omni-2.0) reaching 51.70%, leaving substantial headroom. Moreover, we find that robustness to multimodal hallucination is largely decoupled from general multimodal reasoning performance. We anticipate that TraceAV-Bench will stimulate further research toward OmniLLMs that can reason coherently and faithfully over long-form audio-visual content.
Hengyi Feng, Hao Liang, Mingrui Chen +6
University of Electronic Science and Technology of China · Peking University · Zhongguancun Academy +1