Continuing the Perception Test challenge series, we organised the fourth edition as a workshop at the European Conference on Computer Vision (ECCV) 2026 in Malmö, Sweden. This edition focused on spatial intelligence and featured four different tracks: unified multiple-choice videoQA and grounded videoQA from the original Perception Test benchmark, alongside two new tracks based on city-scale walking-tour videos (KilometerAudio and KilometerVision). In this report, we describe the new benchmarks used for the city-scale tracks and summarise the winning solutions across all tracks, including a generalist model that competed across all tracks with satisfactory performance. The winning solutions in the newly added city-scale tracks demonstrated that complex spatial and multimodal reasoning can be solved by expensive agentic pipelines, but remains difficult for multimodal models used standalone.
Figures & tables
Figure 1: Example of map trace tasks in the KilometerVision benchmark. The model receives a video and a text question ( “Which of these maps correctly represent the path of the person in this video?” ) together with options given as raster images, with or without text labels inpainted on the maps.
Figure 2: Example of landmark task in the KilometerVision benchmark, where the question contains a reference image in addition to the text prompt ( “Does the image show a location in the video, and if so, what is the straight-line distance between the camera location of the image and the camera location of the final video frame?” ).
Figure 3: Example KilometerAudio questions from the validation set (correct answer in green). Each requires reasoning over the audio track of an hour-long walking-tour video.
Figure 4: Multiple-choice audio-video benchmarks. The proposed KilometerAudio benchmark has the longest videos (more than 70 minutes on average) and the questions focus exclusively on audio events; Video-MME (Long) has videos up to 1h long, but only a fraction of the questions focus on audio events.
Split
Num videos
Num questions
Train
–
–
Validation
6
8
Test
218
500
Table 1: Splits of the KilometerAudio track.
Figure 6
Figure 7: Best multiple-choice video QA test results per edition. The 11,528 original questions are common to all editions, and the 1,842 unified questions were added in 2025; 2023, 2024, and the frequency baseline (striped) are therefore on the original questions only, while for 2025 and 2026 we show the accuracy on all 13,370 questions and on each subset.
Rank
Team name
top-1
Baseline
Random
0.313
Runner-up
F423 (zero-shot)
0.910
Best
Njust-KMG (fine-tuned)
0.914
Table 2: Multiple-choice video QA results on the test set.
Figure 8: Left: accuracy of the best 2026 submission (Njust-KMG) on each of the 141 unique questions. Right: distribution across areas of the questions that are quasi-solved (accuracy >95% , top) and of the 10 hardest questions (accuracy ≤80% , bottom).
Rank
Team name
HOTA
DetA
AssA
Baseline
MDETR+static
0.052
0.033
0.091
Runner-up
Huawei SpaceMind
0.654
0.651
0.661
Best
Banana Bats
0.664
0.654
0.677
Table 3: Grounded video QA results on the test set.
Figure 9: Baseline vs best results over time in terms of overall HOTA, detection, and association accuracy for the grounded video QA task.
Split
Num videos
Num questions
Train
–
–
Validation
13
17
Test
54
986
Table 4: Splits of the KilometerVision track.
Rank
Team name
top-1
Baseline
Random
0.196
Runner-up
Huawei SpaceMind
0.723
Best
ohmyyuan
0.753
Best-performing †
CloudAI-evolve
0.905
Table 5: KilometerVision results. † Best-performing entry, not eligible for a prize; see text.
Figure 10: Top-5 public test results (top-1 accuracy) in KilometerAudio (top) and KilometerVision (bottom), including the best-performing entry, CloudAI-evolve.
Rank
Team name
top-1
Baseline
Random
0.174
Runner-up
WW11
0.752
Best
IUCV
0.782
Best-performing †
CloudAI-evolve
0.864
Table 6: KilometerAudio results. † Best-performing entry, not eligible for a prize; see text.