A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers that never see a frame approach full-video accuracy. On 51 benchmarks with temporal probes, shuffled frames keep a median 96% of full-video accuracy. Near-duplicate questions make up at least half the items in 63 benchmarks. We screen 505,518 question-answer pairs from 112 of them into an audited pool. Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks. We release Video-Index, the 210 hardest verified items under these attacks in each of four capability groups, 840 items from 76 sources. With the same fixed input, Claude Opus 5 outscores every open-source model by over 37 percentage points, and agent tools add about 20 more, yet all systems leave room to improve efficiency and accuracy. Blog: https://www.enxinsong.com/blog/video-index/ GitHub: https://github.com/Espere-1119-Song/Video-Index Hugging Face: https://huggingface.co/datasets/Video-Index/Video-Index
Figures & tables
Figure 1: From pool to Video-Index. We filter the 505,518 items in four stages and keep the 840 hardest.
Figure 2: Options and question text replace the video on 30 benchmarks. Each benchmark is placed by the full video’s gain over options alone and over the question text, and shaded bands mark breaks.
Figure 3: Remembering earlier answers narrows the video gap. Each column is a benchmark, with its gap to the strongest pool attacker after 25 to 100% of the items.
Figure 4: Captions recover more of the video score than a single frame. Each column is a benchmark, sorted by its caption gap, with gaps under no video, one frame, and captions.
Figure 5: Much of the score survives temporal perturbation. Each dot is a benchmark, placed by the share of full-video accuracy that each perturbation retains.
Figure 6: Newer benchmarks break earlier. a , Benchmarks that reach and break at each level. b , Breaking levels by release year, with the share breaking before visual input in red.
Figure 7: Each capability group breaks at its own levels. Bars split the benchmarks of each group by breaking level, and segment labels give counts.
Figure 8: Eleven benchmarks gain more from pixels and three from frames. The curve is the density of pixel gain minus frame gain, and dots mark survivors with both ladders.
Figure 9: Most long videos gain from larger budgets. a , Scores per budget with means and quartiles, blue if the reference still gains at 600 frames. b , The six largest gains, 512 to 1024.
Figure 10: Errors concentrate in few capabilities. Columns are the 76 pool-level survivors with 20 or more agreed attributions, grouped by their first breaking level, each summing to 100%.
Figure 11: Near-duplicates shrink effective size. a , Effective items per 200, red where at least half are near-duplicates. b , The largest near-duplicate shares across benchmarks.
Figure 12: Agents differ in accuracy, runtime, and inspection. a , Accuracy and median runtime, Q1–Q3. b , Operations per answer. c , Image medians by duration on shared log axes.
\toprule Model
Video
Blind
Gain
Percep.
Temp.
Spatial
Reason.
\midrule VideoChat3-4B ( li2026videochat3 )
9.4
3.8
5.7
9.5
12.4
6.2
9.5
Qwen3-VL-8B ( bai2025qwen3 )
9.5
1.4
8.1
9.0
11.4
6.7
11.0
Cambrian-S-7B ( yang2026cambrian )
11.3
5.3
6.0
12.4
10.0
3.8
19.0
InternVideo2.5-8B ( wang2025internvideo2 )
12.4
7.5
4.9
12.9
11.0
8.1
17.6
VideoLLaMA3-7B ( zhang2025videollama )
12.4
5.2
7.1
11.9
18.6
8.6
10.5
Video-XL-2 ( qin2025video )
12.4
7.3
5.1
12.4
15.7
9.0
12.4
Table 1: Agents lead open models on Video-Index. Accuracy in %, fixed input then agent tools. Blind sees text only and Gain is Video minus Blind. Highest in bold. Protocols in the appendix.
The inherent complexity of video understanding makes it difficult to determine whether Video-LLM benchmark performance stems from visual perception, linguistic reasoning, or knowledge priors. While many benchmarks have emerged to assess high-level reasoning, shared criteria for evaluating video understanding remain largely overlooked. Instead of introducing yet another benchmark, we take a step back to re-examine the criteria for evaluating video understanding. In this work, we introduce Video-Oasis, a sustainable diagnostic suite for systematically auditing existing video understanding benchmarks. This audit reveals that 55% of existing benchmark samples are solvable without visual input or temporal context. After filtering these shortcuts, the remaining video-native challenges expose a substantial capability gap: state-of-the-art models perform only marginally above random guessing. Building on these findings, we use the distilled challenges as a testbed to investigate which algorithmic design choices contribute to robust video understanding. We hope our work provides a practical foundation for constructing rigorous video benchmarks and evaluating future Video-LLMs. Code is available at https://github.com/sejong-rcv/Video-Oasis.
Geuntaek Lim, Sungjune Park, Jaeyun Lee +5
Sejong University, South Korea · NAVER Cloud, South Korea
Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The usefulness of such models have also been demonstrated on a wide range of benchmarks, with an important caveat - the dominant approach in these benchmarks evaluates multiple choice reasoning via text options. This is a natural way to test text-based reasoning in these models, and has led to significant insights regarding model behavior in the community. In this work, we ask a different question - what happens when the evaluation modality is visual, rather than text? We introduce three new vision-centric evaluation benchmarks in temporal frame retrieval, video future prediction, and causal memory distortion, all designed around evaluating visual understanding capabilities in large video models. Our approach complements the existing approaches to evaluate video understanding in frontier models. We show that current frontier models exhibit significant weakness when attempting to reason through visual queries, rather than text. We conclude with an extended analysis section that provides pointers for future improvements in visual understanding for large video models.
Rwiddhi Chakraborty, Yinong, Wang +7
Oliver · University of Copenhagen · Carnegie Mellon University +4
Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide limited evidence for their scores. We introduce VideoArgus, a unified rubric-grounded framework covering five video generation and editing settings. For each input instance, VideoArgus generates an output-blind, sample-specific rubric once and reuses it to evaluate all corresponding candidate videos. The rubric defines concrete criteria, scoring rules, failure modes, and evidence plans, which guide criterion-specific VLM QA and visual tools to produce evidence-grounded criterion scores, rationales, and a diagnostic report. We further construct VideoArgus-Bench, containing 1,026 curated input instances built from 653 high-quality images and 416 high-quality videos, with all benchmark rubrics pre-generated, frozen, and released. On a separate 1,260-video human-alignment set, VideoArgus achieves higher within-input Spearman and Kendall correlations with human judgments than the corresponding benchmark-specific evaluators across all five tasks. Model rankings also remain largely consistent across different rubric-generation and evaluation-VLM backbones. All code and data are released. Visit our project page: https://zzzmyyzeng.github.io/VideoArgus