cs.CVApr 2, 2026

VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification

Authors: Jiahao Meng, Yue Tan, Qi Xu, Haochen Wang, Zhongwei Ren, Weisong Liu, Yuhao Wang, Renrui Zhang, +4 more

Organizations: PKU · WHU · CASIA · BJTU · CUHK

Abstract

Video multimodal large language models achieve strong results on existing benchmarks, but answer accuracy alone does not establish whether they can locate the evidence needed to answer a question. We introduce VideoZeroBench, a challenging long-video benchmark with manually annotated question-answer pairs spanning 13 video domains. Questions target fine-grained cues, fleeting events, and evidence distributed across multiple segments. Temporal intervals and key-frame boxes are annotated where applicable. All questions undergo two rounds of cross-verification for answer validity and evidence quality. Our five-level diagnostic protocol compares answering with and without evidence hints, then combines answer correctness with independently evaluated temporal and spatial grounding. Across 19 evaluated models, the best standard QA accuracy is 24.8% (Level-3), achieved by Gemini-3.7-Flash. No model exceeds 1.8% when correct answers and accurate spatio-temporal localization are jointly required (Level-5). Analyses of atomic abilities, evidence spans, input modalities, and thinking-with-videos inference further characterize where the evaluated systems struggle. These findings motivate more precise evidence search and localization for long-video question answering. Our code and data are publicly released.

Explore similar work

CardsList
  1. EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence

    Jun 23, 2026Linpeng Huang, Weixing Chen, Zexin Chen +2Long Video Question AnsweringTemporal Grounding

  2. Evidence-Backed Video Question Answering

    Jul 13, 2026Shijie Wang, Honglu Zhou, Ziyang Wang +5Long Video Question AnsweringVideo Understanding

  3. Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

    Jun 1, 2026Xiaolin Liu, Yilun Zhu, Xiangyu Zhao +9Video Multimodal Large Language ModelsMultimodal Large Language Models