Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single input and perform direct inference, which introduces training-inference mismatch, information loss from frame compression, and a lack of explicit cross-video coordination. Meanwhile, current multi-video benchmarks primarily emphasize event-level comparison, leaving identity-level matching, fine-grained discrimination, and structured multi-step reasoning underexplored. To address these gaps, we introduce MVX-Bench, a Multi-Video Cross-Dimension Benchmark that reformulates 11 classical computer vision tasks into a unified multi-video question-answering framework, comprising 1,442 questions over 4,255 videos from diverse real-world datasets. We further propose SAMA, a Skill-Augmented Agentic Framework for Multi-Video Understanding, which integrates visual tools, task-specific skills, and a conflict-aware verification mechanism to enable iterative and structured reasoning. Experimental results show that SAMA outperforms strong open-source baselines and GPT on MVX-Bench, and ablations validate the effectiveness of skill design and conflict resolution.
Figures & tables
Figure 1: Statistics of our benchmark. The benchmark includes 1,442 samples across 11 tasks, spanning low-level perception, mid-level cross-video matching, and high-level reasoning.
Figure 2: Overview of MVX-Bench. MVX-Bench evaluates 11 complementary capabilities spanning low-, mid-, and high-level video understanding and reasoning. Representative examples illustrate tasks ranging from perceptual comparison and cross-video matching to temporal, spatial, counterfactual, and defeasible reasoning.
Figure 3: Overview of the SAMA framework. A text-only LLM planner, guided by task-adaptive skills that define when, how, and in what order to invoke tools, orchestrates six visual tools across three groups: Perception, Detection, and Other tools. Skills determine the invocation strategy based on task type — for example, similarity tasks prioritize Visual Similarity before Video Reader, while counting tasks consult both Video Reader and Scene Graph to enable cross-modal conflict detection. Detected conflicts trigger adaptive re-reading, feeding corrective information back to the planner before the final answer is produced.
Configuration
Acc.
SAMA
60.0%
w/o-Skills
51.7%
w/o-Conflict
55.9%
w/o-Evidence Aggregation
57.3%
Table 2: Ablation studies of our SAMA.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Case study of the conflict detection mechanism.
Figure 5: Case study of skill-guided tool selection.
Planner
SAMA Acc.
End-to-end Acc.
GPT-4o
51.7%
45.3%
GPT-5.2
58.7%
48.3%
DeepSeek-V4-Pro
52.3%
N/A
Appendix
Table 3: Planner ablation on the 300-sample subset. Video Reader is fixed to Qwen2.5-VL-7B.
Video Reader
SAMA Acc.
End-to-end Acc.
Qwen2.5-VL-7B
58.7%
48.3%
InternVL3.5-8B
59.3%
41.0%
Appendix
Table 4: Video Reader ablation on the 300-sample subset. Planner is fixed to GPT-5.2.
Method
Acc.
VideoAgent + Boundary-Aware Concatenation
30.6%
SAMA
60.0%
Appendix
Table 5: Comparison with a boundary-aware adaptation of VideoAgent under the same evaluation protocol.
Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded perception context budgets. Existing agentic methods mitigate this via rule-based preprocessing, yet often suffer from information loss, high cost, and reliance on textual intermediates. We propose MACF, an end-to-end Multi-Agent Collaboration Framework that decouples per-agent perception budgets from global video complexity, enabling scalable video understanding while preserving visual fidelity. MACF partitions videos into segments for locally budgeted agents and enables holistic reasoning via an agent-native latent communication protocol. Each agent encodes partial observations into compact, task-sufficient tokens in a shared embedding space, allowing efficient and information-preserving collaboration by a central coordinator. We introduce a curriculum training strategy that progressively enforces semantic alignment, evidence summarization, and cross-agent coordination. Extensive experiments on diverse video understanding benchmarks show that MACF consistently outperforms state-of-the-art MLLMs and multi-agent systems under identical budget constraints, demonstrating the effectiveness of our latent collaboration for scalable video understanding.
Kerui Chen, Jinglu Wang, Jianrong Zhang +3
Microsoft Research Asia · Zhejiang University · Guanming Lab.
While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over complex narratives remains poorly evaluated. Existing summarization benchmarks fragment supervision across isolated granularities, such as keyframes, key shots, or disjointed text summaries, failing to capture the inherently hierarchical structure of cross-modal alignment. To address this critical gap, we introduce HAVEN, a hierarchically aligned multimodal benchmark for unified video understanding. HAVEN pioneers a fully granular (frame, shot, and video levels) and fully multimodal (video and text) dataset architecture, complete with explicit, continuous alignment between modalities. Built upon this unified annotation paradigm, we propose a comprehensive evaluation suite spanning summarization, temporal reasoning, multimodal grounding, and saliency ranking. Extensive benchmarking of state-of-the-art MLLMs exposes a persistent gap between surface-level textual fluency and grounded multimodal understanding. Ultimately, HAVEN advances the evaluation of multimodal systems beyond traditional QA formats, offering a rigorous, standardized testbed to drive future research in interpretable, hierarchical video understanding. We publicly release the dataset, benchmark suite, and evaluation protocols.
Mengqi Shi, Haopeng Zhang
Department of Information and Computer Sciences · University of Hawaii at Manoa
Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent video streams remains poorly understood. Existing multi-video benchmarks rely largely on human-annotated real-world footage, limiting the precision of spatial, temporal, and physical ground truth and making it difficult to diagnose model failures. We introduce SYNCR, a controlled synthetic benchmark for cross-video reasoning with programmatically verified grounding. Built using Habitat, Kubric, and CLEVRER simulator engines, SYNCR contains 4,000 multi-video question-answer pairs grounded in 4,827 unique videos. It evaluates MLLMs across eight tasks spanning four diagnostic pillars: Temporal Alignment, Spatial Tracking, Comparative Reasoning, and Holistic Synthesis. Our zero-shot evaluation of leading open- and closed-weight MLLMs reveals a substantial gap between current models and humans: the best model achieves only 64.5% average accuracy, compared to an 89.5% human baseline. Models perform relatively well on temporal ordering but struggle with precise physical and spatial reasoning, with the best model reaching only 29.8% accuracy on Kinematic Comparison. We further find that parameter scaling and reasoning-specialized post-training improve temporal alignment capabilities, but do not reliably address fine-grained physical tracking or global spatial synthesis. Finally, a sim-to-real correlation analysis suggests that SYNCR tracks model-level trends on a real-world multi-video benchmark.
Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy +1