Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single input and perform direct inference, which introduces training-inference mismatch, information loss from frame compression, and a lack of explicit cross-video coordination. Meanwhile, current multi-video benchmarks primarily emphasize event-level comparison, leaving identity-level matching, fine-grained discrimination, and structured multi-step reasoning underexplored. To address these gaps, we introduce MVX-Bench, a Multi-Video Cross-Dimension Benchmark that reformulates 11 classical computer vision tasks into a unified multi-video question-answering framework, comprising 1,442 questions over 4,255 videos from diverse real-world datasets. We further propose SAMA, a Skill-Augmented Agentic Framework for Multi-Video Understanding, which integrates visual tools, task-specific skills, and a conflict-aware verification mechanism to enable iterative and structured reasoning. Experimental results show that SAMA outperforms strong open-source baselines and GPT on MVX-Bench, and ablations validate the effectiveness of skill design and conflict resolution.
Figures & tables
Figure 1: Statistics of our benchmark. The benchmark includes 1,442 samples across 11 tasks, spanning low-level perception, mid-level cross-video matching, and high-level reasoning.
Figure 2: Overview of MVX-Bench. MVX-Bench evaluates 11 complementary capabilities spanning low-, mid-, and high-level video understanding and reasoning. Representative examples illustrate tasks ranging from perceptual comparison and cross-video matching to temporal, spatial, counterfactual, and defeasible reasoning.
Figure 3: Overview of the SAMA framework. A text-only LLM planner, guided by task-adaptive skills that define when, how, and in what order to invoke tools, orchestrates six visual tools across three groups: Perception, Detection, and Other tools. Skills determine the invocation strategy based on task type — for example, similarity tasks prioritize Visual Similarity before Video Reader, while counting tasks consult both Video Reader and Scene Graph to enable cross-modal conflict detection. Detected conflicts trigger adaptive re-reading, feeding corrective information back to the planner before the final answer is produced.
Configuration
Acc.
SAMA
60.0%
w/o-Skills
51.7%
w/o-Conflict
55.9%
w/o-Evidence Aggregation
57.3%
Table 2: Ablation studies of our SAMA.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Case study of the conflict detection mechanism.
Figure 5: Case study of skill-guided tool selection.
Planner
SAMA Acc.
End-to-end Acc.
GPT-4o
51.7%
45.3%
GPT-5.2
58.7%
48.3%
DeepSeek-V4-Pro
52.3%
N/A
Appendix
Table 3: Planner ablation on the 300-sample subset. Video Reader is fixed to Qwen2.5-VL-7B.
Video Reader
SAMA Acc.
End-to-end Acc.
Qwen2.5-VL-7B
58.7%
48.3%
InternVL3.5-8B
59.3%
41.0%
Appendix
Table 4: Video Reader ablation on the 300-sample subset. Planner is fixed to GPT-5.2.
Method
Acc.
VideoAgent + Boundary-Aware Concatenation
30.6%
SAMA
60.0%
Appendix
Table 5: Comparison with a boundary-aware adaptation of VideoAgent under the same evaluation protocol.