cs.CVMar 16, 2026

A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding

Authors: Yue Zhang, Liqiang Jing, Jia Li, Yapeng Tian, Xinya Du, Yunhui Guo, Vibhav Gogate

Organizations: University of Texas at Dallas, USA

Abstract

Multimodal Large Language Models have achieved strong performance in single-video understanding, yet their ability to reason across multiple videos remains limited. Existing approaches typically concatenate multiple videos into a single input and perform direct inference, which introduces training-inference mismatch, information loss from frame compression, and a lack of explicit cross-video coordination. Meanwhile, current multi-video benchmarks primarily emphasize event-level comparison, leaving identity-level matching, fine-grained discrimination, and structured multi-step reasoning underexplored. To address these gaps, we introduce MVX-Bench, a Multi-Video Cross-Dimension Benchmark that reformulates 11 classical computer vision tasks into a unified multi-video question-answering framework, comprising 1,442 questions over 4,255 videos from diverse real-world datasets. We further propose SAMA, a Skill-Augmented Agentic Framework for Multi-Video Understanding, which integrates visual tools, task-specific skills, and a conflict-aware verification mechanism to enable iterative and structured reasoning. Experimental results show that SAMA outperforms strong open-source baselines and GPT on MVX-Bench, and ablations validate the effectiveness of skill design and conflict resolution.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

    May 1, 2026Kerui Chen, Jinglu Wang, Jianrong Zhang +3Fine-Grained Video UnderstandingMultimodal Large Language Models

  2. HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

    May 19, 2026Mengqi Shi, Haopeng ZhangMultimodal UnderstandingMultimodal Benchmarks

  3. SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

    May 8, 2026Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy +1Video Multimodal Large Language ModelsLong-Video Benchmarks