Video Understanding

Latest papers 69

All topics
CardsList
  1. P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

    Jun 22, 2026Felix Tristram, Stefano Gasperini, Benjamin Killeen +4Temporal Action SegmentationVideo Understanding

  2. HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning

    Jun 19, 2026Awais Rauf, Ahmed Hasssan, Greg SlabaughVideo UnderstandingLong-Context Modeling

  3. SA-VIS: Sparse frame Annotations for training Video Instance Segmentation

    Jun 18, 2026Edoardo Mello Rella, Ajad Chhatkuli, Shipra Jain +2Video UnderstandingVideo Object Segmentation

  4. Interpretable Temporal Facial-Region Motion Analysis for In-the-Wild Parkinson's Disease Video Classification

    Jun 8, 2026Riyadh AlmushrafyVideo UnderstandingMedical Diagnosis

  5. Momentum-Guided Semantic Forecasting (MoFore) for Self-Supervised Video Representation Learning

    Jun 8, 2026Qinwu XuVideo UnderstandingVideo Representation Learning

  6. MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding

    Jun 8, 2026Jie Zhang, Qilang Ye, Hao Zhou +2Video UnderstandingMulti-Agent Collaboration

  7. CACR:Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning

    Jun 7, 2026Muge Qi, Rong Fu, Pengbin Feng +7Temporal Video GroundingVideo Understanding

  8. Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation

    Jun 5, 2026Danial Hamdi, Fardin Ayar, Mahdi JavanmardiVideo UnderstandingInstance Segmentation

  9. StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

    Jun 4, 2026Zhengqian Wu, Zhixian Liu, Aodong Chen +6Video UnderstandingSynthetic Data Generation

  10. VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

    Jun 3, 2026Lin Fu, Zheyuan Yang, Yang Wang +3Video UnderstandingVideo QA

  11. VCIFBench: Evaluating Complex Instruction Following for Video Understanding

    Jun 3, 2026Huangchen Xu, Yuan Wu, Yi ChangVLM EvaluationVideo Understanding

  12. TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

    May 29, 2026Yeil Jeong, Youngjin Yoo, Jiyoung Bae +5Video UnderstandingAI in Education

  13. EarlyTom: Early Token Compression Completes Fast Video Understanding

    May 28, 2026Hesong Wang, Xin Jin, Lu Lu +4Video UnderstandingEfficient VLM Inference

  14. ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement

    May 28, 2026Jianping Ye, Michel WedelVideo UnderstandingOnline Advertising

  15. MetaphorVU: Towards Metaphorical Video Understanding

    May 25, 2026Zhuoqun Li, Boxi Cao, Guiping Jiang +13VLM EvaluationFigurative Language Understanding

  16. VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

    May 21, 2026Haichen He, Jiayi Zhou, Sifeng Shang +3VLM EvaluationVideo Understanding

  17. Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

    May 21, 2026Lina Zhang, Tonmoy Monsoor, Peizheng Li +23Video UnderstandingMultimodal Large Language Models

  18. RECIPE: Procedural Planning via Grounding in Instructional Video

    May 19, 2026Luigi Seminara, Antonino Furnari, Lorenzo TorresaniVideo Understanding

  19. ViMU: Benchmarking Video Metaphorical Understanding

    May 14, 2026Qi Li, Xinchao WangVideo UnderstandingVideo QA

  20. BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

    May 12, 2026Patrick Knab, Orgest Xhelili, Inis Buzi +7Video UnderstandingScene Graph Generation

  21. Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization

    May 10, 2026Omer Tariq, Syed Muhammad Raza, Jeongbae SonVideo UnderstandingUncertainty Quantification

  22. Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

    May 8, 2026Jiazheng Li, Chi-Hao Wu, Yunze Liu +3Video UnderstandingMultimodal Memory

  23. Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology

    May 5, 2026Lina Zhang, Tonmoy Monsoor, Mehmet Efe Lorasdagi +8Video UnderstandingMultimodal Large Language Models

  24. HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding

    May 4, 2026Haopeng Jin, Hongzhu Yi, Wenlong Zhao +6Video UnderstandingVideo-Language Models

  25. OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

    May 2, 2026Detao Bai, Shimin Yao, Weixuan Chen +4Cross-Modal LearningVideo Understanding

  26. SF20K Competition 2025: Summary and findings

    May 2, 2026Ridouane Ghermi, Xi Wang, Vicky Kalogeiton +1VLM EvaluationVideo Understanding

  27. VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

    May 2, 2026Alejandro Aparcedo, Akash Kumar, Aaryan Garg +5VLM EvaluationVideo Understanding

  28. High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions

    May 1, 2026Yongpeng Cao, Yuji YamakawaVideo UnderstandingZero-Shot Learning

  29. DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation

    Apr 29, 2026Mingji Ge, Qirui Chen, Zeqian Li +1Video UnderstandingLLM-Assisted Annotation

  30. Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations

    Apr 28, 2026Chen Liang, Xirui Jiang, Naihao Deng +2VLM EvaluationVideo Understanding