Video Understanding

Latest papers 69

All topics
CardsList
  1. Video-Index: A Curated Meta-Benchmark for Video Understanding

    Oct 1, 2026Enxin Song, Yinuo Xu, Shusheng Yang +2Video UnderstandingVideo-Language Model Evaluation

  2. CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding

    Sep 30, 2026Yiduo Jia, Muzhi Zhu, Jinchuan Shi +4Temporal Video GroundingVideo Understanding

  3. CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

    Sep 30, 2026Yulong Liu, Xiaotian Han, Junyuan Shang +6Video UnderstandingEfficient VLM Inference

  4. Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution

    Sep 29, 2026Bingjun Luo, Jialin Guo, Siqi LiVideo UnderstandingAgent Harness Optimization

  5. Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark

    Sep 29, 2026Yuedong Tan, Lei Qi, Yu Liu +13Video UnderstandingEgocentric Video QA

  6. Rethinking Visual Token Compression for Video Large Language Models: A Simple Yet Strong Baseline

    Sep 28, 2026Xiao Zhang, Wang Zeng, Sheng Jin +3Video UnderstandingVideo-Language Models

  7. CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

    Sep 23, 2026Shuo Xing, Pooja Verlani, Balu Adsumilli +1Video UnderstandingVideo-Language Model Evaluation

  8. MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos

    Sep 23, 2026Reno Kriz, David Etter, Alexander Martin +11Video UnderstandingVideo Retrieval

  9. Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows

    Sep 23, 2026Xiyuan Shen, Jiuyang Lyu, Seokhyun Hwang +4Video UnderstandingVideo-Language Model Evaluation

  10. Label-Guided Knowledge Distillation for 3D-CNNs in Action Recognition

    Sep 14, 2026Yanjiang Shi, Peng Zhao, Nan Qi +1Video UnderstandingHuman Activity Recognition

  11. Zero-shot video highlight detection based on text descriptions and synthetic images

    Sep 13, 2026Michal Byra, Alberto Presta, Grzegorz Stefanski +1Video UnderstandingTemporal Video Understanding

  12. Video-MOPD: Multi-Teacher On-Policy Distillation for Video Understanding

    Sep 8, 2026Zhenxin Qin, Peng Shi, Cong Han +3Temporal Video GroundingVideo Understanding

  13. Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

    Sep 3, 2026Hongyu Qu, Guangming Yao, Ling Xing +7Memory-Augmented VLMsVideo Understanding

  14. KnowVis: Knowledge-Centric Visual Summarization for Video Lectures

    Sep 3, 2026Yi Xu, Yifan Hou, Xiaoyu ZhangVideo UnderstandingVideo Summarization

  15. DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos

    Aug 30, 2026Bomiao Wang, Zekai Shao, Jiexiang Lan +3Video UnderstandingVideo-Language Model Evaluation

  16. NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

    Aug 13, 2026Yuheng Huang, Jianlang Chen, Jiayang Song +6Video UnderstandingVideo-Language Model Evaluation

  17. DAEP: Difficulty-Aware Evidence Planning for Medical Video Corpus Temporal Answer Grounding

    Aug 7, 2026Tianjian He, Yujie Liu, Zhiping Huang +1Temporal Video GroundingVideo Understanding

  18. QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

    Jul 27, 2026Arian Kheirandish, Fardin Ayar, Ehsan Javanmardi +2Video UnderstandingInstance Segmentation

  19. MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

    Jul 27, 2026Tuan-An To, Yuk-Kwan Wong, Tuan-Anh Vu +2Visual Question AnsweringVideo Understanding

  20. T-STAR: A Large-Scale Benchmark for Spatio-Temporal Panoptic Scene Graph Generation in Satellite Video

    Jul 23, 2026Linlin Wang, Xue Yang, Zhihuang Zhou +3Video UnderstandingScene Graph Generation

  21. VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

    Jul 16, 2026Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24Video UnderstandingStreaming Video Understanding

  22. VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

    Jul 16, 2026Yunfeng Liu, Yuandong Yang, Jiarui Han +5VLM EvaluationVideo Understanding

  23. SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments

    Jul 6, 2026Suryanarayana Reddy Yarrabothula, Manisha Chawla, Kunal Sinha +5VLM EvaluationVideo Understanding

  24. K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos

    Jul 2, 2026Khush Attarde, Yusuf Ali, Megha Thukral +3Video UnderstandingMultimodal Large Language Models

  25. LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

    Jul 2, 2026Shunya Kato, Taiki Miyanishi, Shuhei Kurita +3Temporal Video GroundingVideo Understanding

  26. MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

    Jul 2, 2026Yuan Wang, Shujian Gao, Songtao Jiang +2Video UnderstandingStreaming Video Understanding

  27. Learning to Watch: Active Video Anomaly Understanding via Interleaved Policy Optimization

    Jul 1, 2026Mengjingcheng Mo, Jiaxu Leng, Xinbo GaoVideo UnderstandingDirect Preference Optimization

  28. DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

    Jun 26, 2026Yankai Yang, Yancheng Long, Bin Wen +4Video UnderstandingVideo-Language Models

  29. HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

    Jun 25, 2026Jiajun Wu, Haoyu Kang, Yining Sun +13Video UnderstandingVideo Content Moderation

  30. P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

    Jun 22, 2026Felix Tristram, Stefano Gasperini, Benjamin Killeen +4Temporal Action SegmentationVideo Understanding

  31. HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning

    Jun 19, 2026Awais Rauf, Ahmed Hasssan, Greg SlabaughVideo UnderstandingLong-Context Modeling

  32. SA-VIS: Sparse frame Annotations for training Video Instance Segmentation

    Jun 18, 2026Edoardo Mello Rella, Ajad Chhatkuli, Shipra Jain +2Video UnderstandingVideo Object Segmentation

  33. Interpretable Temporal Facial-Region Motion Analysis for In-the-Wild Parkinson's Disease Video Classification

    Jun 8, 2026Riyadh AlmushrafyVideo UnderstandingMedical Diagnosis

  34. Momentum-Guided Semantic Forecasting (MoFore) for Self-Supervised Video Representation Learning

    Jun 8, 2026Qinwu XuVideo UnderstandingVideo Representation Learning

  35. MAVIS: Multi-Agent Video Retrieval via Structured Video Understanding

    Jun 8, 2026Jie Zhang, Qilang Ye, Hao Zhou +2Video UnderstandingMulti-Agent Collaboration

  36. CACR:Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning

    Jun 7, 2026Muge Qi, Rong Fu, Pengbin Feng +7Temporal Video GroundingVideo Understanding

  37. Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation

    Jun 5, 2026Danial Hamdi, Fardin Ayar, Mahdi JavanmardiVideo UnderstandingInstance Segmentation

  38. StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

    Jun 4, 2026Zhengqian Wu, Zhixian Liu, Aodong Chen +6Video UnderstandingSynthetic Data Generation

  39. VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

    Jun 3, 2026Lin Fu, Zheyuan Yang, Yang Wang +3Video UnderstandingVideo QA

  40. VCIFBench: Evaluating Complex Instruction Following for Video Understanding

    Jun 3, 2026Huangchen Xu, Yuan Wu, Yi ChangVLM EvaluationVideo Understanding

  41. TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

    May 29, 2026Yeil Jeong, Youngjin Yoo, Jiyoung Bae +5Video UnderstandingAI in Education

  42. EarlyTom: Early Token Compression Completes Fast Video Understanding

    May 28, 2026Hesong Wang, Xin Jin, Lu Lu +4Video UnderstandingEfficient VLM Inference

  43. ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement

    May 28, 2026Jianping Ye, Michel WedelVideo UnderstandingOnline Advertising

  44. MetaphorVU: Towards Metaphorical Video Understanding

    May 25, 2026Zhuoqun Li, Boxi Cao, Guiping Jiang +13VLM EvaluationFigurative Language Understanding

  45. VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

    May 21, 2026Haichen He, Jiayi Zhou, Sifeng Shang +3VLM EvaluationVideo Understanding

  46. Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

    May 21, 2026Lina Zhang, Tonmoy Monsoor, Peizheng Li +23Video UnderstandingMultimodal Large Language Models

  47. RECIPE: Procedural Planning via Grounding in Instructional Video

    May 19, 2026Luigi Seminara, Antonino Furnari, Lorenzo TorresaniVideo Understanding

  48. ViMU: Benchmarking Video Metaphorical Understanding

    May 14, 2026Qi Li, Xinchao WangVideo UnderstandingVideo QA

  49. BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

    May 12, 2026Patrick Knab, Orgest Xhelili, Inis Buzi +7Video UnderstandingScene Graph Generation

  50. Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization

    May 10, 2026Omer Tariq, Syed Muhammad Raza, Jeongbae SonVideo UnderstandingUncertainty Quantification

  51. Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

    May 8, 2026Jiazheng Li, Chi-Hao Wu, Yunze Liu +3Video UnderstandingMultimodal Memory

  52. Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology

    May 5, 2026Lina Zhang, Tonmoy Monsoor, Mehmet Efe Lorasdagi +8Video UnderstandingMultimodal Large Language Models

  53. HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding

    May 4, 2026Haopeng Jin, Hongzhu Yi, Wenlong Zhao +6Video UnderstandingVideo-Language Models

  54. OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

    May 2, 2026Detao Bai, Shimin Yao, Weixuan Chen +4Cross-Modal LearningVideo Understanding

  55. SF20K Competition 2025: Summary and findings

    May 2, 2026Ridouane Ghermi, Xi Wang, Vicky Kalogeiton +1VLM EvaluationVideo Understanding

  56. VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

    May 2, 2026Alejandro Aparcedo, Akash Kumar, Aaryan Garg +5VLM EvaluationVideo Understanding

  57. High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions

    May 1, 2026Yongpeng Cao, Yuji YamakawaVideo UnderstandingZero-Shot Learning

  58. DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation

    Apr 29, 2026Mingji Ge, Qirui Chen, Zeqian Li +1Video UnderstandingLLM-Assisted Annotation

  59. Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations

    Apr 28, 2026Chen Liang, Xirui Jiang, Naihao Deng +2VLM EvaluationVideo Understanding