Temporal Video Understanding

Latest papers 220

All topics
CardsList
  1. VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

    May 21, 2026Jinho Park, Youbin Kim, Hogun Park +1VLM EvaluationSpatiotemporal Reasoning

  2. GALAR-TemporalNet v2: Anatomy-Guided Dual-Branch Temporal Classification with Bidirectional Mamba and Dual-Graph GCN for Video Capsule Endoscopy -- after competition results

    May 21, 2026Jiye Won, Seangmin Lee, Soon Ki JungEndoscopyMedical Imaging

  3. Zero-Shot Temporal Action Localization Through Textual Guidance

    May 21, 2026Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero +3Zero-Shot LearningTemporal Action Detection

  4. Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

    May 21, 2026Tomaso Trinci, Henrique Piñeiro Monteagudo, Leonardo TaccariMultimodal Large Language ModelsAutonomous Driving

  5. ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs

    May 21, 2026Bingjun Luo, Tony Wang, Chaoqi Chen +1Multimodal Token CompressionVideo Token Compression

  6. VISTA: Validation-Guided Integration of Spatial and Temporal Foundation Models with Anatomical Decoding for Rare-Pathology VCE Event Detection -- after competition results

    May 21, 2026Bo-Cheng Qiu, Fang-Ying Lin, Ming-Han Sun +3EndoscopyMedical Imaging

  7. Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding

    May 21, 2026Bingjun Luo, Tony Wang, Hanqi Chen +1Video-Language ModelsMultimodal Token Compression

  8. Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

    May 21, 2026Dazhao Du, Jian Liu, Jialong Qin +7Counterfactual EvaluationVideo-Language Models

  9. Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

    May 20, 2026Aditya Chetan, Eric Cai, Peeyush Kushwaha +5VLM EvaluationTemporal Video Grounding

  10. TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos

    May 20, 2026Yakun Yu, Ashley Wiens, Adrián Barahona-Ríos +4VLM EvaluationAutomated Software Testing

  11. VSCD: Video-based Scene Change Detection in Unaligned Scenes

    May 20, 2026Jiae Yoon, Ue-Hwan KimTemporal Video Understanding

  12. HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

    May 19, 2026Mengqi Shi, Haopeng ZhangMultimodal GroundingVideo Summarization

  13. Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models

    May 18, 2026Svetlana Orlova, Niccolò Cavagnero, Gijs DubbelmanStreaming Video UnderstandingVision Foundation Model Adaptation

  14. Visual Timelines of Police Encounters in Body-Worn Camera Footage: Operational Context and Activity Cataloging for Training and Analysis in OpenBWC

    May 16, 2026Angela Srbinovska, Christopher Homan, Adrian Martin +1Human Activity RecognitionVideo Action Recognition

  15. LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection

    May 14, 2026Mitchell Piehl, Muchao YeInterpretable Anomaly DetectionVLM Reasoning

  16. SurgicalMamba: Dual-Path SSD with State Regramming for Online Surgical Phase Recognition

    May 14, 2026Sukju Oh, Sukkyu SunSurgical Video UnderstandingVisual SSMs

  17. Video-Zero: Self-Evolution Video Understanding

    May 14, 2026Ruixu Zhang, Deyi Ji, Lanyun Zhu +4Temporal Video GroundingLong-Video Understanding

  18. BARISTA: A Multi-Task Egocentric Benchmark for Compositional Visual Understanding

    May 12, 2026Patrick Knab, Orgest Xhelili, Inis Buzi +7Video UnderstandingScene Graph Generation

  19. Stabilizing Temporal Inference Dynamics for Online Surgical Phase Recognition

    May 11, 2026Yang Liu, Ning Zhu, Jingjing Peng +4Surgical Video UnderstandingTemporal Consistency

  20. StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video

    May 11, 2026Ao Li, Zihan Xiao, Zihao Yue +7Streaming Video UnderstandingTemporal Video Understanding

  21. OZ-TAL: Online Zero-Shot Temporal Action Localization

    May 11, 2026Chaolei Han, Hongsong Wang, Xin Gong +1Zero-Shot LearningTemporal Action Detection

  22. Fetal Brain Imaging: A Composite Neural Network Approach for Keyframe Detection in Ultrasound Videos

    May 10, 2026Aleksander Zamojski, Kacper Jarczak, Radoslaw RoszczykUltrasound ImagingVideo Frame Selection

  23. SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

    May 10, 2026Ismael Elsharkawi, Ahmed Sait, Silvio Giancola +3VLM EvaluationSports Video Analysis

  24. Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs

    May 10, 2026Yigui Feng, Qinglin Wang, Yang Liu +1Efficient VLM InferenceVisual Tokenization

  25. Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models

    May 9, 2026Tri Cao, Khoi Le, Thong Nguyen +7VLM EvaluationObject Hallucination in VLMs

  26. MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment

    May 9, 2026Qiqi Li, Pengfei Wang, Hongyu Chen +1Multimodal FusionAction Quality Assessment

  27. TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

    May 8, 2026Hengyi Feng, Hao Liang, Mingrui Chen +6Audio-Visual UnderstandingMultimodal Hallucination

  28. Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs

    May 8, 2026Peitao Han, Fei Cheng, Lis K. Pereira +2Vision-Language ModelsTemporal Reasoning in Language Models

  29. Jointly Learning Structured Representations and Stabilized Affinity for Human Motion Segmentation

    May 7, 2026Xianghan Meng, Zhiyuan Huang, Zhengyu Tong +1Representation LearningClustering