Temporal Video Understanding

Latest papers 220

All topics
CardsList
  1. SpaTime: Streaming Vision-Language Models for Spatio-temporal Reasoning

    Oct 6, 2026Hairong Yin, Huangying Zhan, Shin-Fang Chng +2Vision-Language ModelsSpatiotemporal Reasoning

  2. EC-RAG: Event Chain Retrieval-Augmented Generation for Long Video Understanding

    Oct 6, 2026Yuhao Qin, Junbo Wang, Yuke Li +1Multimodal RAGVideo-Language Models

  3. Multi-Label Perceptual Bug Detection in Video Games using Deep Learning on Gameplay Footage

    Oct 6, 2026Nahian Rifaat, Felix Morosov, Loutfouz ZamanAutomated Software TestingMulti-Label Classification

  4. Joint Class-Time Learning for Video Classification with Multi-Instance Partial-Label Learning

    Oct 5, 2026Lingyu Shen, Wei Tang, Fakhri Karray +1Weakly Supervised LearningMultiple Instance Learning

  5. Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

    Oct 1, 2026Mohd Ubaid Wani, Sara Atito, Josef Kittler +1CoT ReasoningInterpretable Anomaly Detection

  6. Video-Index: A Curated Meta-Benchmark for Video Understanding

    Oct 1, 2026Enxin Song, Yinuo Xu, Shusheng Yang +2Video UnderstandingVideo-Language Model Evaluation

  7. RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling

    Sep 30, 2026Can Zhang, Xiaotian Han, Junyuan Shang +5Video-Language ModelsVideo Representation Learning

  8. TSMD: Temporal-Stream Modality Dropout for Robust Video Highlight Detection

    Sep 30, 2026Bo-Yuan Cheng, Kuan-Yu Chen, Po-Han Huang +2Missing-Modality LearningMultimodal Robustness

  9. SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation

    Sep 29, 2026Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy +1Video-Language Model EvaluationVideo Reasoning

  10. EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding

    Sep 29, 2026Yuwei Miao, Xuesheng Zhang, Wenhao Zou +5Streaming Video UnderstandingSelf-Distillation

  11. RoboChrono: A Real Robot Benchmark for Streaming Task Understanding

    Sep 29, 2026Yuzhou Wu, Longteng Fan, Zimeng Li +22Next-Action PredictionRobot Manipulation Benchmarks

  12. Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

    Sep 29, 2026Estela Monserrat Arriaga Santana, Julian Rosas Scull, Ehécatl Sacamch'en Núñez Rico +1Physical Consistency in Video GenerationWorld Model Evaluation

  13. ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering

    Sep 28, 2026Zhen Yao, Likai Wang, Yuming Yang +9Remote Sensing VQAVideo QA

  14. ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models

    Sep 28, 2026Gueter Josmy Faure, Min-Hung Chen, Hao Ping Wang +3Video-Language Model EvaluationTemporal Video Understanding

  15. A Hierarchy-Aware Video-Language Model Evaluation and Hyperbolic Baseline for Surgery

    Sep 22, 2026Ana Manzano Rodríguez, Pascal Mettes, Marlies P. Schijven +1Video-Language Model EvaluationHierarchical Representation Learning

  16. Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models

    Sep 15, 2026Rwiddhi Chakraborty, Yinong, Wang +7Video-Language Model EvaluationVideo QA

  17. BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

    Sep 14, 2026Yolo Y. Tang, Daiki Shimada, Jiayue Meng +14AI Agent BenchmarksTemporal Video Understanding

  18. CapsuleMotion: A Lightweight Real-Time Visual Motion Predictor for Capsule Endoscopy

    Sep 14, 2026Oliver Bause, Julia Werner, Oliver BringmannEndoscopyMedical Imaging

  19. OphBiWSSD: Scaling Temporal Action Localization in Ophthalmic Surgeries with Bidirectional Weight-tied State Space Duality

    Sep 14, 2026Yang Liu, Qionghong Ma, Joongwon Chae +10Surgical Video UnderstandingTemporal Action Detection

  20. Zero-shot video highlight detection based on text descriptions and synthetic images

    Sep 13, 2026Michal Byra, Alberto Presta, Grzegorz Stefanski +1Video UnderstandingTemporal Video Understanding

  21. LiveProBench: Can Streaming Video Models Really Interact Like Humans?

    Sep 11, 2026Kaixuan Du, Xin Wan, Hang Zhang +4Streaming Video UnderstandingVideo-Language Model Evaluation

  22. MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

    Sep 8, 2026Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche +7Video-Language Model EvaluationVideo-Language Models

  23. EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning

    Sep 8, 2026Jingpu Yang, Fengxian Ji, Mingxuan Cui +4Visual Spatial ReasoningEgocentric Video QA

  24. Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

    Sep 8, 2026Ruibo Ming, Lei Sun, Deheng Zhang +8Video-Language Model EvaluationVideo-Language Models

  25. STSG-VQA: Evidence-Grounded Temporal Question Answering from Surgical Spatio-Temporal Scene Graphs

    Sep 8, 2026Jing Li, Duygu SarikayaMedical VQAQuestion Answering

  26. Temporal Self-Distillation: Learning Visual State Tracking in Videos Without Supervision

    Sep 3, 2026Shravan Venkatraman, Wenshuai Zhao, Mohammad Hassan Vali +1Self-Supervised LearningState Tracking

  27. TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrieval

    Sep 2, 2026Uicheol Jung, Juyoung Hong, Hojung Kwon +1Temporal Video UnderstandingMixture of Experts