Video Reasoning

Momentum

13 papers in the last four weeks, up 86% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 135

All topics
CardsList
  1. EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

    Jul 29, 2026Yuyun Chen, Tianao Li, TianQuan Feng +4Egocentric Video UnderstandingVideo Reasoning

  2. Visual prompt engineering for video models

    Jul 28, 2026Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7Video ReasoningVisual Prompt Learning

  3. CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

    Jul 27, 2026Jinlong Yang, Wenhao Zhang, Kuanwei Lin +1Tool Use in VLMsEfficient VLM Inference

  4. DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video

    Jul 25, 2026Muhammad Sulthan Adhipradhana, Ehsan Javanmardi, Naren Bao +1Disaster ResponseVideo Reasoning

  5. BasketEvent: Understanding Who Did What and When in Basketball Videos

    Jul 23, 2026Yu Zhang, Jiayuan Rao, Haoning Wu +1Sports Video AnalysisVideo Reasoning

  6. An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia

    Jul 21, 2026Nassira Ait Mehdi, Milissa Temmam, Slimane LarabiVision-Language ModelsMental Health

  7. ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

    Jul 20, 2026Ting Huang, Zhenyu Zhang, Wenyuan Huang +2Memory-Augmented VLMsVisual Spatial Reasoning

  8. Thinking in Video: Can Video Generators Really Reason About the Real World?

    Jul 20, 2026Yongheng Zhang, Guang Yang, Ruihan Hou +12Video ReasoningCausal Reasoning in Videos

  9. Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding

    Jul 13, 2026Kerui Chen, Jinglu Wang, Xiaoyi Zhang +1Active PerceptionSports Video Analysis

  10. TreeSoc: Tree-Structured Dynamic Reasoning and Tool Synergy for Soccer Video Understanding

    Jul 13, 2026Thanh-Nhan Vo, Thanh-Khoi Nguyen, Trong-Thuan Nguyen +2Tool Use in VLMsSports Video Analysis

  11. Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset

    Jul 11, 2026Zhiyan Zhang, Peipei Song, Jinpeng Hu +3Video ReasoningCausal Reasoning in Videos

  12. OpenCoF: Learning to Reason Through Video Generation

    Jul 9, 2026Xinyan Chen, Ziyu Guo, Renrui Zhang +2Video Diffusion ModelsVideo Generation

  13. TimeThink: Reasoning with Time for Video LLMs

    Jul 6, 2026Handong Li, Longteng Guo, Zikang Liu +8Temporal Video GroundingProcess Reward Models

  14. STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation

    Jul 3, 2026Syed Ariff Syed Hesham, Yun Liu, Guolei Sun +4Video ReasoningReasoning Segmentation

  15. Learning to Evolve Scenes: Reasoning about Human Activities with Scene Graphs

    Jul 2, 2026Francesca Pistilli, Simone Alberto Peirone, Giuseppe AvertaScene Graph GenerationEgocentric Video Understanding

  16. Latent Visual Cache for Video Reasoning

    Jul 1, 2026Yongheng Zhang, Zhipeng Xu, Hao Wu +4Memory-Augmented VLMsVision-Language Models

  17. EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection

    Jul 1, 2026Wenhao Zhang, Kuanwei Lin, Xuyi Yang +2Long-Video UnderstandingVideo Reasoning

  18. Linguistic Relative Policy Optimization for Video Anomaly Reasoning

    Jul 1, 2026Jiaxu Leng, Jiankang Zheng, Mengjingcheng Mo +4VLM AdaptationRL for Language Model Reasoning

  19. HumanMoveVQA: Can Video MLLMs reason about human movement in videos?

    Jun 26, 2026Pulkit Gera, Faegheh Sardari, Asmar Nadeem +4VLM EvaluationSpatial Reasoning Benchmarks

  20. Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding

    Jun 26, 2026Shuimu Chen, Yuteng Chen, Yuanshen Guan +7LLM Self-CorrectionLong-Video Understanding

  21. Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

    Jun 26, 2026Hohin Kwan, Hongyu Li, Ray Zhang +5VLM EvaluationMultimodal Large Language Models

  22. Confidence-Aware Tool Orchestration for Robust Video Understanding

    Jun 25, 2026Yangfan He, Yujin Choi, Jaehong YoonVLM RobustnessTool Use in VLMs

  23. FeVOS: Foresight Expression Video Object Segmentation

    Jun 24, 2026Kehan Lan, Kaining Ying, Henghui DingVideo Object SegmentationReferring Video Object Segmentation

  24. SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

    Jun 23, 2026Sheng Xia, Zhengqin Lai, Tianxiang Jiang +4Temporal Video GroundingVideo Reasoning

  25. VideoLatent: Video-Language Learning via Latent Self-Forcing

    Jun 22, 2026Zi-Yuan Hu, Zicong Tang, Shijia Huang +3Multimodal Large Language ModelsLatent Visual Reasoning

  26. HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning

    Jun 19, 2026Awais Rauf, Ahmed Hasssan, Greg SlabaughVideo UnderstandingLong-Context Modeling

  27. CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs

    Jun 18, 2026Chengwen Liu, Hao Peng, Jisheng Dang +3Reward ShapingVideo Reasoning

  28. Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs

    Jun 16, 2026Chengwen Liu, Zhe Huang, Jisheng Dang +3Multimodal Reward ModelingVideo Reasoning