Video Reasoning

Momentum

13 papers in the last four weeks, up 86% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 135

All topics
CardsList
  1. LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue

    May 19, 2026Chaoyue Li, Yongxue Xu, Jie Feng +1Spatiotemporal ReasoningMultimodal Large Language Models

  2. PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning

    May 16, 2026Sikuan Yan, Sicheng Dong, Haotong Wang +8Multimodal MemoryLong-Video Understanding

  3. TRACE: Evidence Grounding-Guided Multi-Video Event Understanding and Claim Generation

    May 16, 2026Pengyu Yan, Akhil Gorugantu, Mahesh Bhosale +3Video ReasoningEvidence-Grounded Generation

  4. Video Models Can Reason with Verifiable Rewards

    May 14, 2026Tinghui Zhu, Sheng Zhang, James Y. Huang +5Video Diffusion ModelsRL for Video Generation

  5. Video-Zero: Self-Evolution Video Understanding

    May 14, 2026Ruixu Zhang, Deyi Ji, Lanyun Zhu +4Temporal Video GroundingLong-Video Understanding

  6. ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding

    May 13, 2026Xiao Liu, Nayu Liu, Junnan Zhu +6Video QATool-Using Vision-Language Agents

  7. WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

    May 11, 2026Keming Wu, Yijing Cui, Wenhan Xue +11World ModelsVideo Reasoning

  8. CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

    May 9, 2026Joowon Kim, Seungho Shin, Joonhyung Park +1Video GenerationVideo Reasoning

  9. SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

    May 8, 2026Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy +1Multimodal GroundingVideo Reasoning

  10. Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

    May 8, 2026Yuancheng Wei, Linli Yao, Lei Li +4Reward ModelingMultimodal Reward Modeling

  11. Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs

    May 8, 2026Peitao Han, Fei Cheng, Lis K. Pereira +2Vision-Language ModelsTemporal Reasoning in Language Models

  12. RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation

    May 8, 2026Junwei Wen, Deshui Miao, Guangming Lu +2Video Object SegmentationVideo Reasoning

  13. Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios

    May 7, 2026Peizheng Yan, Yu Zhao, Liang Xie +3Retrieval-Augmented GenerationStreaming Video Understanding

  14. VISD: Enhancing Video Reasoning via Structured Self-Distillation

    May 7, 2026Hao Lin, Kunyang Lv, Xu Jiang +5Video-Language ModelsVideo Reasoning

  15. 4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

    May 7, 2026Zhangquan Chen, Manyuan Zhang, Xinlei Yu +9Visual Spatial ReasoningVideo Reasoning

  16. VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

    May 5, 2026Andong Deng, Dawei Du, Zhenfang Chen +7Video Editing BenchmarksVideo Editing

  17. From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs

    May 4, 2026Le Zhang, Jihan Yang, Soundarya Krishnan +11Spatial Reasoning BenchmarksVisual Spatial Reasoning

  18. Act2See: Emergent Active Visual Perception for Video Reasoning

    May 3, 2026Martin Q. Ma, Yuxiao Qu, Aditya Agrawal +4Tool Use in VLMsVideo Reasoning

  19. SoccerRef-Agents: Multi-Agent System for Automated Soccer Refereeing

    Apr 25, 2026Zi Meng, Wanli Song, Yi Hu +2Multi-Agent LLM SystemsSports Video Analysis

  20. ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward

    Apr 23, 2026Jingpei Wu, Xiao Han, Weixiang Shen +3Process Reward ModelsMultimodal QA

  21. Neuro-Symbolic Manipulation Understanding with Enriched Semantic Event Chains

    Apr 22, 2026Fatemeh ZiaeetabarNext-Action PredictionVideo Action Recognition

  22. Video-ToC: Video Tree-of-Cue Reasoning

    Apr 22, 2026Qizhong Tan, Zhuotao Tian, Guangming Lu +2Video ReasoningVLM Reasoning

  23. SurgCoT: Advancing Spatiotemporal Reasoning in Surgical Videos through a Chain-of-Thought Benchmark

    Apr 22, 2026Gui Wang, YongSong Zhou, Kaijun Deng +4Spatiotemporal ReasoningHealthcare

  24. How Far Are Video Models from True Multimodal Reasoning?

    Apr 21, 2026Xiaotian Zhang, Jianhui Wei, Yuan Wang +9Multimodal ReasoningVideo Reasoning

  25. Find, Fix, Reason: Context Repair for Video Reasoning

    Apr 17, 2026Haojian Huang, Chuanyu Qin, Yinchuan Li +1Video Reasoning

  26. RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees

    Apr 17, 2026Yichen Xu, Yuanhang Liu, Chuhan Wang +5Sports Video AnalysisVideo Reasoning