Video Reasoning

Momentum

13 papers in the last four weeks, up 86% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 135

All topics
CardsList
  1. Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding

    Apr 16, 2026Zhixuan Wu, Quanxing Zha, Teng Wang +6Video ReasoningMultimodal CoT Reasoning

  2. GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models

    Apr 12, 2026Nicolae Cudlenco, Mihai Masala, Marius LeordeanuSynthetic Data GenerationVideo Representation Learning

  3. From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation

    Mar 16, 2026Yibin Liu, Yaxing Lyu, Daqi Gao +5RoboticsRobotic Manipulation

  4. A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding

    Mar 16, 2026Yue Zhang, Liqiang Jing, Jia Li +4Video UnderstandingMultimodal QA

  5. Are Video Reasoning Models Ready to Go Outside?

    Mar 11, 2026Yangfan He, Changgyu Boo, Jaehong YoonVLM RobustnessVideo Reasoning

  6. What if? Emulative Simulation with World Models for Situated Reasoning

    Mar 6, 2026Ruiping Liu, Yufan Chen, Yuheng Zhang +8Visual Spatial ReasoningVideo Reasoning

  7. A Very Big Video Reasoning Suite

    Feb 23, 2026Maijunxian Wang, Ruisi Wang, Juyi Lin +53Reasoning EvaluationVideo Reasoning

  8. TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs

    Jan 30, 2026Baiqi Li, Kangyi Zhao, Ce Zhang +3VLM EvaluationVideo Understanding

  9. CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation

    Jan 12, 2026Chaoyu Li, Fei Tao, Pooyan FazliTest-Time Scaling for VLMsTest-Time Scaling

  10. VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

    Nov 25, 2025Tianxiang Jiang, Sheng Xia, Yicheng Xu +5VLM EvaluationMultimodal Large Language Models

  11. TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

    May 2, 2025Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou +11Temporal Action SegmentationTemporal Video Grounding

  12. Learning to Visually Connect Actions and their Effects

    Jan 19, 2024Paritosh Parmar, Eric Peh, Basura FernandoVideo ReasoningSelf-Supervised Visual Representation Learning