Long-Video QA

QA: Question Answering

Momentum

12 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 61

All topics
CardsList
  1. CASE: Cost-Aware Stopping for Efficient Long-Video Agents

    Oct 4, 2026Yiming Du, Chenghao Liu, Zhiyuan Liu +4Efficient VLM InferenceVideo Reasoning

  2. LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

    Sep 30, 2026Juyi Lin, Zhiqiang Lao, Jiali Cui +9Audio-Visual QALong-Video QA

  3. Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

    Sep 29, 2026Hui Ren, Lei Fan, Henry Pao +5Multimodal MemoryQuestion Answering

  4. Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

    Sep 29, 2026Ziheng Huang, Yicheng Bao, Xueheng Li +8Streaming Video UnderstandingVideo QA

  5. MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering

    Sep 27, 2026Ashim Dahal, Bikramjit BanerjeeMultimodal Large Language ModelsAdaptive Sampling

  6. Long-to-Short Video Evidence Reasoning for Grounded Question Answering

    Sep 14, 2026Kaiyan Chen, Junbin Xiao, Xun YangTemporal Video GroundingVideo QA

  7. One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering

    Sep 14, 2026Jian Hu, Zixu Cheng, Da Li +3LLM Agent Skill LearningLong-Video QA

  8. Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

    Sep 11, 2026Weitong Cai, Hang Zhang, Yukai Huang +6Memory-Augmented VLMsEfficient VLM Inference

  9. EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

    Sep 1, 2026Yijun Chen, Yaqi Zheng, Yanya Li +11Multimodal RAGEfficient Multimodal Inference

  10. From Intent to Evidence: Policy-Steered Multi-Strategy Retrieval for Long-Video Agents

    Aug 31, 2026Can Zhang, Baofeng Zhang, Xiaotian Han +5Long-Video UnderstandingLong-Video QA

  11. NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

    Aug 13, 2026Yuheng Huang, Jianlang Chen, Jiayang Song +6Video UnderstandingVideo-Language Model Evaluation

  12. R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video

    Aug 11, 2026Ke Ma, Yamin Mao, Weiming Li +5Egocentric Video QAVideo QA

  13. REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering

    Aug 9, 2026Caijun Yan, Yang Zhou, Meixing Shi +4Video QALong-Video QA

  14. Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding

    Aug 7, 2026Yeeun Choi, Youngbeom Yoo, Joon-Young Lee +2Memory-Augmented VLMsEpisodic Memory

  15. Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering

    Aug 3, 2026Fan Wei, Siru Zhong, Runmin Dong +3Long-Video UnderstandingVideo Frame Selection

  16. Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding

    Jul 30, 2026Bowen Liu, Shuning Wang, Xinpeng Ding +3Video-Language ModelsLong-Video Understanding

  17. CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

    Jul 27, 2026Jinlong Yang, Wenhao Zhang, Kuanwei Lin +1Tool Use in VLMsEfficient VLM Inference

  18. ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning

    Jul 21, 2026Santiram Tiwari, Nishant Sinha, Kunal KislayMemory-Augmented VLMsKV Caching

  19. Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

    Jul 17, 2026Ce Zhang, Ziyang Wang, Yulu Pan +6Temporal Video GroundingTool-Using Vision-Language Agents

  20. SLVMBench: Skill Learning from Video Memory

    Jul 13, 2026Yudong Yang, Guangzhi Sun, Yixuan Li +1Memory-Augmented VLMsVideo-Language Models

  21. Incentivizing Vision Language Models to Search for Long Video Question Answering

    Jul 3, 2026Harsh Goel, S P Sharan, Sahil Shah +4Tool Use in VLMsLong-Video QA

  22. ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

    Jul 2, 2026Minkuk Kim, Suyong Yun, Young Tae Kim +3Efficient VLM InferenceVideo QA

  23. QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding

    Jul 1, 2026Jun Peng, Baiyang Song, Jie Li +4Long-Video UnderstandingVideo Frame Selection

  24. Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning

    Jul 1, 2026Yixin Ji, Fanghua Ye, Juntao Li +5Long-Video UnderstandingMemory-Augmented Video Understanding

  25. Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs

    Jun 24, 2026Agnese Taluzzi, Riccardo Santambrogio, Simone Mentasti +2Egocentric Video QAVideo QA