Long-Video QA

QA: Question Answering

Momentum

12 papers in the last four weeks, up 200% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 61

All topics
CardsList
  1. CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams

    Jun 22, 2026Xu Liu, Guikun Chen, Zihao Yan +2Streaming Video UnderstandingLong-Video Understanding

  2. TimeProVe: Propose, then Verify for Efficient Long Video Temporal Reasoning in Activities of Daily Living

    Jun 18, 2026Arkaprava Sinha, Dominick Reilly, Siddharth Krishnan +2Temporal Video GroundingEfficient VLM Inference

  3. VL-MemKnG: Hybrid Memory with a Spatio-Temporal Knowledge Graph for Question Answering over Long Egocentric Navigation Trajectories

    Jun 15, 2026Svetlana Lukina, Mohamad Al Mdfaa, Gloria Haro +2Egocentric Video QAVideo QA

  4. Rethinking RAG in Long Videos: What to Retrieve and How to Use It?

    Jun 11, 2026Yuho Lee, Jisu Shin, Nicole Hee-Yeon Kim +5Multimodal RAGEgocentric Video Understanding

  5. From Content to Knowledge: Lightning Fast Long-Video Understanding with Neural Knowledge Representations

    Jun 10, 2026Yuchen Guan, Xiao Li, Zongyu Guo +4Efficient VLM InferenceLong-Video Understanding

  6. See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding

    Jun 8, 2026Shuning Wang, Zhiheng Wu, YiNuo Lu +6Video QAVideo-Language Models

  7. StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

    Jun 4, 2026Zhengqian Wu, Zhixian Liu, Aodong Chen +6Video UnderstandingSynthetic Data Generation

  8. MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering

    Jun 4, 2026Qing Yang, Pengcheng Huang, Xinze Li +6Video SummarizationVideo QA

  9. Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA

    May 31, 2026Yuyang Sun, Yongliang Wu, Xingyu Zhu +8Video QAStructured Reasoning

  10. SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

    May 30, 2026Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx +4Multimodal MemoryEgocentric Video QA

  11. CASTLE2026 Team WDL Technical Report

    May 30, 2026Zhengyang Li, Zhenglin Du, Yi Wen +3Multimodal RAGEgocentric Video QA

  12. Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval

    May 22, 2026Michal Shlapentokh-Rothman, Prachi Garg, Yu-Xiong Wang +1Tool Use in VLMsTemporal IR

  13. Swift Sampling: Selecting Temporal Surprises via Taylor Series

    May 21, 2026Dahye Kim, Bhuvan Sachdeva, Karan Uppal +3Video Frame SelectionLong-Video QA

  14. MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering

    May 21, 2026Junbin Xiao, Jiajun Chen, Tianxiang Sun +2KV CachingVideo QA

  15. SurgLQA: Scalable Long-Horizon Surgical Video Question Answering

    May 18, 2026Diandian Guo, Xikai Yang, Ruiyang Li +2Temporal Video GroundingTest-Time Scaling

  16. AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding

    May 13, 2026Xiao Yang, Yingzhe Ma, Haoxuan Yu +2Efficient VLM InferenceLong-Video Understanding

  17. VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority

    May 12, 2026Chenhao Qiu, Yechao Zhang, Xin Luo +2Large Vision-Language ModelsVideo QA

  18. Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

    May 8, 2026Jiazheng Li, Chi-Hao Wu, Yunze Liu +3Video UnderstandingMultimodal Memory

  19. Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios

    May 7, 2026Peizheng Yan, Yu Zhao, Liang Xie +3Retrieval-Augmented GenerationStreaming Video Understanding

  20. SF20K Competition 2025: Summary and findings

    May 2, 2026Ridouane Ghermi, Xi Wang, Vicky Kalogeiton +1VLM EvaluationVideo Understanding

  21. VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning

    Mar 26, 2026Zhe Gao, Shiyu Shen, Taifeng Chai +7Tool Use in VLMsEfficient VLM Inference

  22. HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering

    Mar 19, 2026Dan Ben-Ami, Gabriele Serussi, Kobi Cohen +1Multimodal Large Language ModelsMultimodal QA

  23. FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering

    Mar 4, 2026Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov +4Efficient VLM InferenceEgocentric Video QA

  24. Agentic Very Long Video Understanding

    Jan 26, 2026Aniket Rege, Arka Sadhu, Yuliang Li +5Question AnsweringEgocentric Video Understanding

  25. Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding

    Date pendingTianyue Wang, Xuying Wu, Yuxiang Ma +7Adaptive RetrievalLong-Video Understanding