Video QA

QA: Question Answering

Momentum

10 papers in the last four weeks, up 11% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 116

All topics
CardsList
  1. ViMU: Benchmarking Video Metaphorical Understanding

    May 14, 2026Qi Li, Xinchao WangVideo UnderstandingVideo QA

  2. ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding

    May 13, 2026Xiao Liu, Nayu Liu, Junnan Zhu +6Video QATool-Using Vision-Language Agents

  3. VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority

    May 12, 2026Chenhao Qiu, Yechao Zhang, Xin Luo +2Large Vision-Language ModelsVideo QA

  4. ChronoSC: Task-Oriented Semantic Communication via Temporal-to-Color Encoding

    May 11, 2026Phuc H. Nguyen, Trung T. Nguyen, Quy N. Duong +1Semantic CommunicationVideo QA

  5. Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

    May 8, 2026Hang Wu, Sherin Mary Mathews, Yujun Cai +2Streaming Video UnderstandingEfficient VLM Inference

  6. VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA

    May 6, 2026Haibin He, Maoyuan Ye, Jing Zhang +2Video QAVideo Frame Selection

  7. KARMA-MV: A Benchmark for Causal Question Answering on Music Videos

    May 5, 2026Archishman Ghosh, Abhinaba Roy, Dorien HerremansAudio-Visual UnderstandingVideo QA

  8. VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition

    May 4, 2026Tanush Yadav, Mohammadreza Salehi, Jae Sung Park +6VLM EvaluationVLM Adaptation

  9. SF20K Competition 2025: Summary and findings

    May 2, 2026Ridouane Ghermi, Xi Wang, Vicky Kalogeiton +1VLM EvaluationVideo Understanding

  10. CurEvo: Curriculum-Guided Self-Evolution for Video Understanding

    Apr 29, 2026Guiyi Zeng, Junqing Yu, Yi-Ping Phoebe Chen +3Video QACurriculum Learning

  11. Interactive Episodic Memory with User Feedback

    Apr 27, 2026Nikesh Subedi, Loris Bazzani, Ziad Al-HalahEgocentric Video QAEpisodic Memory

  12. Don't Pause! Every prediction matters in a streaming video

    Apr 27, 2026Dibyadip Chatterjee, Zhanzhong Pang, Fadime Sener +2Streaming Video UnderstandingEfficient VLM Inference

  13. UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks

    Apr 25, 2026Jason Nguyen, Ameet Rao, Alexander Chang +2Multimodal Large Language ModelsVideo QA

  14. Grounding Video Reasoning in Physical Signals

    Apr 23, 2026Alibay Osmanli, Zixu Cheng, Shaogang GongTemporal Video GroundingVideo QA

  15. CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs

    Apr 22, 2026Xingcheng Zhou, Hao Guo, Rui Song +5Visual Question AnsweringVideo QA

  16. Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

    Apr 20, 2026Kecheng Zhang, Zongxin Yang, Mingfei Han +6Streaming Video UnderstandingEfficient VLM Inference

  17. When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models

    Apr 19, 2026Cui Yakun, Xingqun Qi, TianTian Geng +3VLM EvaluationVLM Hallucination

  18. ForeSea: AI Forensic Search with Multi-modal Queries for Video Surveillance

    Mar 24, 2026Hyojin Park, Yi Li, Janghoon Cho +8Multimodal RAGMultimodal IR

  19. Overreliance on AI in Information-seeking from Video Content

    Mar 20, 2026Anders Giovanni Møller, Elisa Bassignana, Francesco Pierri +1Video QATrust in AI

  20. DynaTokens: Controlling Token Dynamics for Continual Video-Language Understanding

    Mar 2, 2026Toan Nguyen, Yang Liu, Celso De Melo +1Continual LearningVideo QA

  21. NeMo: Needle in a Montage for Video-Language Understanding

    Sep 29, 2025Zi-Yuan Hu, Shuo Liang, Duo Zheng +10Temporal Video GroundingVideo-Language Model Evaluation

  22. ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering

    Aug 28, 2025Paritosh Parmar, Eric Peh, Basura FernandoVideo QAMultimodal CoT Reasoning

  23. CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes

    Apr 1, 2024Paritosh Parmar, Eric Peh, Ruirui Chen +4Video QACausal Reasoning

  24. StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs

    Date pendingJoya Chen, Zeyun Zhong, Mike Zheng ShouMemory-Augmented VLMsStreaming Video Understanding