Query-Aware Video Frame Selection

Latest papers 36

All topics
CardsList
  1. RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding

    Oct 6, 2026Yiyang Huang, Yitian Zhang, Yizhou Wang +6Long-Video UnderstandingQuery-Aware Video Frame Selection

  2. Less Is More: Genetic Frame Selection for Efficient Novel View Synthesis

    Sep 28, 2026Diego E. Farchione, Ramzi Idoughi, Alberto Jaspe-Villanueva +1Novel View SynthesisFeed-Forward 3D Reconstruction

  3. MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding

    Sep 14, 2026Hongchang Shi, Jinpeng Hu, Ao Wang +4Long-Video UnderstandingVideo Frame Selection

  4. One Skill Does Not Fit All: Automatic Discovery and Taxonomy-Guided Routing of Frame-Selection Skills for Long-Video Question Answering

    Sep 14, 2026Jian Hu, Zixu Cheng, Da Li +3LLM Agent Skill LearningLong-Video QA

  5. Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

    Sep 11, 2026Weitong Cai, Hang Zhang, Yukai Huang +6Memory-Augmented VLMsEfficient VLM Inference

  6. CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding

    Sep 3, 2026Jing Jiang, Yiran Ling, Ruonan Li +2Streaming Video UnderstandingEfficient VLM Inference

  7. RIDGE: Region-Informed Derivative-Guided Evidence Selection for Long Video Understanding

    Aug 30, 2026Shanqing Xu, Meng Luo, Mengchen Qian +7Large Vision-Language ModelsLong-Video Understanding

  8. Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

    Aug 6, 2026Bo Zhang, Wenxin Wang, Feng Chen +4Efficient VLM InferenceLong-Video Understanding

  9. One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

    Aug 6, 2026Wang Chen, Yu Chen, Xiang Wang +3Long-Video UnderstandingVideo Frame Selection

  10. When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

    Aug 4, 2026Ke Li, Jiayu Chen, Maoliang Li +5Efficient VLM InferenceLarge Vision-Language Models

  11. Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering

    Aug 3, 2026Fan Wei, Siru Zhong, Runmin Dong +3Long-Video UnderstandingVideo Frame Selection

  12. Coverage-Driven Adaptive Keyframe Selection for Video Understanding

    Aug 1, 2026Junyang Zhang, Puhan Luo, Chen Tang +2Efficient VLM InferenceLarge Vision-Language Models

  13. VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding

    Jul 30, 2026Haiyue Zhang, Yi Bin, Xun Jiang +5Large Vision-Language ModelsLong-Video Understanding

  14. FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

    Jul 28, 2026Ghazal Kaviani, Ghassan AlRegibEfficient Multimodal InferenceMultimodal Large Language Models

  15. LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos

    Jul 27, 2026Ce Zhang, Jinxi He, Katia Sycara +1Multimodal Large Language ModelsLong-Video Understanding

  16. QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding

    Jul 6, 2026Wei Ao, Lan Wang, Vishnu Naresh BoddetiTemporal Video GroundingStreaming Video Understanding

  17. ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

    Jul 2, 2026Minkuk Kim, Suyong Yun, Young Tae Kim +3Efficient VLM InferenceVideo QA

  18. QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding

    Jul 1, 2026Jun Peng, Baiyang Song, Jie Li +4Long-Video UnderstandingVideo Frame Selection

  19. DART: Difficulty-Adaptive Routing for Zero-Shot Video Temporal Grounding

    Jul 1, 2026Zhengbo Zhang, Mark He Huang, Zhigang Tu +1Temporal Video GroundingVLM Reasoning

  20. Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval

    May 22, 2026Michal Shlapentokh-Rothman, Prachi Garg, Yu-Xiong Wang +1Tool Use in VLMsTemporal IR

  21. CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering

    May 18, 2026Mahesh Bhosale, Abdul Wasi, Vishvesh Trivedi +3Claim VerificationVideo QA

  22. AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding

    May 13, 2026Xiao Yang, Yingzhe Ma, Haoxuan Yu +2Efficient VLM InferenceLong-Video Understanding

  23. LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs

    May 12, 2026Jingfeng Chen, Jiawen Qian, Wendi Deng +5Efficient VLM InferenceMultimodal Large Language Models

  24. GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs

    May 11, 2026Mohamed Eltahir, Lama Ayash, Ali Habibullah +2Efficient VLM InferenceLong-Video Understanding

  25. CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding

    May 9, 2026Mehrajul Abadin Miraj, Abdul Mohaimen Al Radi, Shariful Islam Rayhan +4Data SelectionLong-Video Understanding