Video Frame Selection

Momentum

5 papers in the last four weeks, up 67% on the four weeks before. 0.0% of all new papers.

Jul 13Week of Sep 28

Latest papers 46

All topics
CardsList
  1. MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering

    Sep 27, 2026Ashim Dahal, Bikramjit BanerjeeMultimodal Large Language ModelsAdaptive Sampling

  2. Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed

    Sep 24, 2026Ignat Romanov, Andreas Hadjipieris, Neofytos DimitriouTraining Data SelectionVideo Frame Selection

  3. MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding

    Sep 14, 2026Hongchang Shi, Jinpeng Hu, Ao Wang +4Long-Video UnderstandingVideo Frame Selection

  4. Detecting and Explaining Fake News Short Videos with Multimodal Content and Real-World Evidence

    Sep 14, 2026Yifeng Luo, Yupeng Li, Ming Tang +2Fake News DetectionAutomated Fact-Checking

  5. CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding

    Sep 3, 2026Jing Jiang, Yiran Ling, Ruonan Li +2Streaming Video UnderstandingEfficient VLM Inference

  6. RIDGE: Region-Informed Derivative-Guided Evidence Selection for Long Video Understanding

    Aug 30, 2026Shanqing Xu, Meng Luo, Mengchen Qian +7Large Vision-Language ModelsLong-Video Understanding

  7. Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

    Aug 6, 2026Bo Zhang, Wenxin Wang, Feng Chen +4Efficient VLM InferenceLong-Video Understanding

  8. One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

    Aug 6, 2026Wang Chen, Yu Chen, Xiang Wang +3Long-Video UnderstandingVideo Frame Selection

  9. When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

    Aug 4, 2026Ke Li, Jiayu Chen, Maoliang Li +5Efficient VLM InferenceLarge Vision-Language Models

  10. Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models

    Aug 4, 2026Yuxin Cao, Wei Song, Jingling Xue +1Selective PredictionVideo-Language Model Evaluation

  11. Ground, Cover, and Refine: Evidence-Centric Frame Selection for Long-Video Question Answering

    Aug 3, 2026Fan Wei, Siru Zhong, Runmin Dong +3Long-Video UnderstandingVideo Frame Selection

  12. Coverage-Driven Adaptive Keyframe Selection for Video Understanding

    Aug 1, 2026Junyang Zhang, Puhan Luo, Chen Tang +2Efficient VLM InferenceLarge Vision-Language Models

  13. FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

    Jul 28, 2026Ghazal Kaviani, Ghassan AlRegibEfficient Multimodal InferenceMultimodal Large Language Models

  14. PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

    Jul 22, 2026Zihan Song, Shuo Ye, Bo Zhao +4Video-Language ModelsVideo Frame Selection

  15. Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding

    Jul 14, 2026Yifan Lu, Ziqi Zhang, Chunfeng Yuan +3Large Vision-Language ModelsVideo Frame Selection

  16. ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

    Jul 2, 2026Minkuk Kim, Suyong Yun, Young Tae Kim +3Efficient VLM InferenceVideo QA

  17. QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding

    Jul 1, 2026Jun Peng, Baiyang Song, Jie Li +4Long-Video UnderstandingVideo Frame Selection

  18. Foundation Model-driven Key Anatomy Frame Selection for Blind-sweep Ultrasound Fetal Birth Weight Estimation

    Jul 1, 2026Le Ou, Xiliang Zhu, Huanwen Liang +8Ultrasound ImagingVideo Frame Selection

  19. Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling

    Jun 23, 2026Kun Zhang, Chenxin Fang, Tao Chen +4Efficient Multimodal InferenceMultimodal Large Language Models

  20. DBT-Bleed: Dual-Branch Temporal Modeling with Key-Frame Selection for Surgical Bleeding Detection

    Jun 22, 2026Sudhanshu Mishra, Jialang Xu, Jensen Ang +3Surgical Video UnderstandingVideo Frame Selection

  21. FetSelect: Task-Specific Architectures and Self-Supervised Learning for Automated Fetal Ultrasound Frame Selection

    Jun 21, 2026Mahmood Alzubaidi, Raden Muaz, Uzair Shah +4Ultrasound ImagingVideo Frame Selection

  22. EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

    Jun 18, 2026Ganlin Yang, Zhangzheng Tu, Yuqiang Yang +10Long-Horizon Robotic ManipulationVision-Language-Action Models

  23. Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding

    Jun 10, 2026Biao Tang, Xu Chen, Shuxiang Gou +3Multimodal Large Language ModelsLong-Video Understanding

  24. Dot-Flik: A Scalable Edge AI Architecture for Distributed Insect Monitoring

    May 27, 2026Mattia Consani, Denisa-Andreea Constantinescu, Åse Håtveit +4Edge ComputingInternet of Things

  25. Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval

    May 22, 2026Michal Shlapentokh-Rothman, Prachi Garg, Yu-Xiong Wang +1Tool Use in VLMsTemporal IR