cs.CVSep 27, 2026

MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering

Authors: Ashim Dahal, Bikramjit Banerjee

Organizations: University of Southern Mississippi Hattiesburg, MS, USA

Abstract

Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-kk embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We introduce \textbf{MetaSampling}, a training-free, plug-and-play sampling strategy that can be applied on top of existing frame selectors. MetaSampling improves downstream VQA efficiency by dynamically reducing the number of frames passed to the MLLM while preserving, and in some cases improving, answer accuracy. We evaluate MetaSampling across 36 paired frame-selector--MLLM-backbone--VQA-benchmark configurations. MetaSampling reduces the number of selected frames in all 36 configurations and improves accuracy in 25 of them, yielding an average frame reduction of 8.9%8.9\% while slightly improving accuracy overall.

Figures & tables

Explore similar work

CardsList
  1. HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering

    Mar 19, 2026Dan Ben-Ami, Gabriele Serussi, Kobi Cohen +1Long Video Question AnsweringVideo Multimodal Large Language Models

  2. ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

    Jul 2, 2026Minkuk Kim, Suyong Yun, Young Tae Kim +3Long Video Question AnsweringKeyframe Selection

  3. TempCore: Are Video QA Benchmarks Temporally Grounded?

    Sep 1, 2025Hyunjong Ok, Jaeho Lee