Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-k embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We introduce \textbf{MetaSampling}, a training-free, plug-and-play sampling strategy that can be applied on top of existing frame selectors. MetaSampling improves downstream VQA efficiency by dynamically reducing the number of frames passed to the MLLM while preserving, and in some cases improving, answer accuracy. We evaluate MetaSampling across 36 paired frame-selector--MLLM-backbone--VQA-benchmark configurations. MetaSampling reduces the number of selected frames in all 36 configurations and improves accuracy in 25 of them, yielding an average frame reduction of 8.9% while slightly improving accuracy overall.
Figures & tables
Figure 4 : Overview of MetaSampling. Given any frozen frame selector, MetaSampling first obtains a small set of global anchors, uses their temporal distribution to allocate local sampling budgets, applies the same selector independently within each interval, and unions the global and local selections. This preserves broad temporal coverage while concentrating evidence in anchor-relevant regions and can reduce the number of unique frames passed to the answer model.
Figure 5 : Qualitative LongVideoBench examples for AKS, FOCUS, VideoITG, and AdaQ. Each panel compares frames selected by the base selector (top) and with MetaSampling (bottom). Timelines show the complete selected frame sets and light-green highlights indicate the frames shown in the representative filmstrips. Gold stars mark displayed frames selected only by MetaSampling.
Model
Method
VideoMME [ 6 ]
LongVideoBench [ 22 ]
LVBench [ 17 ]
Overall
Acc. ↑
Time ↓
# Frames ↓
Acc. ↑
Time ↓
# Frames ↓
Acc. ↑
Time ↓
# Frames ↓
Acc. ↑
Time ↓
# Frames ↓
Qwen3.5-9B [ 14 ]
AKS [ 15 ]
69.81
7.92
62.28
67.46
11.71
53.09
51.52
13.10
64.00
64.18
10.26
60.56
AKS + ours
69.96
6.82
54.91
68.51
10.59
45.91
51.97
10.47
52.12
64.62
8.73
51.98
FOCUS [ 30 ]
66.81
7.13
55.52
64.85
11.16
48.96
52.55
13.10
64.00
62.39
9.75
56.30
FOCUS + ours
67.37
6.73
53.15
65.07
10.75
46.21
53.07
12.42
61.08
62.85
9.27
53.69
VideoITG [ 16 ]
73.04
7.94
62.28
70.91
11.68
53.09
59.07
13.08
64.00
68.66
10.26
60.56
Table 1: Accuracy (%), inference time (s/q), and mean number of frames passed to the answer model for four selectors across three MLLM/VLM backbones. Overall columns are weighted by the number of evaluated questions in each dataset (2,700 VideoMME, 1,337 LongVideoBench, and 1,549 LVBench). The Δ Avg. row reports mean accuracy change, time speedup, and relative frame-count change. Green cells mark paired improvements of at least 1.0 percentage point in accuracy or 1.15× in speed; red cells mark paired accuracy losses greater than 1.0 percentage point; gray cells denote per-column aggregate statistics per backbone. † denotes a probabilistic selector.
Figure 6 : Accuracy–efficiency trade-offs across eight model–dataset settings. Ninth one is on Fig. 2 . The green shading indicates the favorable upper-left direction. Axis scales vary by panel.
Method
VMME
LVB
LVBench
AKS
69.67
67.61
50.23
AKS + ours
69.96
68.51
51.97
Δ
+0.29
+0.90
+1.74
FOCUS
66.81
65.82
52.36
FOCUS + ours
67.37
65.07
53.07
Δ
+0.56
-0.75
+0.71
Table 2 : Matched-budget accuracy (%) with Qwen3.5-9B. Budgets are matched by mean evidence count from cached selector outputs; AdaQ retains its 64-proposal runs. Gray Δ rows report paired accuracy changes, with green and red denoting gains and losses, respectively. † denotes a probabilistic sampler.
Method
Route Acc. (%)
Full Acc. (%)
Time (s/q) ↓
Frames ↓
Plain AKS
63.05
67.46
15.19
64.00
Ours, no backfill
64.72
68.51
13.40
52.55
Ours, backfill
65.55
69.04
15.20
64.00
Δ
+0.83
+0.52
+13.4%
+21.8%
Table 3: Backfill control on LongVideoBench using Qwen3.5-9B and AKS. Route metrics use the same 839 MetaSampling-route questions; full accuracy includes the 498 unchanged short-route questions. Green cells mark the better value for each metric; the Δ row reports backfill relative to no backfill.
Allocation
Route Acc. (%)
Full Acc. (%)
Time (s/q) ↓
Frames ↓
Equal
62.46
67.09
13.48
52.91
Anchor-weighted (ours)
64.72
68.51
13.40
52.55
Δ
+2.26
+1.42
1.01×
−0.68%
Table 4: Allocation ablation on LongVideoBench using Qwen3.5-9B and AKS. Route accuracy is measured on the 839 MetaSampling-route questions; full accuracy includes the 498 unchanged short-route questions. Both variants disable backfill. Green cells mark the better paired value; the Δ row reports anchor-weighted minus equal allocation.
b0
Route Acc. (%)
Full Acc. (%)
Time (s/q) ↓
Frames ↓
4
64.84
68.59
14.69
60.09
12
64.72
68.51
13.40
52.55
20
61.86
66.72
12.53
46.98
Table 5 : Full-LongVideoBench sensitivity to the global-anchor budget b0 using Qwen3.5-9B and AKS. Route metrics use all 839 union-route questions; full accuracy includes the 498 unchanged short-route questions. All other settings are fixed at B=64 , K=6 , m=4 , and α=1 . Confidence intervals use 10,000 video-level bootstrap resamples over 326 videos. Paired differences and W/L/T are reported as the default b0=12 minus each alternative.
K
Route Acc. (%)
Full Acc. (%)
Time (s/q) ↓
Frames ↓
4
62.93
67.39
13.33
52.41
6
64.72
68.51
13.40
52.55
8
63.53
67.76
13.45
52.84
Table 6 : Full-LongVideoBench sensitivity to the number of temporal intervals K using Qwen3.5-9B and AKS. Route metrics use all 839 union-route questions; full accuracy includes the 498 unchanged short-route questions. All settings use B=64 , b0=12 , m=4 , and α=1 . Confidence intervals use 10,000 video-level bootstrap resamples over 326 videos. Paired differences and W/L/T are reported as the default K=6 minus each alternative.
Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bound by finite context windows. Within the controlled frame-budget regime that governs practical deployment, prior selectors score frames against a single global query embedding; as a result, compositional multimodal questions that involve temporal ordering or cross-modal cues such as ``what happens on screen right after the narrator mentions the reaction?'' are flattened into a representation that loses sub-event ordering and modality bindings. We introduce \textbf{HiMu}, a training-free framework for compositional multimodal frame selection. A single text-only LLM call decomposes the query into a hierarchical logic tree whose leaves are atomic predicates, each routed to a lightweight expert spanning vision (CLIP, open-vocabulary detection, OCR) and audio (speech recognition and non-speech sound matching). Expert signals are normalized, smoothed to align across modalities, and composed bottom-up through fuzzy-logic operators that enforce temporal sequencing and adjacency, yielding a continuous per-frame satisfaction curve. Under the standard 16-frame budget on Video-MME, LongVideoBench, and HERBench-Lite, HiMu achieves state-of-the-art accuracy among frame selection methods and improves over uniform sampling across seven diverse MLLMs as a drop-in module, matching the accuracy of uniform sampling at 4× its frame budget, without retraining and without multiple iterative MLLM calls during selection.
Dan Ben-Ami, Gabriele Serussi, Kobi Cohen +1
INSIGHT Lab, Ben-Gurion University of the Negev, Israel · Ben-Gurion University of the Negev, Israel
Recent multimodal large language models (MLLMs) have substantially advanced video understanding, yet long-form video QA remains challenging under fixed input token budgets, where uniform sampling can be inefficient for evidence localization. We propose ReQuest , an uncertainty-driven, question-adaptive keyframe selection pipeline that aligns question intent with relevant video content through selective computation. ReQuest integrates (i) a lightweight question-aware selector distilled from MLLM-generated supervision, (ii) Re-thinking Routing that triggers additional inference only when the model is uncertain with a length-adaptive criterion, and (iii) uncertainty-guided adaptive non-maximum suppression that selects temporally diverse frames while adjusting spacing based on question difficulty. As a plug-andplay method, ReQuest improves long-video QA without modifying or fine-tuning the underlying MLLM. Experiments on Video-MME, MLVU, and LongVideoBench demonstrate consistent accuracy gains with competitive computational cost, with particularly strong improvements in medium and long video regimes.
Minkuk Kim, Suyong Yun, Young Tae Kim +3
Kyung Hee University, Republic of Korea · Electronics and Telecommunications Research Institute (ETRI), Republic of Korea
Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current Video QA benchmarks genuinely require temporal frame selection, or can most questions be answered regardless of which frames are shown? We introduce Frame Selection Sensitivity (FSS), a per-sample diagnostic that measures how much VLM accuracy changes when the most relevant frames are replaced with the least relevant ones. Across six benchmarks and eight VLMs, we find that a large majority of samples are frame-agnostic: only a minority are genuinely sensitive to frame choice. Combining FSS with a Language Independence Score (LIS) reveals that merely 5.5--31% of samples are Temporally Sensitive. We construct TempCore, compact evaluation subsets that isolate these temporal samples from existing benchmarks, and will release code and per-sample annotations upon publication.