Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-k embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We introduce \textbf{MetaSampling}, a training-free, plug-and-play sampling strategy that can be applied on top of existing frame selectors. MetaSampling improves downstream VQA efficiency by dynamically reducing the number of frames passed to the MLLM while preserving, and in some cases improving, answer accuracy. We evaluate MetaSampling across 36 paired frame-selector--MLLM-backbone--VQA-benchmark configurations. MetaSampling reduces the number of selected frames in all 36 configurations and improves accuracy in 25 of them, yielding an average frame reduction of 8.9% while slightly improving accuracy overall.
Figures & tables
Figure 4 : Overview of MetaSampling. Given any frozen frame selector, MetaSampling first obtains a small set of global anchors, uses their temporal distribution to allocate local sampling budgets, applies the same selector independently within each interval, and unions the global and local selections. This preserves broad temporal coverage while concentrating evidence in anchor-relevant regions and can reduce the number of unique frames passed to the answer model.
Figure 5 : Qualitative LongVideoBench examples for AKS, FOCUS, VideoITG, and AdaQ. Each panel compares frames selected by the base selector (top) and with MetaSampling (bottom). Timelines show the complete selected frame sets and light-green highlights indicate the frames shown in the representative filmstrips. Gold stars mark displayed frames selected only by MetaSampling.
Model
Method
VideoMME [ 6 ]
LongVideoBench [ 22 ]
LVBench [ 17 ]
Overall
Acc. ↑
Time ↓
# Frames ↓
Acc. ↑
Time ↓
# Frames ↓
Acc. ↑
Time ↓
# Frames ↓
Acc. ↑
Time ↓
# Frames ↓
Qwen3.5-9B [ 14 ]
AKS [ 15 ]
69.81
7.92
62.28
67.46
11.71
53.09
51.52
13.10
64.00
64.18
10.26
60.56
AKS + ours
69.96
6.82
54.91
68.51
10.59
45.91
51.97
10.47
52.12
64.62
8.73
51.98
FOCUS [ 30 ]
66.81
7.13
55.52
64.85
11.16
48.96
52.55
13.10
64.00
62.39
9.75
56.30
FOCUS + ours
67.37
6.73
53.15
65.07
10.75
46.21
53.07
12.42
61.08
62.85
9.27
53.69
VideoITG [ 16 ]
73.04
7.94
62.28
70.91
11.68
53.09
59.07
13.08
64.00
68.66
10.26
60.56
Table 1: Accuracy (%), inference time (s/q), and mean number of frames passed to the answer model for four selectors across three MLLM/VLM backbones. Overall columns are weighted by the number of evaluated questions in each dataset (2,700 VideoMME, 1,337 LongVideoBench, and 1,549 LVBench). The Δ Avg. row reports mean accuracy change, time speedup, and relative frame-count change. Green cells mark paired improvements of at least 1.0 percentage point in accuracy or 1.15× in speed; red cells mark paired accuracy losses greater than 1.0 percentage point; gray cells denote per-column aggregate statistics per backbone. † denotes a probabilistic selector.
Figure 6 : Accuracy–efficiency trade-offs across eight model–dataset settings. Ninth one is on Fig. 2 . The green shading indicates the favorable upper-left direction. Axis scales vary by panel.
Method
VMME
LVB
LVBench
AKS
69.67
67.61
50.23
AKS + ours
69.96
68.51
51.97
Δ
+0.29
+0.90
+1.74
FOCUS
66.81
65.82
52.36
FOCUS + ours
67.37
65.07
53.07
Δ
+0.56
-0.75
+0.71
Table 2 : Matched-budget accuracy (%) with Qwen3.5-9B. Budgets are matched by mean evidence count from cached selector outputs; AdaQ retains its 64-proposal runs. Gray Δ rows report paired accuracy changes, with green and red denoting gains and losses, respectively. † denotes a probabilistic sampler.
Method
Route Acc. (%)
Full Acc. (%)
Time (s/q) ↓
Frames ↓
Plain AKS
63.05
67.46
15.19
64.00
Ours, no backfill
64.72
68.51
13.40
52.55
Ours, backfill
65.55
69.04
15.20
64.00
Δ
+0.83
+0.52
+13.4%
+21.8%
Table 3: Backfill control on LongVideoBench using Qwen3.5-9B and AKS. Route metrics use the same 839 MetaSampling-route questions; full accuracy includes the 498 unchanged short-route questions. Green cells mark the better value for each metric; the Δ row reports backfill relative to no backfill.
Allocation
Route Acc. (%)
Full Acc. (%)
Time (s/q) ↓
Frames ↓
Equal
62.46
67.09
13.48
52.91
Anchor-weighted (ours)
64.72
68.51
13.40
52.55
Δ
+2.26
+1.42
1.01×
−0.68%
Table 4: Allocation ablation on LongVideoBench using Qwen3.5-9B and AKS. Route accuracy is measured on the 839 MetaSampling-route questions; full accuracy includes the 498 unchanged short-route questions. Both variants disable backfill. Green cells mark the better paired value; the Δ row reports anchor-weighted minus equal allocation.
b0
Route Acc. (%)
Full Acc. (%)
Time (s/q) ↓
Frames ↓
4
64.84
68.59
14.69
60.09
12
64.72
68.51
13.40
52.55
20
61.86
66.72
12.53
46.98
Table 5 : Full-LongVideoBench sensitivity to the global-anchor budget b0 using Qwen3.5-9B and AKS. Route metrics use all 839 union-route questions; full accuracy includes the 498 unchanged short-route questions. All other settings are fixed at B=64 , K=6 , m=4 , and α=1 . Confidence intervals use 10,000 video-level bootstrap resamples over 326 videos. Paired differences and W/L/T are reported as the default b0=12 minus each alternative.
K
Route Acc. (%)
Full Acc. (%)
Time (s/q) ↓
Frames ↓
4
62.93
67.39
13.33
52.41
6
64.72
68.51
13.40
52.55
8
63.53
67.76
13.45
52.84
Table 6 : Full-LongVideoBench sensitivity to the number of temporal intervals K using Qwen3.5-9B and AKS. Route metrics use all 839 union-route questions; full accuracy includes the 498 unchanged short-route questions. All settings use B=64 , b0=12 , m=4 , and α=1 . Confidence intervals use 10,000 video-level bootstrap resamples over 326 videos. Paired differences and W/L/T are reported as the default K=6 minus each alternative.