Organizations: Department of Electrical and Computer Engineering, Northeastern University · Khoury College of Computer Science, Northeastern University
Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation--Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.
Figures & tables
Figure 1: Two key challenges in long-video frame selection. Answering “Why did the milk spill?” requires evidence of the cat tipping the glass and the resulting spill. Query Comprehension Gap (left): similarity-based methods overlook implicit causal requirements, retrieving spilled-milk frames while missing the cat tipping the glass. Interpretation–Selection Gap (right): judgment-based methods understand implicit evidence requirements but fail to perform selection reliably, choosing a milk-pouring frame instead of the causal action.
Figure 2: Empirical analysis of two gaps on LongVideoBench subsets. (a) Embedding models such as CLIP handle single-evidence queries effectively but struggle with implicit requirements in complex queries, which query decomposition helps make explicit. (b) VLMs are more effective at interpreting complex queries and describing relevant evidence than at directly selecting frames, as using these interpretations to guide tool-based selection yields higher accuracy and lower latency than direct VLM selection, especially for long videos.
Figure 3: Overview of RACER. RACER begins with (a) Query Interpretation , where a lightweight Vid-LLM uses a uniformly sampled overview to decompose the query into complementary sub-queries, making implicit evidence requirements explicit. Then, (b) Evidence Localization employs an embedding-based retrieval tool to score candidate frames against the original query and sub-queries, followed by entropy-adaptive score fusion and DPP selection to obtain relevant and visually diverse key frames. Finally, (c) Reflection merges feedback frames with the overview, removing duplicates and the least visually aligned overview frames to maintain the frame budget. The resulting hybrid overview guides sub-query refinement for the next selection round.
Table 4
Vid-LLMs
Frame Selection Method
Video-MME (1–60 min)
MLVU (3–120 min)
LongVideoBench (1–60 min)
Average
LLaVA-OV-7B
Uniform
58.5 (+0.0)
62.4 (+0.0)
56.6 (+0.0)
59.2 (+0.0)
AKS ( Tang et al., 2025 )
58.4 (-0.1)
66.8 (+4.4)
59.3 (+2.7)
61.5 (+2.3)
BOLT ( Liu et al., 2025 )
59.9 (+1.4)
66.8 (+4.4)
59.6 (+3.0)
62.1 (+2.9)
MDP 3 ( Sun et al., 2025 )
60.5 (+2.0)
68.3 (+5.9)
59.0 (+2.4)
62.6 (+3.4)
AdaQ ( Zhang et al., 2026 )
59.5 (+1.0)
69.4 (+7.0)
60.0 (+3.4)
63.0 (+3.8)
GIFT ( Ma et al., 2026 )
61.2 (+2.7)
69.4 (+7.0)
59.6 (+3.0)
63.4 (+4.2)
Table 3: Comparison with state-of-the-art frame-selection methods across three benchmarks and four Vid-LLMs: accuracy (%) for Video-MME and LongVideoBench, and M-Avg for MLVU. Parentheses show gains over uniform sampling (positive in green). Average is the unweighted mean; bold / underline mark best/second-best results per Vid-LLM. RACER achieves the best or tied-best performance across all benchmarks and answering models.
Answering Vid-LLM
Frame Selection
Retrieval Tool
Query Interpretation Model
Video-MME
Short
Medium
Long
Overall
Qwen2.5-VL-3B
Uniform Sampling
N/A
N/A
69.5
57.0
47.1
57.8
LLaVA-OV-7B
Uniform Sampling
N/A
N/A
70.2
57.1
48.2
58.5
+ Tool-based Retrieval
CLIP ViT-L/14
N/A
71.2
58.7
51.8
60.6
+ Query Interpretation
CLIP ViT-L/14
Qwen2.5-VL-3B
72.3
60.1
52.1
61.5
+ Reflection (RACER)
CLIP ViT-L/14
Qwen2.5-VL-3B
72.6
60.6
52.7
61.9
Table 4: Module ablation of RACER on Video-MME with components added cumulatively. Scores report accuracy (%) by duration and overall; bold marks column bests. Each component improves performance, with full RACER performing best. Despite weaker direct answering capability, Qwen2.5-VL-3B improves LLaVA-OV-7B when assigned Query Interpretation.
Table 7
Figure 4: Qualitative visualization of RACER. (a) Query interpretation guides tool-based evidence localization to recover the explanation that water expands upon freezing, missed by direct CLIP retrieval and VLM judgment. Full RACER retains this evidence with less redundancy. (b) Reflection refines sub-queries to target the blue object, recovering a previously missed close-up and correcting the answer from a clipper to a dremel .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Perception
Recognition
Reasoning
Information Synopsis
OCR Problems
Counting Problem
Uniform
68.9
60.8
53.2
74.3
66.2
36.6
CLIP Retrieve
69.2
59.5
53.9
66.9
66.2
37.3
VLM Judgment
70.7
62.1
55.8
75.2
66.2
36.2
RACER (ours)
71.9
63.4
57.8
75.8
67.6
41.0
Appendix
Table 7: Accuracy (%) across six Video-MME question categories with LLaVA-OV-7B.
Method
Short (1–3 min)
Medium (3–30 min)
Long (30–60 min)
Overall
Acc.
Time (s)
#Frames
Acc.
Time (s)
#Frames
Acc.
Time (s)
#Frames
Acc.
Time (s)
#Frames
Uniform
70.2
1.1
0
57.1
1.1
0
48.2
1.1
0
58.5
1.1
0
CLIP Retrieval
68.8
3.5
0
58.6
7.5
0
45.8
32.8
0
57.7
14.6
0
VLM judgment
70.7
22.8
157.7
58.0
58.4
516.3
51.7
286.2
2465.9
60.1
122.4
1046.6
RACER
72.6
22.4
57.6
60.6
29.0
64.0
52.7
54.9
67.1
61.9
35.4
62.9
Appendix
Table 8: Efficiency comparison on Video-MME across video durations. Time is measured in seconds, and #Frames denotes the average number of unique frames analyzed by the VLM.
DPP Relevance
Entropy Sensitivity
Original-Query Weight
Generation Temperature
α
Acc.
β
Acc.
w0
Acc.
τ
Acc.
0.55
60.41
2.0
62.11
0.4
61.37
0.6
60.30
0.65
60.63
2.1
61.96
0.5
61.22
0.7
61.00
0.75†
61.93
2.2†
61.93
0.6†
61.93
0.8†
61.93
0.85
60.89
2.3
61.89
0.7
61.22
0.9
60.59
0.95
61.52
2.4
61.93
0.8
60.48
1.0
60.89
Appendix
Table 9: Hyperparameter ablation on Video-MME with LLaVA-OV-7B. Each column pair reports a separate parameter sweep. Acc. denotes accuracy (%); † marks the default setting; bold indicates the best observed result.
Design
Configuration
Acc. (%)
Sub-query aggregation
Uniform weighting
61.67
Entropy-based weighting †
61.93
Reflection steps
0
61.52
1
61.89
2†
61.93
3
61.96
Appendix
Table 10: Ablations of sub-query aggregation, reflection depth, and reflection input on Video-MME with LLaVA-OV-7B. Acc. denotes accuracy (%). † marks the default setting, and bold indicates the best result within each group. Each reflection step refines the sub-queries after retrieval; the default two steps yield three retrieval rounds, including the initial round.
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, P.R. China · Peng Cheng Laboratory, Shenzhen, P.R. China · School of Electronic and Computer Engineering, Peking University, P.R. China