Organizations: Department of Electrical and Computer Engineering, Northeastern University · Khoury College of Computer Science, Northeastern University
Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation--Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.
Figures & tables
Figure 1: Two key challenges in long-video frame selection. Answering “Why did the milk spill?” requires evidence of the cat tipping the glass and the resulting spill. Query Comprehension Gap (left): similarity-based methods overlook implicit causal requirements, retrieving spilled-milk frames while missing the cat tipping the glass. Interpretation–Selection Gap (right): judgment-based methods understand implicit evidence requirements but fail to perform selection reliably, choosing a milk-pouring frame instead of the causal action.
Figure 2: Empirical analysis of two gaps on LongVideoBench subsets. (a) Embedding models such as CLIP handle single-evidence queries effectively but struggle with implicit requirements in complex queries, which query decomposition helps make explicit. (b) VLMs are more effective at interpreting complex queries and describing relevant evidence than at directly selecting frames, as using these interpretations to guide tool-based selection yields higher accuracy and lower latency than direct VLM selection, especially for long videos.
Figure 3: Overview of RACER. RACER begins with (a) Query Interpretation , where a lightweight Vid-LLM uses a uniformly sampled overview to decompose the query into complementary sub-queries, making implicit evidence requirements explicit. Then, (b) Evidence Localization employs an embedding-based retrieval tool to score candidate frames against the original query and sub-queries, followed by entropy-adaptive score fusion and DPP selection to obtain relevant and visually diverse key frames. Finally, (c) Reflection merges feedback frames with the overview, removing duplicates and the least visually aligned overview frames to maintain the frame budget. The resulting hybrid overview guides sub-query refinement for the next selection round.
Table 4
Vid-LLMs
Frame Selection Method
Video-MME (1–60 min)
MLVU (3–120 min)
LongVideoBench (1–60 min)
Average
LLaVA-OV-7B
Uniform
58.5 (+0.0)
62.4 (+0.0)
56.6 (+0.0)
59.2 (+0.0)
AKS ( Tang et al., 2025 )
58.4 (-0.1)
66.8 (+4.4)
59.3 (+2.7)
61.5 (+2.3)
BOLT ( Liu et al., 2025 )
59.9 (+1.4)
66.8 (+4.4)
59.6 (+3.0)
62.1 (+2.9)
MDP 3 ( Sun et al., 2025 )
60.5 (+2.0)
68.3 (+5.9)
59.0 (+2.4)
62.6 (+3.4)
AdaQ ( Zhang et al., 2026 )
59.5 (+1.0)
69.4 (+7.0)
60.0 (+3.4)
63.0 (+3.8)
GIFT ( Ma et al., 2026 )
61.2 (+2.7)
69.4 (+7.0)
59.6 (+3.0)
63.4 (+4.2)
Table 3: Comparison with state-of-the-art frame-selection methods across three benchmarks and four Vid-LLMs: accuracy (%) for Video-MME and LongVideoBench, and M-Avg for MLVU. Parentheses show gains over uniform sampling (positive in green). Average is the unweighted mean; bold / underline mark best/second-best results per Vid-LLM. RACER achieves the best or tied-best performance across all benchmarks and answering models.
Answering Vid-LLM
Frame Selection
Retrieval Tool
Query Interpretation Model
Video-MME
Short
Medium
Long
Overall
Qwen2.5-VL-3B
Uniform Sampling
N/A
N/A
69.5
57.0
47.1
57.8
LLaVA-OV-7B
Uniform Sampling
N/A
N/A
70.2
57.1
48.2
58.5
+ Tool-based Retrieval
CLIP ViT-L/14
N/A
71.2
58.7
51.8
60.6
+ Query Interpretation
CLIP ViT-L/14
Qwen2.5-VL-3B
72.3
60.1
52.1
61.5
+ Reflection (RACER)
CLIP ViT-L/14
Qwen2.5-VL-3B
72.6
60.6
52.7
61.9
Table 4: Module ablation of RACER on Video-MME with components added cumulatively. Scores report accuracy (%) by duration and overall; bold marks column bests. Each component improves performance, with full RACER performing best. Despite weaker direct answering capability, Qwen2.5-VL-3B improves LLaVA-OV-7B when assigned Query Interpretation.
Table 7
Figure 4: Qualitative visualization of RACER. (a) Query interpretation guides tool-based evidence localization to recover the explanation that water expands upon freezing, missed by direct CLIP retrieval and VLM judgment. Full RACER retains this evidence with less redundancy. (b) Reflection refines sub-queries to target the blue object, recovering a previously missed close-up and correcting the answer from a clipper to a dremel .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Perception
Recognition
Reasoning
Information Synopsis
OCR Problems
Counting Problem
Uniform
68.9
60.8
53.2
74.3
66.2
36.6
CLIP Retrieve
69.2
59.5
53.9
66.9
66.2
37.3
VLM Judgment
70.7
62.1
55.8
75.2
66.2
36.2
RACER (ours)
71.9
63.4
57.8
75.8
67.6
41.0
Appendix
Table 7: Accuracy (%) across six Video-MME question categories with LLaVA-OV-7B.
Method
Short (1–3 min)
Medium (3–30 min)
Long (30–60 min)
Overall
Acc.
Time (s)
#Frames
Acc.
Time (s)
#Frames
Acc.
Time (s)
#Frames
Acc.
Time (s)
#Frames
Uniform
70.2
1.1
0
57.1
1.1
0
48.2
1.1
0
58.5
1.1
0
CLIP Retrieval
68.8
3.5
0
58.6
7.5
0
45.8
32.8
0
57.7
14.6
0
VLM judgment
70.7
22.8
157.7
58.0
58.4
516.3
51.7
286.2
2465.9
60.1
122.4
1046.6
RACER
72.6
22.4
57.6
60.6
29.0
64.0
52.7
54.9
67.1
61.9
35.4
62.9
Appendix
Table 8: Efficiency comparison on Video-MME across video durations. Time is measured in seconds, and #Frames denotes the average number of unique frames analyzed by the VLM.
DPP Relevance
Entropy Sensitivity
Original-Query Weight
Generation Temperature
α
Acc.
β
Acc.
w0
Acc.
τ
Acc.
0.55
60.41
2.0
62.11
0.4
61.37
0.6
60.30
0.65
60.63
2.1
61.96
0.5
61.22
0.7
61.00
0.75†
61.93
2.2†
61.93
0.6†
61.93
0.8†
61.93
0.85
60.89
2.3
61.89
0.7
61.22
0.9
60.59
0.95
61.52
2.4
61.93
0.8
60.48
1.0
60.89
Appendix
Table 9: Hyperparameter ablation on Video-MME with LLaVA-OV-7B. Each column pair reports a separate parameter sweep. Acc. denotes accuracy (%); † marks the default setting; bold indicates the best observed result.
Design
Configuration
Acc. (%)
Sub-query aggregation
Uniform weighting
61.67
Entropy-based weighting †
61.93
Reflection steps
0
61.52
1
61.89
2†
61.93
3
61.96
Appendix
Table 10: Ablations of sub-query aggregation, reflection depth, and reflection input on Video-MME with LLaVA-OV-7B. Acc. denotes accuracy (%). † marks the default setting, and bold indicates the best result within each group. Each reflection step refines the sub-queries after retrieval; the default two steps yield three retrieval rounds, including the initial round.
Multimodal Large Language Models (MLLMs) have made strong progress in video understanding, yet long videos remain difficult: the visual token budget grows with video length, so temporally sparse evidence is easily lost. Existing methods compress the input through uniform sampling or frame selection, but these strategies optimize different objectives, either broad temporal coverage or local question relevance, and neither preserves both global storyline context and fine-grained evidence. We propose VideoRouter (VR), which rethinks long-video understanding as coordinating complementary evidence views rather than selecting a single subset of frames. VideoRouter first organizes each video into a question-agnostic temporal hierarchy that partitions it into coarse-to-fine temporally coherent segments. Upper-level nodes capture broad storyline context and event progression, while lower-level nodes preserve fine-grained local details and evidence-bearing moments. This gives rise to two complementary views: a global view for coverage-oriented reasoning and a local view for detail-oriented evidence recovery. We further introduce a verification-guided router that judges which view is better supported by its own selected evidence and decides the final answer. Across six backbones, routing improves over both views in all settings, and the choice of view is shown to be dataset-dependent, confirming that no single evidence granularity is universally preferable. On VideoMME, our method outperforms state-of-the-art frame selection methods by 2.5 points, under the LLaVA-Video-7B backbone. We will release the code.
Long videos contain far more visual content than Large Vision-Language Models (LVLMs) can process under a fixed visual-token budget, making frame selection essential. Existing query-aware selectors usually estimate frame-query relevance and build a compact subset from high-scoring frames. Although their mechanisms differ, the similarity sequence is still often treated primarily as values to rank or sample from, rather than as an ordered signal whose shape reflects how query-relevant evidence emerges, peaks, and fades over time. This can obscure frames that explain, contextualize, or follow an event, because such evidence may lie on the rising or falling sides of a nearby relevance peak and receive lower absolute scores. We propose RIDGE, a frame selection framework that reads the frame-query similarity curve as a temporal signal. By using local changes and curvature, RIDGE partitions the timeline into structural regions and applies region-specific selection to preserve event cores, transitions, buildup, aftermath, and contextual frames under a fixed budget. It is a lightweight post-processing step on precomputed frame-query scores and requires neither training nor iterative LVLM calls. Across four long-video benchmarks and three backbones, RIDGE achieves the best performance in most settings and remains competitive in the others.
Shanqing Xu, Meng Luo, Mengchen Qian +7
Huazhong University of Science and Technology · National University of Singapore · Fudan University
Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computationally expensive. This challenge is further amplified when only a small subset of frames is truly relevant to the given query. In this paper, we propose a Query- and Content-Aware (QCA) keyframe selection framework that can select a compact yet information-rich set of frames from long videos. QCA first partitions the video into temporal segments and estimates the information contribution of each segment by jointly modeling query relevance and content deviation, and dynamically allocates keyframe budget to each segment. Within each segment, QCA anchors on the most query-relevant frame and iteratively incorporates additional frames to maximize diversity while maintaining high semantic relevance to the query. Crucially, our method requires no additional training and can be seamlessly integrated into existing Video-LLMs. Extensive experiments across multiple long video understanding benchmarks demonstrate that our proposed approach achieves state-of-the-art performance and has strong generalization ability. For instance, QCA achieves 67.8% on LongVideoBench using 128 frames, while GPT-4o achieves 66.7% using 256 frames. Our codes are available in GitHub.
Jun Peng, Baiyang Song, Jie Li +4
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, P.R. China · Peng Cheng Laboratory, Shenzhen, P.R. China · School of Electronic and Computer Engineering, Peking University, P.R. China