Understanding long videos is crucial for embodied intelligent agents, as their performance depends on effectively accumulating and using long-horizon perceptual memories. Multimodal large language models (MLLMs) are increasingly used for long-video understanding, but their performance degrades and inference time increases as more frames are provided. Therefore, selecting informative keyframes is essential for efficient question answering over long videos. In this work, we develop FocusGraph, a framework for keyframe selection in egocentric long-video question answering. It includes a lightweight Scene-Graph LLM Selector that identifies query-relevant clips from compact graph-based captions, avoiding the need to process raw frame sequences at question time. From these clips, we extract keyframes using Patch-wise Sparse-Flow Retention (PSFR), an offline program-evolved method with no learned parameters at inference time, before passing them to an MLLM for answer generation. FocusGraph achieves state-of-the-art performance on FindingDory and HourVideo while reducing question-time inference cost compared with existing approaches.
Figures & tables
Fig. 1: FocusGraph retrieves relevant intervals from a reusable scene-graph memory and uses PSFR to select non-redundant original frames for answer generation.
Fig. 2: FocusGraph overview. FocusGraph partitions an egocentric video into temporal clips represented by object-centric scene graphs generated with the SG-Ego pipeline [ 33 ] . The graph triplets are encoded into a compact chronological memory and processed jointly with the question by Scene-Graph LLM Selector, which generates a temporally grounded answer and identifies relevant video intervals. PSFR then selects up to K informative, non-redundant frames for final answer generation by a downstream MLLM.
Frame Selection Method
Scene Representation
QA MLLM
Frame Number
All tasks
Single-Goal Spatial Tasks
Single-Goal Temporal Tasks
Multi-Goal Tasks
full val.
dev.
full val.
dev.
full val.
dev.
full val.
dev.
Uniform Sampling
image
Gemini-3-Flash [ 46 ]
8
-
26.42
-
31.82
-
30.60
-
2.1
Uniform Sampling
image
Qwen-2.5-VL-7B [ 47 ]
8
12.6
18.4
16.4
22.6
11.8
19.4
2.0
3.2
Uniform Sampling
image
GPT-5-Mini [ 48 ]
8
15.49
24.25
20.37
29.27
14.04
27.38
2.71
2.78
Uniform Sampling
image
Qwen3.8-27B [ 49 ]
8
18.12
28.97
23.38
34.57
17.08
34.08
3.11
2.55
MaxInfo [ 13 ]
image
Qwen-2.5-VL-7B [ 47 ]
16
13.7
17.2
17.5
19.5
13.4
22.2
1.9
2.1
TABLE I: Performance comparison across frame sampling methods and LLM settings on FindingDory [ 6 ] .
Frame Selection Method
Frame No.
Overall
Nav.
Perc.
Reason.
Summ.
Uniform Sampling
8
29.86
10.0
33.7
27.9
37.8
Uniform Sampling
16
31.22
15.0
36.8
27.5
43.9
MaxInfo [ 13 ]
16
29.61
17.5
33.9
26.6
40.2
ReMEmbR [ 2 ]
8 1 1 1 For ReMEmbR, frame number is the average number of documents retrieved from the database per query.
24.28
17.5
24.54
27.47
0.0
VideoTree [ 51 ]
8
31.4
22.5
32.4
30.4
39.0
LVNet [ 52 ]
8
24.79
27.50
22.72
23.93
40.24
TABLE II: Results on HourVideo [ 4 ] . All methods use Qwen-2.5-VL-7B [ 47 ] except ReMEmbR [ 2 ] , which uses Llama3.1-8B [ 50 ] .
Method / stage
Scope
Tok./fr.
Time, ms
Reusable window processing
Frame-level relation extraction
Window
–
59180±14430
Object grounding and frame-graph construction
Window
–
10440±6200
Temporal tracking and graph consolidation
Window
–
16230±9890
MaxInfo [ 13 ]
Window
–
3420
PSFR visual feature extraction (ours)
Window
–
4530±2440
TABLE III: Inference latency on HourVideo [ 4 ] .
Fig. 3: Comparison of frame sampling methods. Left: FindingDory [ 6 ] . Green-bordered frames are present in the ground truth; red-bordered frames are not. Right: HourVideo [ 4 ] . Correct answers are highlighted in green, incorrect ones in red.
Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant frames. Existing score-based methods, however, typically search within a fixed, uniformly sampled candidate pool, preventing evidence outside this pool from ever being selected. Given a limited relevance-scoring budget, the key challenge is to allocate evaluations adaptively to promising frames while continuing to explore underrepresented temporal regions. We introduce FORTE, a training-free framework that addresses this challenge through two stages: adaptive relevance scoring and global keyframe optimization. Starting from sparse, uniformly distributed observations, our efficient Gaussian-process relevance predictor estimates relevance for unscored frames, exploiting temporal locality and the approximately banded kernel structure to reduce the core computation from cubic to linear time in the number of frames for fixed bandwidth. The scoring stage then selects which frames to score next by balancing predicted relevance with temporal coverage, prioritizing promising regions while also exploring less-represented parts of the video. The optimization stage selects the final keyframes by maximizing an objective that jointly captures measured relevance and temporal coverage. We derive an exact algorithm that leverages the logarithmic coverage structure to identify the optimal subset of the scored candidate pool in time linear in the pool size, for a fixed final-frame budget. Experiments on four long-video question-answering benchmarks show that FORTE achieves the highest observed mean accuracy among the compared selectors under every tested scoring budget. Further evaluations demonstrate its consistent effectiveness across different relevance scorers and downstream MLLMs.
Video understanding is often plagued by severe temporal redundancy, where processing dense frame sequences is both semantically inefficient and computationally expensive. This challenge is further amplified when only a small subset of frames is truly relevant to the given query. In this paper, we propose a Query- and Content-Aware (QCA) keyframe selection framework that can select a compact yet information-rich set of frames from long videos. QCA first partitions the video into temporal segments and estimates the information contribution of each segment by jointly modeling query relevance and content deviation, and dynamically allocates keyframe budget to each segment. Within each segment, QCA anchors on the most query-relevant frame and iteratively incorporates additional frames to maximize diversity while maintaining high semantic relevance to the query. Crucially, our method requires no additional training and can be seamlessly integrated into existing Video-LLMs. Extensive experiments across multiple long video understanding benchmarks demonstrate that our proposed approach achieves state-of-the-art performance and has strong generalization ability. For instance, QCA achieves 67.8% on LongVideoBench using 128 frames, while GPT-4o achieves 66.7% using 256 frames. Our codes are available in \href{https://github.com/hktk07/QCA}{GitHub}.
Jun Peng, Baiyang Song, Jie Li +4
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, P.R. China · Peng Cheng Laboratory, Shenzhen, P.R. China · School of Electronic and Computer Engineering, Peking University, P.R. China
Recent multimodal large language models (MLLMs) have substantially advanced video understanding, yet long-form video QA remains challenging under fixed input token budgets, where uniform sampling can be inefficient for evidence localization. We propose ReQuest , an uncertainty-driven, question-adaptive keyframe selection pipeline that aligns question intent with relevant video content through selective computation. ReQuest integrates (i) a lightweight question-aware selector distilled from MLLM-generated supervision, (ii) Re-thinking Routing that triggers additional inference only when the model is uncertain with a length-adaptive criterion, and (iii) uncertainty-guided adaptive non-maximum suppression that selects temporally diverse frames while adjusting spacing based on question difficulty. As a plug-andplay method, ReQuest improves long-video QA without modifying or fine-tuning the underlying MLLM. Experiments on Video-MME, MLVU, and LongVideoBench demonstrate consistent accuracy gains with competitive computational cost, with particularly strong improvements in medium and long video regimes.
Minkuk Kim, Suyong Yun, Young Tae Kim +3
Kyung Hee University, Republic of Korea · Electronics and Telecommunications Research Institute (ETRI), Republic of Korea