FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
Organizations: AXXX · MIRAI · Yandex · FusionBrain Lab
Abstract
Understanding long videos is crucial for embodied intelligent agents, as their performance depends on effectively accumulating and using long-horizon perceptual memories. Multimodal large language models (MLLMs) are increasingly used for long-video understanding, but their performance degrades and inference time increases as more frames are provided. Therefore, selecting informative keyframes is essential for efficient question answering over long videos. In this work, we develop FocusGraph, a framework for keyframe selection in egocentric long-video question answering. It includes a lightweight Scene-Graph LLM Selector that identifies query-relevant clips from compact graph-based captions, avoiding the need to process raw frame sequences at question time. From these clips, we extract keyframes using Patch-wise Sparse-Flow Retention (PSFR), an offline program-evolved method with no learned parameters at inference time, before passing them to an MLLM for answer generation. FocusGraph achieves state-of-the-art performance on FindingDory and HourVideo while reducing question-time inference cost compared with existing approaches.
Figures & tables
| Frame Selection Method | Scene Representation | QA MLLM | Frame Number | All tasks | Single-Goal Spatial Tasks | Single-Goal Temporal Tasks | Multi-Goal Tasks | ||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| full val. | dev. | full val. | dev. | full val. | dev. | full val. | dev. | ||||
| Uniform Sampling | image | Gemini-3-Flash [ 46 ] | 8 | - | 26.42 | - | 31.82 | - | 30.60 | - | 2.1 |
| Uniform Sampling | image | Qwen-2.5-VL-7B [ 47 ] | 8 | 12.6 | 18.4 | 16.4 | 22.6 | 11.8 | 19.4 | 2.0 | 3.2 |
| Uniform Sampling | image | GPT-5-Mini [ 48 ] | 8 | 15.49 | 24.25 | 20.37 | 29.27 | 14.04 | 27.38 | 2.71 | 2.78 |
| Uniform Sampling | image | Qwen3.8-27B [ 49 ] | 8 | 18.12 | 28.97 | 23.38 | 34.57 | 17.08 | 34.08 | 3.11 | 2.55 |
| MaxInfo [ 13 ] | image | Qwen-2.5-VL-7B [ 47 ] | 16 | 13.7 | 17.2 | 17.5 | 19.5 | 13.4 | 22.2 | 1.9 | 2.1 |
| Frame Selection Method | Frame No. | Overall | Nav. | Perc. | Reason. | Summ. |
|---|---|---|---|---|---|---|
| Uniform Sampling | 8 | 29.86 | 10.0 | 33.7 | 27.9 | 37.8 |
| Uniform Sampling | 16 | 31.22 | 15.0 | 36.8 | 27.5 | 43.9 |
| MaxInfo [ 13 ] | 16 | 29.61 | 17.5 | 33.9 | 26.6 | 40.2 |
| ReMEmbR [ 2 ] | 8 1 1 1 For ReMEmbR, frame number is the average number of documents retrieved from the database per query. | 24.28 | 17.5 | 24.54 | 27.47 | 0.0 |
| VideoTree [ 51 ] | 8 | 31.4 | 22.5 | 32.4 | 30.4 | 39.0 |
| LVNet [ 52 ] | 8 | 24.79 | 27.50 | 22.72 | 23.93 | 40.24 |
| Method / stage | Scope | Tok./fr. | Time, ms |
| Reusable window processing | |||
| Frame-level relation extraction | Window | – | |
| Object grounding and frame-graph construction | Window | – | |
| Temporal tracking and graph consolidation | Window | – | |
| MaxInfo [ 13 ] | Window | – | 3420 |
| PSFR visual feature extraction (ours) | Window | – | |
| Method | Sent. Sim. | Score | ||||
| GPT-4o [ 54 ] (60 frames, z.s.) | 18.9 | 24.4 | 12.0 | 12.2 | 73.7 | 5.4 |
| InternVL2-8B [ 55 ] (30 frames, z.s.) | 11.8 | 24.0 | 6.3 | 6.6 | 71.9 | 4.5 |
| LLaVA-NeXT-Video-7B [ 56 ] (32 frames, z.s.) | – | – | – | – | 62.1 | 4.2 |
| TimeChat-7B [ 57 ] (96 frames, z.s.) | 10.2 | 5.6 | 3.0 | 3.6 | 58.9 | 3.3 |
| VTimeLLM-7B [ 58 ] (100 frames, z.s.) | 12.4 | 28.2 | 8.8 | 9.2 | 70.5 | 4.3 |
| LLaVA-NeXT-7B [ 56 ] Llama-3.1-8B [ 50 ] (180 frames, z.s.) | 21.4 | 22.3 | 10.1 | 9.7 | 63.6 | 3.5 |
| Clip selection | Frame sampling | Frame number | All tasks | Single-goal spatial | Single-goal temporal | Multi-goal tasks |
|---|---|---|---|---|---|---|
| Uniform image baseline | Uniform | 8 | 12.60 | 16.40 | 11.80 | 2.00 |
| Scene-Graph LLM Selector | Uniform | 8 | 15.76 | 21.72 | 12.65 | 3.15 |
| Scene-Graph LLM Selector | PSFR | 8 | 17.25 | 23.47 | 14.50 | 3.00 |
| FindingDory (FD) | GenS | Time (s) | |||||
| Keyframe selector | Opt. | Inter. | Incl. | w-Inter | Incl | ||
| Uniform | - | 0.111 | 0.715 | 0.0136 | 0.470 | 0.890 | – |
| Hand-Crafted | FD | 0.075 | 0.380 | 0.008 | 0.501 | 0.840 | |
| Grid-Search | FD | 0.132 | 0.714 | 0.016 | 0.488 | 0.832 | |
| Evolved on Inter. | FD | 0.212 | 0.645 | 0.0289 | 0.449 | 0.803 | |
| Evolved on | FD | 0.181 | 0.834 | 0.0250 | 0.426 | 0.770 | |
| Short | Medium | Long | ||||
| Questions | 46 | 52 | 65 | |||
| Method | Score | EM | Score | EM | Score | EM |
| Uniform | 2.000 | 10.9 | 1.865 | 15.4 | 2.308 | 12.3 |
| FocusGraph | 3.717 | 45.7 | 2.673 | 28.8 | 2.508 | 10.8 |