cs.CVMar 4, 2026

FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering

Authors: Tatiana Zemskova, Solomon Andryushenko, Ilya Obrubov, Viktoriia Khoruzhaia, Ekaterina Eroshenko, Ekaterina Derevyanka, Dmitry Yudin

Organizations: AXXX · MIRAI · Yandex · FusionBrain Lab

Abstract

Understanding long videos is crucial for embodied intelligent agents, as their performance depends on effectively accumulating and using long-horizon perceptual memories. Multimodal large language models (MLLMs) are increasingly used for long-video understanding, but their performance degrades and inference time increases as more frames are provided. Therefore, selecting informative keyframes is essential for efficient question answering over long videos. In this work, we develop FocusGraph, a framework for keyframe selection in egocentric long-video question answering. It includes a lightweight Scene-Graph LLM Selector that identifies query-relevant clips from compact graph-based captions, avoiding the need to process raw frame sequences at question time. From these clips, we extract keyframes using Patch-wise Sparse-Flow Retention (PSFR), an offline program-evolved method with no learned parameters at inference time, before passing them to an MLLM for answer generation. FocusGraph achieves state-of-the-art performance on FindingDory and HourVideo while reducing question-time inference cost compared with existing approaches.

Figures & tables

Explore similar work

CardsList
  1. QCA: Query- and Content-Aware Keyframe Selection for Long Video Understanding

    Jul 1, 2026Jun Peng, Baiyang Song, Jie Li +4Video UnderstandingKeyframe Selection

  2. ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA

    Jul 2, 2026Minkuk Kim, Suyong Yun, Young Tae Kim +3Long Video Question AnsweringKeyframe Selection