cs.CVSep 30, 2026

MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding

Authors: Yinying Li, Yuqian Fu, Yulin Dai, Jingyu Gong, Tianwen Qian, Xiaoling Wang

Organizations: East China Normal University · King Abdullah University of Science and Technology

Abstract

Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into fixed-size representations or by extending storage beyond GPU memory. However, these methods largely rely on global or coarse-grained representations, inevitably losing fine-grained visual information. In this work, we argue that streaming video memory should explicitly encode structured and semantically meaningful representations, particularly at the entity level. To this end, we propose MEMO, a novel framework that models streaming video through multi-level, entity-aware structured memory. MEMO performs multi-level perception to jointly capture global semantics, entity dynamics, and spatial structures, partitioning streaming video into semantically coherent chunks. Each chunk is organized into a structured memory, where lightweight global and entity-level representations serve as retrieval indices, while the corresponding high-resolution visual content is retained separately for on-demand access. At inference time, MEMO performs query-specific retrieval over the structured memory and selectively recalls relevant visual evidence for downstream reasoning. Notably, MEMO is training-free and plug-and-play with existing multimodal large language models. Extensive experiments on StreamingBench and OVO-Bench demonstrate that MEMO consistently improves multiple base models and achieves state-of-the-art performance.

Figures & tables

Explore similar work

CardsList
  1. StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

    Aug 11, 2026Muxin Fu, Yifan Zhang, Wentao Zhang +5Streaming Video UnderstandingLong-Video Benchmarks

  2. Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

    Sep 3, 2026Hongyu Qu, Guangming Yao, Ling Xing +7Streaming Video UnderstandingModern Data-Streaming Systems

  3. FOLIO: Focused Semantic Memory for Streaming Video Understanding

    Jul 14, 2026Haoyang Fan, Dhruv Parikh, Anvitha Ramachandran +4Streaming Video UnderstandingVideo Understanding