cs.CVOct 5, 2026

ReMem: Streaming Video Understanding With Long Context Retention

Authors: Li Yiheng, He Xu, Wang Shaobo, Shao Ling, Lu Shijian

Organizations: Nanyang Technological University · Shanghai Jiaotong University · University of the Chinese Academy of Sciences

Abstract

Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offline scenarios and struggle to handle online streaming videos that demand low latency response. Several studies have explored memory and token compression strategies in an attempt to adapt offline VLMs for streaming video understanding tasks. However, through our probing experiment, we identify that most existing works tend to progressively lose long context information as length of input stream increases. To address this, we propose ReMem, a novel training-free adaptation technique that enables VLMs to process streaming videos of arbitrary lengths while improving their long context information retention capability. ReMem exploits memory from two perspectives, implemented as two core components. The Streaming Context Memory (SCM) continuously compresses historical context with query-independent attention. The Retrieved Vision Memory (RVM) then retrieves the most salient, query-relevant context from memory to augment the VLM's input. Comprehensive experiments demonstrate that the proposed ReMem achieves state-of-the-art (SOTA) performance across a variety of widely used benchmarks, spanning both streaming video and general long video understanding tasks.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

    Aug 11, 2026Muxin Fu, Yifan Zhang, Wentao Zhang +5Streaming Video UnderstandingLong-Video Benchmarks

  2. Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

    Sep 3, 2026Hongyu Qu, Guangming Yao, Ling Xing +7Streaming Video UnderstandingModern Data-Streaming Systems

  3. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    Jul 30, 2026Mingkang Dong, Muxin Pu, Jie Li +8Streaming Video UnderstandingVideo Understanding