cs.AISep 29, 2026

MemEvo: Automatic Discovery of Streaming Video Memory Mechanisms

Authors: Guohong Liu, Jialei Ye, Shanhui Zhao, Yunxin Liu, Yuanchun Li

Organizations: Institute for AI Industry Research, Tsinghua University · Peking University

Abstract

Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism--what observations to preserve, how to represent and consolidate them, and what information to retrieve when a query eventually arrives. Rather than designing a single memory architecture by hand, we formulate memory design as a search problem over executable memory programs. We introduce a lightweight domain-specific language that expresses memory mechanisms through structured primitives for representation, admission, retention, consolidation, budgeting, and retrieval, while enforcing causal and bounded-memory constraints. Although structured, the derived program space remains large and contains heterogeneous, conditionally dependent design choices whose effects can only be assessed via downstream execution. We therefore propose MemEvo, an LLM-driven auto-research framework that uses pretrained LLM as a semantics-aware proposal model to iteratively generate and refine candidate memory programs based on accumulated experimental feedback. At runtime, a deterministic evaluation pipeline validates and evaluates each candidate, while the underlying vision-language model remains frozen throughout discovery. We finally produce a training-free, bounded-memory mechanism. Extensive experiments on StreamingBench and OVO-Bench demonstrate strong streaming video understanding performance together with substantial context and inference efficiency.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding

    Jun 24, 2026Baiyang Song, Yuli Lin, Qiong Wu +5Streaming Video UnderstandingVideo Understanding

  2. MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding

    Sep 30, 2026Yinying Li, Yuqian Fu, Yulin Dai +3Streaming Video UnderstandingStreaming Video

  3. Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

    May 8, 2026Hang Wu, Sherin Mary Mathews, Yujun Cai +2Streaming Video UnderstandingModern Data-Streaming Systems