MemEvo: Automatic Discovery of Streaming Video Memory Mechanisms
Organizations: Institute for AI Industry Research, Tsinghua University · Peking University
Abstract
Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism--what observations to preserve, how to represent and consolidate them, and what information to retrieve when a query eventually arrives. Rather than designing a single memory architecture by hand, we formulate memory design as a search problem over executable memory programs. We introduce a lightweight domain-specific language that expresses memory mechanisms through structured primitives for representation, admission, retention, consolidation, budgeting, and retrieval, while enforcing causal and bounded-memory constraints. Although structured, the derived program space remains large and contains heterogeneous, conditionally dependent design choices whose effects can only be assessed via downstream execution. We therefore propose MemEvo, an LLM-driven auto-research framework that uses pretrained LLM as a semantics-aware proposal model to iteratively generate and refine candidate memory programs based on accumulated experimental feedback. At runtime, a deterministic evaluation pipeline validates and evaluates each candidate, while the underlying vision-language model remains frozen throughout discovery. We finally produce a training-free, bounded-memory mechanism. Extensive experiments on StreamingBench and OVO-Bench demonstrate strong streaming video understanding performance together with substantial context and inference efficiency.
Figures & tables
| StreamingBench | OVO-Bench | |||||
| Method | Size | #Frames | RTVU | RT | BT | Avg. |
| Training-based Methods | ||||||
| VideoLLM-online | 8B | 2 fps | 36.0 | 20.8 | 17.7 | 19.3 |
| Flash-VStream | 7B | 1 fps | 23.2 | 28.4 | 27.4 | 27.9 |
| Dispider | 7B | 1 fps | 67.6 | 54.6 | 36.1 | 45.4 |
| TimeChat-Online | 7B | 1 fps | 75.4 | 61.9 | 41.7 | 51.8 |
| Method | OP | CR | CS | ATP | EU | TR | PR | SU | ACP | CT | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Training-based Methods | |||||||||||
| VideoLLM-Online | 39.1 | 40.1 | 34.5 | 31.1 | 46.0 | 32.4 | 31.5 | 34.2 | 42.5 | 27.9 | 36.0 |
| Flash-VStream | 25.9 | 43.6 | 24.9 | 23.9 | 27.3 | 13.1 | 18.5 | 25.2 | 23.9 | 48.7 | 23.2 |
| Dispider | 74.9 | 75.5 | 74.1 | 73.1 | 74.4 | 59.9 | 76.1 | 62.9 | 62.2 | 45.8 | 67.6 |
| TimeChat-Online | 80.8 | 79.7 | 80.8 | 83.3 | 74.8 | 78.8 | 78.7 | 64.2 | 68.8 | 58.0 | 75.3 |
| StreamForest | 83.1 | 82.8 | 82.7 | 84.3 | 77.5 | 78.2 | 76.9 | 69.1 | 75.6 | 54.4 | 77.3 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Search round / DSL edit | Trials | Best Score | Main outcome |
| E1: : capacity retention | 11 | 0.748356 | Explore capacity and FIFO/temporal/reservoir retention | |
| E2: | 12 | 0.752949 | Compare interval, histogram, SSIM, and optical-flow triggers | |
| E3: admission gap constraints | 5 | 0.753435 | Refine histogram- admission with minimum/maximum gaps | |
| E4: : coverage-change retention | 9 | 0.753435 | Compare histogram/SSIM/flow coverage criteria and weights | |
| E5: local program recombination | 3 | 0.740043 | Recombine promising admission, gap, retention, and capacity choices | |
| Retrieval | , | 15 | 0.771990 | Select |
| Method | Base Model | Method Source | Data Source |
|---|---|---|---|
| Training-based Methods | |||
| VideoLLM-Online | Llama-3-8B | CVPR 2024 | StreamForest paper |
| Flash-VStream | Qwen2-VL-7B | ICCV 2025 | OASIS paper |
| Dispider | Qwen2-7B | CVPR 2025 | OASIS paper |
| TimeChat-Online | Qwen2.5-VL-7B | ACM MM 2025 | HERMES paper |
| StreamForest | Qwen2-7B | NeurIPS 2025 Spotlight | Official |
| # Frames | Retrieval mean (ms) | Packing mean/p95 (s) | Peak GPU Memory (GiB) |
|---|---|---|---|
| 64 | 33.2 | 1.139 / 1.289 | 18.875 |
| 128 | 33.3 | 1.080 / 1.203 | 18.880 |
| 256 | 33.9 | 1.142 / 1.283 | 18.878 |
| 512 | 34.0 | 1.144 / 1.239 | 18.890 |
| 1024 | 32.7 | 1.164 / 1.247 | 18.893 |
| Metric | Mean | Median | p95 |
|---|---|---|---|
| Online (ms/frame) | 911.35 | 937.39 | 1061.25 |
| Sustainable FPS | 1.124 | 1.067 | 1.483 |
| Streaming update (ms/frame) | 571.07 | 568.52 | 640.43 |
| Post-vision memory (ms/frame) | 472.87 | 483.46 | 543.54 |