VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding
Authors: Yongchao Xu, Bowen Ye, Jiefeng Gan, Junkai Ma, Wenzhao Li, Sen Tao, Yi Wei, Jiawei Liu
Organizations: Alibaba Group, Hangzhou, China · University of Science and Technology of China, Hefei, China · Shanghai Jiao Tong University, Shanghai, China · University of Chinese Academy of Sciences, Beijing, China · Beihang University, Beijing, China
Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations. However, most memory-based methods dynamically adapt how information is retrieved for different questions, while largely fixing what is remembered. This mismatch makes missing details costly to recover, whereas stored information is valuable only when it can be reliably retrieved. To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding. Specifically, starting from a coarse low-frame-rate overview, VideoEvolve couples a Memory Evolver for selective memory augmentation with a Retrieval Evolver for adaptive retrieval over the evolving memory. We then co-evolve the two Evolvers through alternating agentic reinforcement learning (Agentic RL), updating one while freezing the other. To steer this alternating evolution, Bottleneck-Aware Evolution Feedback (BEF) identifies whether the current bottleneck lies in memory or retrieval and directs optimization toward the more limiting side. Furthermore, VideoEvolve introduces Capability-Aware Evolution Feedback (CEF) to alleviate downstream feedback from over-specializing memory to a fixed set of training questions, shifting training toward underdeveloped yet learnable video capabilities. By integrating Agentic RL with BEF and CEF, VideoEvolve transforms downstream reasoning experience into transferable capability updates, providing a concrete path from static long-video systems toward experience-driven, self-improving multimodal intelligence. Extensive experiments on multiple long video understanding benchmarks demonstrate the effectiveness of VideoEvolve.
Figures & tables
Figure 1: Comparison between conventional memory-based methods and VideoEvolve. Conventional methods largely rely on predefined memory construction and decoupled construction–retrieval optimization. VideoEvolve instead co-evolves what to remember and how to retrieve through alternating Agentic RL and downstream reasoning feedback. BEF targets the more limiting component, while CEF prioritizes underdeveloped yet learnable capabilities.
Figure 2: Overview of the VideoEvolve. VideoEvolve alternates between two stages: (a) Memory Evolver learns what to remember with the Retrieval Evolver frozen, while (b) Retrieval Evolver learns how to retrieve over the updated memory. Bottleneck-Aware Evolution Feedback (BEF) and Capability-Aware Evolution Feedback (CEF) further diagnose the current bottleneck and capability frontier to guide the next-cycle evolution.
Model
VideoMME w/o sub
VideoMME w/ sub
LongVideo- Bench
LVBench
MLVU
MMVU
Proprietary Video-MLLMs: best-setting numbers from official reports
GPT-4o ( Hurst et al., 2024 )
71.9
77.2
66.7
34.7
64.6
66.7
Gemini 1.5 Pro ( Team et al., 2024 )
75.0
81.3
64.4
33.1
74.3
65.8
Open Reasoning Video-MLLMs: reasoning-enhanced setting ( <think> → <answer> )
Video-R1-7B ( Feng et al., 2026 )
57.6
66.0
57.4
36.9
61.6
61.3
VideoChat-R1-7B ( Li et al., 2025 )
50.4
58.2
49.2
23.8
58.7
65.0
Table 1: Video understanding comparison with Video-MLLMs. See Appendix A.8 for baseline sources and protocols; proprietary models retain their best settings. Table 2 compares memory-based methods. VideoEvolve-8B denotes an 8B Retrieval backbone. Bold / underlined : best/second-best listed open-source results; “–”: unavailable.
Method
Memory Construction
Retrieval
LVBench
Video-MME (Long)
Training-free memory-based methods
EgoRAG ( Yang et al., 2025 )
–
–
32.2
41.1
HippoMM ( Lin et al., 2026 )
Qwen2.5-VL
GPT-4o
38.2
41.6
WorldMM-GPT ( Yeo et al., 2026 )
GPT-5-mini
GPT-5
61.9
76.6
MERIT-GPT ( Choi et al., 2026 )
GPT-5-mini
GPT-5
71.8
77.7
VideoEvolve (Ours)
Qwen3.8-27B
Qwen3.8-27B
76.7
78.8
Table 2: Comparison with memory-based methods. Accuracy (%) follows source-specific settings, detailed in Appendix A.8 . † The memory model is fine-tuned on episodic annotations synthesized by Gemini 1.5 Pro and GPT-4o. “–” denotes an unreported result or unspecified model.
Model
Video-MME
LongVideoBench
LVBench
MLVU
MMVU
overall
long
overall
long
VideoEvolve
73.9
67.1
70.2
65.5
58.9
72.4
75.1
(a) Co-evolution scheme: which Evolvers are trained and how to train
Base (no evolution)
66.8
57.2
62.1
57.3
50.3
64.7
67.5
Memory evolution only
68.8
62.1
63.0
58.8
51.1
67.3
69.2
Retrieval evolution only
69.9
62.5
66.3
61.0
52.7
68.8
70.1
Table 3: Ablation of co-evolution and key components. Overall Video-MME uses the with-subtitles setting.
Cycle
Accuracy (%)
Δ Acc.
(pp)
Next-cycle training shares (%)
Memory-only
Revisit-enabled
Memory
Retrieval
0
45.6
50.3
4.7
50
50
1
48.5
52.6
4.1
35
65
2
54.9
57.1
2.2
45
55
3
57.5
58.9
1.4
60
40
Table 4: Co-evolution dynamics on LVBench. Δ Acc. is Revisit-enabled minus Memory-only accuracy. Next-cycle training shares sum to 100%.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: From Base Memory to augmented memory. Left: low-frame-rate sampling provides the fixed initial memory. Middle: the Memory Evolver selects where to observe, how many frames to request, and what information to extract. Right: accepted supplemental evidence enriches the retrieval view while preserving the Base. The timestamps and evidence cards are schematic, not a measured trajectory or a guarantee of complete coverage.
Tool
Arguments and returned evidence
get_macro_events
No filter, or super_id , or macro_ids ; enumerates Macro summaries and intervals. Filters are mutually exclusive.
get_subgraph
macro_id ; returns local event, entity, and OCR records.
Table 5: Retrieval tool interface. All temporal arguments are absolute video seconds; availability and parameter bounds are enforced by the runtime schema.
Table 6: Memory construction and optimization ablations. Accuracy (%, ↑ ) on the Long subsets of Video-MME (VMME) and LongVideoBench (LVB).
Variant
LVB Overall
LVBench
(a) Memory construction
Learned augmentation
70.2
58.9
Uniform augmentation
65.4
48.6
Random augmentation
64.7
47.9
Base memory only
64.1
47.5
(b) Optimization objective
Appendix
Table 7: Additional metrics for design ablations. Accuracy (%, ↑ ) on LongVideoBench (LVB) Overall and LVBench, complementing Table 6 . Bold marks the best result within each group.
Model
Video-MME
LongVideoBench
LVBench
MLVU
MMVU
overall
long
overall
long
VideoEvolve
73.9
67.1
70.2
65.5
58.9
72.4
75.1
No cold-start SFT
54.4
48.2
45.9
45.4
32.8
57.1
62.5
w/o Stage-I SFT
70.2
63.5
67.7
62.5
54.2
69.9
70.1
w/o Stage-II SFT
64.5
57.9
62.1
51.9
40.6
60.3
65.7
Appendix
Table 8: Cold-start initialization of the Retrieval Evolver. Overall Video-MME uses the with-subtitles setting.
Long-video understanding places substantial demands on memory, as answering questions often requires retrieving information distributed across extended temporal spans. Existing approaches broadly follow two paradigms: query-driven exploration, which is sensitive to localization errors, and query-independent memory construction, which may omit question-specific details. We introduce VideoTapestry, a training-free multi-agent framework that adapts a preconstructed hierarchical video memory through coarse-to-fine, query-driven refinement. The preconstructed memory organizes video content into three levels, capturing global narrative context, event-level temporal structure, and fine-grained relational evidence, respectively. To support coarse-to-fine localization and observation, we assign a specialized agent to each level, keeping retrieval and refinement within a scale-specific context. Guided by the query, these agents revisit relevant video regions and enrich layer-wise memories with targeted multimodal observations. Their refinements are assembled according to the original hierarchy into a composite query-adaptive memory, preserving global context in a compact form while retaining fine-grained evidence along query-relevant branches for final reasoning. Compared with direct GPT-5.5 inference, VideoTapestry achieves absolute accuracy gains of 17.2%, 14.9%, 9.8%, and 7.0% on LVBench, LongVideoBench (Long), Video-MME (Long), and EgoSchema, respectively, achieving the state-of-the-art results among all competitors.
Yucheng Liu, Yufei Yin, Mingxiao Feng +3
University of Science and Technology of China · Jilin University · Hangzhou Dianzi University +1
When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction to pre-model high-level relations in videos, despite not knowing the downstream query at build time. We instead prioritize high-recall retrievability during memory building, and defer query-specific, high-level relation composition to inference time. To this end, we propose MERIT(Multi-key Episodic Retrieval with Inference-time Temporal expansion), a simple yet effective agentic framework for ultra-long video understanding. First, we formulate an episodic multi-key representation that enables precise retrieval of fine-grained memories through a simple key-matching mechanism. Second, we introduce a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction. This is achieved by expanding the temporal scope exclusively around the retrieved segments at inference time. By leveraging simple key-matching with this on-demand temporal expansion, MERIT achieves state-of-the-art performance across three long-video benchmarks: EgoLifeQA, LVBench, and Video-MME (Long).
Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps it searchable over long horizons, and grounds later reasoning in recoverable observations rather than compressed latent state alone. We propose Visual Agentic Memory (VAM), a training-free framework with three components. Online Indexing supports selective evidence retention under streaming constraints. Hierarchical Memory organises retained evidence in a Parallel Representation that aligns temporal context with spatial observations. Agentic Retrieval searches, inspects, and verifies candidate evidence before producing a grounded answer. On OVO-Bench, VAM achieves the highest RT+BT average (68.41) across all reported baselines, improving over end-to-end use of the same underlying MLLM (Gemini 3 Flash, 67.46). On the month-scale split of MM-Lifelong train@month (105.6 hours over 51 days), VAM reaches 17.11%, second only to ReMA with GPT-5 (17.62%). These results suggest that long-horizon video understanding benefits from treating visual memory as an explicit, inspectable, and queryable substrate. Code is available at https://github.com/yiliu-li/Visual-Agentic-Memory.