Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (long). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing: on the hardest quarter of VideoMME (long) questions, a blind judge that reads only the agent's context answers 81.1% correctly, versus 60.9% for the append-only agent. With Gemini 3.1 Pro, VideoLoop reaches 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).
Figures & tables
Figure 1: Existing Append-only working memory leads to semantic thrashing, our VideoLoop sustains stable retrieval via a dual-loop bounded working memory. (a) A single-loop video agent appends every tool result to a working memory bounded by the LVLM context. As context accumulates, attention to key evidence is diluted, and effective retrieval first peaks and then collapses as context continues to grow. The append-only agent achieves an 81.9% final accuracy on VideoMME ( long ). (b) VideoLoop adds an inner orchestration loop. Tool results are written to an unbounded filesystem, and a memory orchestrator retrieves from the filesystem to rewrite a bounded working memory after every outer step. Our method improves task accuracy by +3.9 points on VideoMME ( long ) to 85.8%. Both configurations use Gemini 3 Flash.
Figure 2: Overview of VideoLoop’s dual-loop architecture. Working memory starts empty. The outer loop reasons about the task and explores the video with transcription and visual analysis tools as needed. After each outer step, the inner loop reads the question and the new observation, retrieves relevant evidence from the external store, and rewrites a compact working memory.
Method
VideoMME (long)
VideoMMMU
LongVideoBench (long)
Per.
Comp.
Adapt.
Overall
Video agentic systems
VideoAgent ( Wang et al., 2024 )
49.0
–
–
–
–
–
LLoVi ( Zhang et al., 2024a )
50.2
–
–
–
–
–
VideoTree ( Wang et al., 2025b )
54.2
–
–
–
–
–
QuoTA ( Luo et al., 2026 )
55.7
–
–
–
–
–
Table 1: Benchmark comparison (accuracy, %). Bold denotes column best; parentheses give overall gains over the preceding native backbone in percentage points. VideoMME and LongVideoBench use their long subsets with the official subtitles. VideoMMMU uses all 900 questions (300 per track): Perception (Per.), Comprehension (Comp.), Adaptation (Adapt.), and Overall. Results marked with * are reproduced by us under the same settings. Dashes denote unavailable results.
Figure 4
Figure 4: Illustration of semantic thrashing on VideoMME ( long ). Four panels present fixed task difficulty quartiles Q1–Q4 from left to right (easiest to hardest), with the same grouping used across memory designs. The curves show blind-judge answer accuracy on frozen context snapshots, used as a proxy for evidence retrievability. Diamonds and labels at the right end give final-snapshot accuracy. Semantic thrashing occurs when retrievability early plateaus or degrades despite continued agentic iterations.
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing online methods either retain compact visual representations that lack semantic structure, or build higher-level memory stores organized around temporal proximity rather than explicit causal links, leaving multi-hop narrative reasoning to be reconstructed by the LLM at every query. We bridge this gap with \textsc{Homer}, a Hierarchical Online Memory Exploration and Reasoning framework. \textsc{Homer}'s memory mirrors the multi-scale structure of long videos, ranging from raw perception, to recurring entities, to events connected by explicit temporal and causal relations. Its agentic reasoner then explores this memory the way humans do, locating the relevant scene, looking up details, and composing the answer through multi-round memory retrieval, with a harness that verifies and corrects each step. \textsc{Homer} outperforms the previous best agent method by +5.5, +10.8, and +4.4 points on M3-Bench-robot, M3-Bench-web, and Video-MME-Long, and consistently lifts three various LLM backbones, indicating a model-agnostic structural capability for grounded retrieval over long videos.
Yixin Ji, Fanghua Ye, Juntao Li +5
1Soochow University · 2Tencent Hunyuan Multimodal Department
Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and this pipeline is costly at both ends: with few questions, building memory for the whole video costs far more than answering them; with many questions, the memory is never updated, so what is learned while answering questions is lost to the next question. To alleviate these, we introduce Sprout, an agentic framework that builds memory while reasoning: a temporal tree that sprouts detailed nodes as questions are answered. The agent watches the video segment by segment at a low frame rate, stopping when the current question can be answered, remembers each segment as a coarse node of the tree, and revisits key intervals at a higher frame rate to refine the tree with the recovered details. Once a segment is recorded as text, its video input is removed from the context history, while the original video remains reachable through the video tools. The memory tree and prior question--answer records persist across questions, so the memory is online and dynamic: built from the first question onward and updated by every question thereafter. We find that replacing accumulated video inputs with textual memory substantially reduces context usage while maintaining accuracy, with slight improvements in some settings. Across benchmarks on three models, Sprout achieves competitive or improved accuracy relative to representative offline memory methods, with no upfront construction stage and lower context cost per question.
Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps it searchable over long horizons, and grounds later reasoning in recoverable observations rather than compressed latent state alone. We propose Visual Agentic Memory (VAM), a training-free framework with three components. Online Indexing supports selective evidence retention under streaming constraints. Hierarchical Memory organises retained evidence in a Parallel Representation that aligns temporal context with spatial observations. Agentic Retrieval searches, inspects, and verifies candidate evidence before producing a grounded answer. On OVO-Bench, VAM achieves the highest RT+BT average (68.41) across all reported baselines, improving over end-to-end use of the same underlying MLLM (Gemini 3 Flash, 67.46). On the month-scale split of MM-Lifelong train@month (105.6 hours over 51 days), VAM reaches 17.11%, second only to ReMA with GPT-5 (17.62%). These results suggest that long-horizon video understanding benefits from treating visual memory as an explicit, inspectable, and queryable substrate. Code is available at https://github.com/yiliu-li/Visual-Agentic-Memory.