Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and this pipeline is costly at both ends: with few questions, building memory for the whole video costs far more than answering them; with many questions, the memory is never updated, so what is learned while answering questions is lost to the next question. To alleviate these, we introduce Sprout, an agentic framework that builds memory while reasoning: a temporal tree that sprouts detailed nodes as questions are answered. The agent watches the video segment by segment at a low frame rate, stopping when the current question can be answered, remembers each segment as a coarse node of the tree, and revisits key intervals at a higher frame rate to refine the tree with the recovered details. Once a segment is recorded as text, its video input is removed from the context history, while the original video remains reachable through the video tools. The memory tree and prior question--answer records persist across questions, so the memory is online and dynamic: built from the first question onward and updated by every question thereafter. We find that replacing accumulated video inputs with textual memory substantially reduces context usage while maintaining accuracy, with slight improvements in some settings. Across benchmarks on three models, Sprout achieves competitive or improved accuracy relative to representative offline memory methods, with no upfront construction stage and lower context cost per question.
Figures & tables
Figure 1: (a) Offline memory is first built for the whole video before the first question; Q1–Q6 then reason over the same memory without updating it. Nodes used by a question take that question’s color, and the rest are never touched. (b) Per-question tokens with Gemini 3.8 Flash as the model. Direct feeds the video and question directly to the model. Offline memory is Qwen-MM-Plugins: its construction cost, amortized over the benchmark’s questions. Our method has no construction stage. (c) Sprout builds memory while reasoning. The agent watches the video segment by segment at a low frame rate and stops once the question has enough evidence, so the tail of the video is never watched and costs nothing. Each watched segment is remembered as a coarse textual node written into memory; revisits at a higher frame rate add child nodes beneath the intervals a question needs.
Figure 2: Pipeline of Sprout. Three questions about one video are answered in order, building one memory. Each row shows the rounds of one question: tool calls (circles) and their results (boxes). Video frames stay in the context for one round only; text results are kept. Q1 watches the first two chunks, remembers them as coarse nodes, and revisits n2 at a higher frame rate to add finer child nodes. Q2 reuses the memory: search returns a node and Q1’s answer, and a short revisit adds one more detail, without watching new video. Q3 finds that the memory does not cover what it asks, so it watches two more chunks. The last two chunks are never watched. The memory tree (bottom) is drawn on the same time axis as the video; node colors show which question added each node.
Method
Model
LVBench
Video-MME(L)
Video-Holmes
Direct
GPT-5
60.4
74.3
44.1
Direct
Qwen3.8-Max
81.8
80.4
67.3
Direct
Gemini 3.8 Flash
87.1
90.1
72.0
WorldMM ( Yeo et al., 2026 )
GPT-5
61.9
76.6
–
MERIT ( Choi et al., 2026 )
GPT-5
71.8
77.7
–
VideoSeek ( Lin et al., 2026a )
GPT-5
68.4
81.2
47.3
Table 1: Accuracy (%) on three video benchmarks. Rows are grouped into direct inference, agentic and memory methods, and Sprout; Video-MME(L) is the Long subset of Video-MME.
Method
Model
Ent.
Event
Habit
Rel.
Task
Overall
Direct
Gemini 2.5 Pro
43.2
40.5
41.0
55.2
52.4
46.4
Ego-R1 ( Tian et al., 2025 )
3B
51.2
53.2
63.9
50.4
50.8
53.0
M3-Agent ( Long et al., 2026a )
7B
44.4
54.8
62.3
56.8
54.0
53.5
WorldMM ( Yeo et al., 2026 )
Qwen3-VL-8B
49.6
56.4
63.9
58.4
58.7
56.4
MERIT ( Choi et al., 2026 )
Gemini 2.5 Pro
60.8
61.1
65.6
65.6
73.0
64.2
Qwen-MM-Plugins
Qwen3.8-Max
72.0
77.0
80.3
72.8
82.5
75.8
Table 2: Accuracy (%) on EgoLifeQA. Ent., Event, Habit, Rel., and Task denote the EntityLog, EventRecall, HabitInsight, RelationMap, and TaskMaster categories; Overall is the accuracy over all questions, not the mean of the categories.
Figure 3: Effect of releasing frames with Gemini 3.8 Flash. (a) The context of one LVBench question in each round, followed by the model output of that round; the bar chart gives the measured context size of each round. Without releasing frames, every frame stays, so the context only grows and its last round reaches 47.8K tokens. Sprout releases frames once the model has written them as text, so its context rises and falls: the rounds that watch a chunk (r1, r3) have the most tokens, and even the peak (24.3K) is about half of the variant’s last round. (b) Context window and accuracy of the two variants. (c) Context window of Direct, Qwen-MM-Plugins, and Sprout. The context window is the peak context size over all rounds of a question (the single request for Direct), in thousands of tokens (K); we report its mean and maximum over all questions.
Table 6
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Action
Tool (arguments)
Role
Watch
read_next ()
Attach the next unwatched chunk at the low frame rate fc .
Remember
write_chunk_memory (segments)
Store the pending chunk as contiguous coarse memory nodes.
Revisit
search_memory (pattern, window)
Retrieve facts and prior questions by literal alternatives and/or a time window.
view_node (node_id, focus, fps)
Re-attach the whole interval of an existing node at the chosen frame rate fr .
view_time (start, end, focus, fps)
Attach a short interval inside the watched video at the chosen frame rate fr .
refine_node (node_id, segments)
Subdivide the leaf just viewed into contiguous children.
Appendix
Table 5: Tool interface of Sprout. Frames attached by a tool stay in the context for the current round only; each write-back tool accepts only the chunk or interval watched in the preceding round.
Search
LVBench
Video-MME(L)
Video-Holmes
Embedding retrieval
86.1
85.6
72.1
Literal matching (ours)
86.9
86.0
71.6
Appendix
Table 6: Literal matching versus embedding retrieval in search_memory , with Qwen3.8-Max. Accuracy (%); Video-MME(L) is the Long subset of Video-MME. The embedding variant uses the same text embedding model as Qwen-MM-Plugins ( qwen3-vl-embedding ); everything else is unchanged.
Figure 4: The memory tree that Sprout builds on one LVBench video (Gemini 3.8 Flash, the video of Figures 5 and 6 ), after all 17 of its questions, all answered correctly. Each question has its own color. Q1 watches the first 10-minute chunk and records it as n1–n4; Q2–Q13 watch nothing further, Q14 watches the next two chunks, and Q17 the rest. Each child below a coarse node is written by a revisit and carries the color of the question that wrote it; Q2, Q3, Q4, Q6, Q7, and Q9 answer without writing anything. The child of n2 written by Q5 is the one shown in Figure 6 .
Figure 5: The first three questions on one LVBench video (Sprout with Gemini 3.8 Flash; Figure 6 continues with questions 4 and 5). Each round shows the tool call ( ▹ ) and, in the shaded block, its result. Q1 watches the first chunk and records it as four nodes; no later question in this figure watches further. Q2 is answered by one search. Q3 finds where the start list is shown but not its length, so it revisits that interval at 1 fps, first clamped at the boundary of node_1, and counts the athletes on the frames. Tool calls keep only their key arguments and tool results are abridged; memory text is verbatim from the request log, and token counts are total input tokens per question.
Figure 6: Questions 4 and 5 on the video of Figure 5 . Q4 finds that the bib number is not in memory, by pattern and then by time window, and revisits the whole node_2 at 1 fps to read it. Q5 revisits the start of the race, finds the gun at 03:56 rather than the 03:50 recorded at 0.1 fps, and writes the corrected time back as the child node_2_1.
Figure 7: Question 14 on the video of Figure 5 . Q2–Q13 are all answered within the first chunk. Q14 asks about a moment after 10:00, so it watches and records the next two chunks, then revisits the start of Heat 4 at 1 fps to tell 20:28 from 20:38 and writes what it sees as node_9_1.
Figure 8: Questions 15 and 16 on the video of Figure 5 . Q15 finds a winning time of 10.08s recorded at 0.1 fps, which is not among the options; the revisit reads 10.09s from the result graphic and writes it back as node_9_2. Q16 finds the winner already in memory and revisits the finish to confirm it before answering.
Figure 9: Question 17 on the video of Figure 5 . Q17 asks about Heat 6, whose race lies past the watched prefix, so it watches and records the last chunk; the new memory already names Su Bingtian third, and four revisits of the result graphics confirm it without writing anything.
Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps it searchable over long horizons, and grounds later reasoning in recoverable observations rather than compressed latent state alone. We propose Visual Agentic Memory (VAM), a training-free framework with three components. Online Indexing supports selective evidence retention under streaming constraints. Hierarchical Memory organises retained evidence in a Parallel Representation that aligns temporal context with spatial observations. Agentic Retrieval searches, inspects, and verifies candidate evidence before producing a grounded answer. On OVO-Bench, VAM achieves the highest RT+BT average (68.41) across all reported baselines, improving over end-to-end use of the same underlying MLLM (Gemini 3 Flash, 67.46). On the month-scale split of MM-Lifelong train@month (105.6 hours over 51 days), VAM reaches 17.11%, second only to ReMA with GPT-5 (17.62%). These results suggest that long-horizon video understanding benefits from treating visual memory as an explicit, inspectable, and queryable substrate. Code is available at https://github.com/yiliu-li/Visual-Agentic-Memory.
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing online methods either retain compact visual representations that lack semantic structure, or build higher-level memory stores organized around temporal proximity rather than explicit causal links, leaving multi-hop narrative reasoning to be reconstructed by the LLM at every query. We bridge this gap with \textsc{Homer}, a Hierarchical Online Memory Exploration and Reasoning framework. \textsc{Homer}'s memory mirrors the multi-scale structure of long videos, ranging from raw perception, to recurring entities, to events connected by explicit temporal and causal relations. Its agentic reasoner then explores this memory the way humans do, locating the relevant scene, looking up details, and composing the answer through multi-round memory retrieval, with a harness that verifies and corrects each step. \textsc{Homer} outperforms the previous best agent method by +5.5, +10.8, and +4.4 points on M3-Bench-robot, M3-Bench-web, and Video-MME-Long, and consistently lifts three various LLM backbones, indicating a model-agnostic structural capability for grounded retrieval over long videos.
Yixin Ji, Fanghua Ye, Juntao Li +5
1Soochow University · 2Tencent Hunyuan Multimodal Department
Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multimodal agent with two coupled loops. The outer loop reasons over the video and the inner loop, after each step, retrieves artifacts from an unbounded filesystem of past observations and intermediate analysis, and rewrites a bounded working memory. Extensive experiments demonstrate the effectiveness of VideoLoop, which improves four popular LVLM backbones in a plug-and-play manner, with an average gain of 4.2% points over baseline on VideoMME (long). Further analysis of working memory suggests that VideoLoop mitigates semantic thrashing: on the hardest quarter of VideoMME (long) questions, a blind judge that reads only the agent's context answers 81.1% correctly, versus 60.9% for the append-only agent. With Gemini 3.1 Pro, VideoLoop reaches 88.3% on VideoMME (long), 88.8% on VideoMMMU, and 80.9% on LongVideoBench (long).