Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and this pipeline is costly at both ends: with few questions, building memory for the whole video costs far more than answering them; with many questions, the memory is never updated, so what is learned while answering questions is lost to the next question. To alleviate these, we introduce Sprout, an agentic framework that builds memory while reasoning: a temporal tree that sprouts detailed nodes as questions are answered. The agent watches the video segment by segment at a low frame rate, stopping when the current question can be answered, remembers each segment as a coarse node of the tree, and revisits key intervals at a higher frame rate to refine the tree with the recovered details. Once a segment is recorded as text, its video input is removed from the context history, while the original video remains reachable through the video tools. The memory tree and prior question--answer records persist across questions, so the memory is online and dynamic: built from the first question onward and updated by every question thereafter. We find that replacing accumulated video inputs with textual memory substantially reduces context usage while maintaining accuracy, with slight improvements in some settings. Across benchmarks on three models, Sprout achieves competitive or improved accuracy relative to representative offline memory methods, with no upfront construction stage and lower context cost per question.
Figures & tables
Figure 1: (a) Offline memory is first built for the whole video before the first question; Q1–Q6 then reason over the same memory without updating it. Nodes used by a question take that question’s color, and the rest are never touched. (b) Per-question tokens with Gemini 3.8 Flash as the model. Direct feeds the video and question directly to the model. Offline memory is Qwen-MM-Plugins: its construction cost, amortized over the benchmark’s questions. Our method has no construction stage. (c) Sprout builds memory while reasoning. The agent watches the video segment by segment at a low frame rate and stops once the question has enough evidence, so the tail of the video is never watched and costs nothing. Each watched segment is remembered as a coarse textual node written into memory; revisits at a higher frame rate add child nodes beneath the intervals a question needs.
Figure 2: Pipeline of Sprout. Three questions about one video are answered in order, building one memory. Each row shows the rounds of one question: tool calls (circles) and their results (boxes). Video frames stay in the context for one round only; text results are kept. Q1 watches the first two chunks, remembers them as coarse nodes, and revisits n2 at a higher frame rate to add finer child nodes. Q2 reuses the memory: search returns a node and Q1’s answer, and a short revisit adds one more detail, without watching new video. Q3 finds that the memory does not cover what it asks, so it watches two more chunks. The last two chunks are never watched. The memory tree (bottom) is drawn on the same time axis as the video; node colors show which question added each node.
Method
Model
LVBench
Video-MME(L)
Video-Holmes
Direct
GPT-5
60.4
74.3
44.1
Direct
Qwen3.8-Max
81.8
80.4
67.3
Direct
Gemini 3.8 Flash
87.1
90.1
72.0
WorldMM ( Yeo et al., 2026 )
GPT-5
61.9
76.6
–
MERIT ( Choi et al., 2026 )
GPT-5
71.8
77.7
–
VideoSeek ( Lin et al., 2026a )
GPT-5
68.4
81.2
47.3
Table 1: Accuracy (%) on three video benchmarks. Rows are grouped into direct inference, agentic and memory methods, and Sprout; Video-MME(L) is the Long subset of Video-MME.
Method
Model
Ent.
Event
Habit
Rel.
Task
Overall
Direct
Gemini 2.5 Pro
43.2
40.5
41.0
55.2
52.4
46.4
Ego-R1 ( Tian et al., 2025 )
3B
51.2
53.2
63.9
50.4
50.8
53.0
M3-Agent ( Long et al., 2026a )
7B
44.4
54.8
62.3
56.8
54.0
53.5
WorldMM ( Yeo et al., 2026 )
Qwen3-VL-8B
49.6
56.4
63.9
58.4
58.7
56.4
MERIT ( Choi et al., 2026 )
Gemini 2.5 Pro
60.8
61.1
65.6
65.6
73.0
64.2
Qwen-MM-Plugins
Qwen3.8-Max
72.0
77.0
80.3
72.8
82.5
75.8
Table 2: Accuracy (%) on EgoLifeQA. Ent., Event, Habit, Rel., and Task denote the EntityLog, EventRecall, HabitInsight, RelationMap, and TaskMaster categories; Overall is the accuracy over all questions, not the mean of the categories.
Figure 3: Effect of releasing frames with Gemini 3.8 Flash. (a) The context of one LVBench question in each round, followed by the model output of that round; the bar chart gives the measured context size of each round. Without releasing frames, every frame stays, so the context only grows and its last round reaches 47.8K tokens. Sprout releases frames once the model has written them as text, so its context rises and falls: the rounds that watch a chunk (r1, r3) have the most tokens, and even the peak (24.3K) is about half of the variant’s last round. (b) Context window and accuracy of the two variants. (c) Context window of Direct, Qwen-MM-Plugins, and Sprout. The context window is the peak context size over all rounds of a question (the single request for Direct), in thousands of tokens (K); we report its mean and maximum over all questions.
Table 6
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Action
Tool (arguments)
Role
Watch
read_next ()
Attach the next unwatched chunk at the low frame rate fc .
Remember
write_chunk_memory (segments)
Store the pending chunk as contiguous coarse memory nodes.
Revisit
search_memory (pattern, window)
Retrieve facts and prior questions by literal alternatives and/or a time window.
view_node (node_id, focus, fps)
Re-attach the whole interval of an existing node at the chosen frame rate fr .
view_time (start, end, focus, fps)
Attach a short interval inside the watched video at the chosen frame rate fr .
refine_node (node_id, segments)
Subdivide the leaf just viewed into contiguous children.
Appendix
Table 5: Tool interface of Sprout. Frames attached by a tool stay in the context for the current round only; each write-back tool accepts only the chunk or interval watched in the preceding round.
Search
LVBench
Video-MME(L)
Video-Holmes
Embedding retrieval
86.1
85.6
72.1
Literal matching (ours)
86.9
86.0
71.6
Appendix
Table 6: Literal matching versus embedding retrieval in search_memory , with Qwen3.8-Max. Accuracy (%); Video-MME(L) is the Long subset of Video-MME. The embedding variant uses the same text embedding model as Qwen-MM-Plugins ( qwen3-vl-embedding ); everything else is unchanged.
Figure 4: The memory tree that Sprout builds on one LVBench video (Gemini 3.8 Flash, the video of Figures 5 and 6 ), after all 17 of its questions, all answered correctly. Each question has its own color. Q1 watches the first 10-minute chunk and records it as n1–n4; Q2–Q13 watch nothing further, Q14 watches the next two chunks, and Q17 the rest. Each child below a coarse node is written by a revisit and carries the color of the question that wrote it; Q2, Q3, Q4, Q6, Q7, and Q9 answer without writing anything. The child of n2 written by Q5 is the one shown in Figure 6 .
Figure 5: The first three questions on one LVBench video (Sprout with Gemini 3.8 Flash; Figure 6 continues with questions 4 and 5). Each round shows the tool call ( ▹ ) and, in the shaded block, its result. Q1 watches the first chunk and records it as four nodes; no later question in this figure watches further. Q2 is answered by one search. Q3 finds where the start list is shown but not its length, so it revisits that interval at 1 fps, first clamped at the boundary of node_1, and counts the athletes on the frames. Tool calls keep only their key arguments and tool results are abridged; memory text is verbatim from the request log, and token counts are total input tokens per question.
Figure 6: Questions 4 and 5 on the video of Figure 5 . Q4 finds that the bib number is not in memory, by pattern and then by time window, and revisits the whole node_2 at 1 fps to read it. Q5 revisits the start of the race, finds the gun at 03:56 rather than the 03:50 recorded at 0.1 fps, and writes the corrected time back as the child node_2_1.
Figure 7: Question 14 on the video of Figure 5 . Q2–Q13 are all answered within the first chunk. Q14 asks about a moment after 10:00, so it watches and records the next two chunks, then revisits the start of Heat 4 at 1 fps to tell 20:28 from 20:38 and writes what it sees as node_9_1.
Figure 8: Questions 15 and 16 on the video of Figure 5 . Q15 finds a winning time of 10.08s recorded at 0.1 fps, which is not among the options; the revisit reads 10.09s from the result graphic and writes it back as node_9_2. Q16 finds the winner already in memory and revisits the finish to confirm it before answering.
Figure 9: Question 17 on the video of Figure 5 . Q17 asks about Heat 6, whose race lies past the watched prefix, so it watches and records the last chunk; the new memory already names Su Bingtian third, and four revisits of the result graphics confirm it without writing anything.