Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence jointly and progressively build a shared understanding that supports the complete question set. We introduce \textbf{VAMR} (\textbf{V}ideo \textbf{A}gent for \textbf{M}ulti-Question \textbf{R}easoning), which coordinates all questions about a video through one shared tool-use trajectory. At each round, a persistent policy model can invoke tools for one or more unresolved questions and submit answers for questions with sufficient evidence. Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions. After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds. Specifically, a question-level critic estimates the value of each active question, while round alignment maps each question advantage to the rounds that directly serve it before the aligned advantages are aggregated to optimize the shared actor. Across LVBench, Video-Holmes, and LongVideoBench, VAMR achieves the highest accuracy overall and the fewest reasoning rounds among iterative methods. On LVBench, it reaches 62.1% accuracy, exceeding VideoARM by \textbf{4.3} points while reducing reasoning rounds and processed frames by \textbf{85.9%} and \textbf{61.4%}.
Figures & tables
Figure 1: Motivation for multi-question video reasoning. Existing video agents use separate online reasoning loops for the questions associated with a video, even when video-level preprocessing is reusable. VAMR instead coordinates the question set in one shared loop, allowing each tool call to enrich a shared video story and the separate clues maintained for individual questions.
Figure 2: Overview of VAMR. One persistent policy model coordinates the active question window, submits resolved answers, and invokes ASR or question-conditioned visual tools. Tool observations update the shared story layer, question clue layer, and observation buffer, while completed slots are replenished from the backlog.
Method
LVBench
Video-Holmes
LongVideoBench
Acc. ↑
Rounds/Q ↓
Frames/Q ↓
Acc. ↑
Rounds/Q ↓
Frames/Q ↓
Acc. ↑
Rounds/Q ↓
Frames/Q ↓
Backbone and Loop Controls
Qwen3.5-9B (Direct)
28.8
N/A
32.0
43.8
N/A
32.0
58.9
N/A
32.0
Qwen3.7-Plus (Direct)
43.7
N/A
32.0
49.8
N/A
32.0
62.4
N/A
32.0
Qwen3.5-9B + Tools
45.2
4.8
210.7
46.1
3.4
126.7
60.3
4.5
161.0
Ind. Loops (SFT+PPO)
56.4
4.1
224.0
53.8
3.1
138.6
65.1
4.2
192.4
Table 1: Main results on three held-out video benchmarks. The best and second-best results are shown in bold and underlined, respectively.
Figure 3: Ablation studies on LVBench. Each point plots overall accuracy against reasoning rounds per question. Panel (a) isolates the shared loop, QA-inspector, and layered memory using the same Qwen3.5-9B model. Panel (b) compares the backbone, SFT, PPO, QHPO w/o round alignment, and QHPO. Lines connect configurations separated by one controlled architectural or training change. Moving toward the upper left indicates higher accuracy with fewer reasoning rounds.
Figure 4: Analysis of cross-question sharing on LVBench. Panel (a) reports accuracy and Frames/Q as more questions share one loop under controlled input information. Panel (b) reports visual-frame reduction and accuracy gain over independent loops across relative evidence-overlap quartiles. Mean gold evidence pair overlap is shown in parentheses.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Partition
Videos
Question Instances
SFT
610
5,775
QHPO
584
6,001
Validation
25
353
Total
1,219
12,129
Appendix
Table 2: Statistics of the CG-Bench training and validation partitions.
Dataset
Split
Videos
Questions
Questions/Video
Avg. Duration (min)
CG-Bench
Full
1,219
12,129
10.0
27.7
LVBench
Full
103
1,549
15.0
68.4
Video-Holmes
Test
270
1,837
6.8
2.8
LongVideoBench
Validation
753
1,337
1.8
7.9
Appendix
Table 3: Video and question statistics for the training and evaluation benchmarks. Average duration is measured over unique videos.
Configuration
LVBench
Video-Holmes
LongVideoBench
VideoARM
57.8±0.5
54.6±0.4
65.9±0.4
Ind. Loops (SFT+PPO)
56.4±0.4
53.8±0.3
65.1±0.3
SFT
57.9±0.3
54.8±0.2
65.7±0.3
PPO
58.7±0.5
55.6±0.4
65.9±0.5
QHPO w/o round alignment
60.0±0.4
56.4±0.4
66.1±0.3
VAMR / QHPO
62.1±0.3
58.2±0.3
66.5±0.4
Appendix
Table 4: Accuracy over three inference runs, reported as mean ± standard deviation.
Method
LVBench
Video-Holmes
LongVideoBench
DVD
6.3
5.3
6.5
VideoHV-Agent
4.7
4.3
4.4
VideoARM
2.4
2.3
2.6
VAMR
1.2
1.5
2.3
Appendix
Table 5: External model calls per question.
Method
Backbone / Native Config.
LVBench
Video-Holmes
LongVideoBench
LOVE-R1
Qwen2.5-VL-7B
48.2
N/A
60.1
LongVT
Qwen2.5-VL-7B
41.3
N/A
N/A
MR. Video
Gemini-2.0-Flash + GPT-4o
60.8
N/A
61.6 †
AVA
Qwen2.5-VL-7B + Qwen2.5-32B + Gemini-1.5-Pro
62.3
N/A
N/A
LongVideo-R1
Qwen3-8B + Qwen2.5-VL-72B
50.0
N/A
N/A
EVA
Qwen2.5-VL-7B
43.3
37.2
55.0
Appendix
Table 6: Accuracy under native model configurations. Prior-method values are published results, whereas VAMR is evaluated with the model stack shown.
Configuration
Acc.
Rounds/Q
Frames/Q
Independent loops
45.2
4.8
210.7
Shared loop
49.6
1.6
98.7
Shared loop + QA-inspector
53.9
1.5
82.5
Shared loop + layered memory
52.3
1.4
88.9
Full VAMR architecture
56.8
1.3
72.4
Appendix
Table 7: Numerical results for the architecture ablation on LVBench. Figure 3 (a) visualizes these results. All variants use the same Qwen3.5-9B model.
Training Objective
Acc.
Rounds/Q
Frames/Q
Backbone
56.8
1.3
72.4
SFT
57.9
0.8
76.1
PPO
58.7
1.0
82.6
QHPO w/o round alignment
60.0
1.0
93.7
QHPO
62.1
0.9
84.9
Appendix
Table 8: Numerical results for the training ablation on LVBench. Figure 3 (b) visualizes these results. All variants use the VAMR architecture.
Questions/loop
Loops
Accuracy
Rounds/Q
Frames/Q
1
768
58.1
4.4
219.2
2
384
59.4
2.5
142.6
4
192
60.2
1.6
110.3
8
96
61.7
1.1
92.8
16
48
62.8
0.9
87.7
Appendix
Table 9: Numerical results for the question-density scaling analysis in Figure 4 (a). Every setting evaluates the same 768 questions and exposes the same 16 question-option pairs per video to the controller.
Quartile
Videos
Pair overlap (%)
Ind. Acc.
VAMR Acc.
Δ Acc.
Ind. Frames/Q
VAMR Frames/Q
Reduction (%)
Q1 (lowest)
26
1.4
54.0
61.0
+7.0
228.4
83.5
63.4
Q2
26
12.4
52.4
56.1
+3.7
226.0
91.0
59.7
Q3
26
23.3
60.5
67.0
+6.5
221.7
85.5
61.4
Q4 (highest)
25
38.1
59.2
64.9
+5.7
219.6
78.5
64.3
Overall
103
18.6
56.4
62.1
+5.7
224.0
84.9
62.1
Appendix
Table 10: Numerical results for the evidence-overlap analysis in Figure 4 (b). Videos are ranked by gold evidence pair overlap and divided into relative quartiles from lowest (Q1) to highest (Q4). Ind. denotes the independent-loop policy trained through SFT followed by standard PPO, while VAMR is trained through SFT followed by QHPO.
Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and this pipeline is costly at both ends: with few questions, building memory for the whole video costs far more than answering them; with many questions, the memory is never updated, so what is learned while answering questions is lost to the next question. To alleviate these, we introduce Sprout, an agentic framework that builds memory while reasoning: a temporal tree that sprouts detailed nodes as questions are answered. The agent watches the video segment by segment at a low frame rate, stopping when the current question can be answered, remembers each segment as a coarse node of the tree, and revisits key intervals at a higher frame rate to refine the tree with the recovered details. Once a segment is recorded as text, its video input is removed from the context history, while the original video remains reachable through the video tools. The memory tree and prior question--answer records persist across questions, so the memory is online and dynamic: built from the first question onward and updated by every question thereafter. We find that replacing accumulated video inputs with textual memory substantially reduces context usage while maintaining accuracy, with slight improvements in some settings. Across benchmarks on three models, Sprout achieves competitive or improved accuracy relative to representative offline memory methods, with no upfront construction stage and lower context cost per question.
Long video understanding requires more than large context windows. It also needs a memory mechanism that decides what visual evidence to retain, keeps it searchable over long horizons, and grounds later reasoning in recoverable observations rather than compressed latent state alone. We propose Visual Agentic Memory (VAM), a training-free framework with three components. Online Indexing supports selective evidence retention under streaming constraints. Hierarchical Memory organises retained evidence in a Parallel Representation that aligns temporal context with spatial observations. Agentic Retrieval searches, inspects, and verifies candidate evidence before producing a grounded answer. On OVO-Bench, VAM achieves the highest RT+BT average (68.41) across all reported baselines, improving over end-to-end use of the same underlying MLLM (Gemini 3 Flash, 67.46). On the month-scale split of MM-Lifelong train@month (105.6 hours over 51 days), VAM reaches 17.11%, second only to ReMA with GPT-5 (17.62%). These results suggest that long-horizon video understanding benefits from treating visual memory as an explicit, inspectable, and queryable substrate. Code is available at https://github.com/yiliu-li/Visual-Agentic-Memory.
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing online methods either retain compact visual representations that lack semantic structure, or build higher-level memory stores organized around temporal proximity rather than explicit causal links, leaving multi-hop narrative reasoning to be reconstructed by the LLM at every query. We bridge this gap with \textsc{Homer}, a Hierarchical Online Memory Exploration and Reasoning framework. \textsc{Homer}'s memory mirrors the multi-scale structure of long videos, ranging from raw perception, to recurring entities, to events connected by explicit temporal and causal relations. Its agentic reasoner then explores this memory the way humans do, locating the relevant scene, looking up details, and composing the answer through multi-round memory retrieval, with a harness that verifies and corrects each step. \textsc{Homer} outperforms the previous best agent method by +5.5, +10.8, and +4.4 points on M3-Bench-robot, M3-Bench-web, and Video-MME-Long, and consistently lifts three various LLM backbones, indicating a model-agnostic structural capability for grounded retrieval over long videos.
Yixin Ji, Fanghua Ye, Juntao Li +5
1Soochow University · 2Tencent Hunyuan Multimodal Department