Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence jointly and progressively build a shared understanding that supports the complete question set. We introduce \textbf{VAMR} (\textbf{V}ideo \textbf{A}gent for \textbf{M}ulti-Question \textbf{R}easoning), which coordinates all questions about a video through one shared tool-use trajectory. At each round, a persistent policy model can invoke tools for one or more unresolved questions and submit answers for questions with sufficient evidence. Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions. After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds. Specifically, a question-level critic estimates the value of each active question, while round alignment maps each question advantage to the rounds that directly serve it before the aligned advantages are aggregated to optimize the shared actor. Across LVBench, Video-Holmes, and LongVideoBench, VAMR achieves the highest accuracy overall and the fewest reasoning rounds among iterative methods. On LVBench, it reaches 62.1% accuracy, exceeding VideoARM by \textbf{4.3} points while reducing reasoning rounds and processed frames by \textbf{85.9%} and \textbf{61.4%}.
Figures & tables
Figure 1: Motivation for multi-question video reasoning. Existing video agents use separate online reasoning loops for the questions associated with a video, even when video-level preprocessing is reusable. VAMR instead coordinates the question set in one shared loop, allowing each tool call to enrich a shared video story and the separate clues maintained for individual questions.
Figure 2: Overview of VAMR. One persistent policy model coordinates the active question window, submits resolved answers, and invokes ASR or question-conditioned visual tools. Tool observations update the shared story layer, question clue layer, and observation buffer, while completed slots are replenished from the backlog.
Method
LVBench
Video-Holmes
LongVideoBench
Acc. ↑
Rounds/Q ↓
Frames/Q ↓
Acc. ↑
Rounds/Q ↓
Frames/Q ↓
Acc. ↑
Rounds/Q ↓
Frames/Q ↓
Backbone and Loop Controls
Qwen3.5-9B (Direct)
28.8
N/A
32.0
43.8
N/A
32.0
58.9
N/A
32.0
Qwen3.7-Plus (Direct)
43.7
N/A
32.0
49.8
N/A
32.0
62.4
N/A
32.0
Qwen3.5-9B + Tools
45.2
4.8
210.7
46.1
3.4
126.7
60.3
4.5
161.0
Ind. Loops (SFT+PPO)
56.4
4.1
224.0
53.8
3.1
138.6
65.1
4.2
192.4
Table 1: Main results on three held-out video benchmarks. The best and second-best results are shown in bold and underlined, respectively.
Figure 3: Ablation studies on LVBench. Each point plots overall accuracy against reasoning rounds per question. Panel (a) isolates the shared loop, QA-inspector, and layered memory using the same Qwen3.5-9B model. Panel (b) compares the backbone, SFT, PPO, QHPO w/o round alignment, and QHPO. Lines connect configurations separated by one controlled architectural or training change. Moving toward the upper left indicates higher accuracy with fewer reasoning rounds.
Figure 4: Analysis of cross-question sharing on LVBench. Panel (a) reports accuracy and Frames/Q as more questions share one loop under controlled input information. Panel (b) reports visual-frame reduction and accuracy gain over independent loops across relative evidence-overlap quartiles. Mean gold evidence pair overlap is shown in parentheses.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Partition
Videos
Question Instances
SFT
610
5,775
QHPO
584
6,001
Validation
25
353
Total
1,219
12,129
Appendix
Table 2: Statistics of the CG-Bench training and validation partitions.
Dataset
Split
Videos
Questions
Questions/Video
Avg. Duration (min)
CG-Bench
Full
1,219
12,129
10.0
27.7
LVBench
Full
103
1,549
15.0
68.4
Video-Holmes
Test
270
1,837
6.8
2.8
LongVideoBench
Validation
753
1,337
1.8
7.9
Appendix
Table 3: Video and question statistics for the training and evaluation benchmarks. Average duration is measured over unique videos.
Configuration
LVBench
Video-Holmes
LongVideoBench
VideoARM
57.8±0.5
54.6±0.4
65.9±0.4
Ind. Loops (SFT+PPO)
56.4±0.4
53.8±0.3
65.1±0.3
SFT
57.9±0.3
54.8±0.2
65.7±0.3
PPO
58.7±0.5
55.6±0.4
65.9±0.5
QHPO w/o round alignment
60.0±0.4
56.4±0.4
66.1±0.3
VAMR / QHPO
62.1±0.3
58.2±0.3
66.5±0.4
Appendix
Table 4: Accuracy over three inference runs, reported as mean ± standard deviation.
Method
LVBench
Video-Holmes
LongVideoBench
DVD
6.3
5.3
6.5
VideoHV-Agent
4.7
4.3
4.4
VideoARM
2.4
2.3
2.6
VAMR
1.2
1.5
2.3
Appendix
Table 5: External model calls per question.
Method
Backbone / Native Config.
LVBench
Video-Holmes
LongVideoBench
LOVE-R1
Qwen2.5-VL-7B
48.2
N/A
60.1
LongVT
Qwen2.5-VL-7B
41.3
N/A
N/A
MR. Video
Gemini-2.0-Flash + GPT-4o
60.8
N/A
61.6 †
AVA
Qwen2.5-VL-7B + Qwen2.5-32B + Gemini-1.5-Pro
62.3
N/A
N/A
LongVideo-R1
Qwen3-8B + Qwen2.5-VL-72B
50.0
N/A
N/A
EVA
Qwen2.5-VL-7B
43.3
37.2
55.0
Appendix
Table 6: Accuracy under native model configurations. Prior-method values are published results, whereas VAMR is evaluated with the model stack shown.
Configuration
Acc.
Rounds/Q
Frames/Q
Independent loops
45.2
4.8
210.7
Shared loop
49.6
1.6
98.7
Shared loop + QA-inspector
53.9
1.5
82.5
Shared loop + layered memory
52.3
1.4
88.9
Full VAMR architecture
56.8
1.3
72.4
Appendix
Table 7: Numerical results for the architecture ablation on LVBench. Figure 3 (a) visualizes these results. All variants use the same Qwen3.5-9B model.
Training Objective
Acc.
Rounds/Q
Frames/Q
Backbone
56.8
1.3
72.4
SFT
57.9
0.8
76.1
PPO
58.7
1.0
82.6
QHPO w/o round alignment
60.0
1.0
93.7
QHPO
62.1
0.9
84.9
Appendix
Table 8: Numerical results for the training ablation on LVBench. Figure 3 (b) visualizes these results. All variants use the VAMR architecture.
Questions/loop
Loops
Accuracy
Rounds/Q
Frames/Q
1
768
58.1
4.4
219.2
2
384
59.4
2.5
142.6
4
192
60.2
1.6
110.3
8
96
61.7
1.1
92.8
16
48
62.8
0.9
87.7
Appendix
Table 9: Numerical results for the question-density scaling analysis in Figure 4 (a). Every setting evaluates the same 768 questions and exposes the same 16 question-option pairs per video to the controller.
Quartile
Videos
Pair overlap (%)
Ind. Acc.
VAMR Acc.
Δ Acc.
Ind. Frames/Q
VAMR Frames/Q
Reduction (%)
Q1 (lowest)
26
1.4
54.0
61.0
+7.0
228.4
83.5
63.4
Q2
26
12.4
52.4
56.1
+3.7
226.0
91.0
59.7
Q3
26
23.3
60.5
67.0
+6.5
221.7
85.5
61.4
Q4 (highest)
25
38.1
59.2
64.9
+5.7
219.6
78.5
64.3
Overall
103
18.6
56.4
62.1
+5.7
224.0
84.9
62.1
Appendix
Table 10: Numerical results for the evidence-overlap analysis in Figure 4 (b). Videos are ranked by gold evidence pair overlap and divided into relative quartiles from lowest (Q1) to highest (Q4). Ind. denotes the independent-loop policy trained through SFT followed by standard PPO, while VAMR is trained through SFT followed by QHPO.