Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence. We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference. During reinforcement post-training, paired dense and sparse views are optimized with grounding rewards and a reliability-gated reference reward, encouraging sparse view predictions to preserve task-relevant temporal evidence. Notably, SAVER is trained only on 1,250 randomly sampled temporal grounding examples, without using any video question answering annotations. Across three temporal grounding benchmarks and six video question-answering benchmarks, SAVER consistently improves performance across frame budgets. In particular, SAVER can match or surpass dense-frame Qwen3.5 baselines while using substantially fewer frames. These results show that temporal grounding can serve as an effective evidence-localization proxy for learning sparse video reasoning that transfers to broader video understanding tasks.
Figures & tables
Figure 2 : Performance–efficiency trade-off across frame rates. Lower frame rates reduce the number of sampled frames but enlarge the sparse–dense performance gap for Qwen3.5. SAVER mitigates this degradation; for example, for video-QA, SAVER -2B at 0.1 fps nearly matches dense Qwen3.5-2B at 2.0 fps.
Method
FPS
Avg. Frames
Charades-STA
ActivityNet
NExT-GQA
Avg. mIoU
Prior temporal grounding models: 7B model
Qwen2.5-VL-7B
2.0
110
52.9
26.9
20.2
33.3
TimeChat
2.0
110
31.2
30.4
17.4
26.3
Temporal-RLT
2.0
110
57.0
39.0
37.3
44.4
VideoAuto-R1-7B
2.0
110
60.0
47.6
36.7
48.1
Controlled frame-rate evaluation: 0.8B model
Table 1 : Temporal grounding results under controlled frame budgets. For each Qwen3.5 setting, Qwen3.5 and SAVER are shown side by side under the same fps and average frame budget. Numbers in parentheses indicate the absolute improvement over Qwen3.5 at the same fps.
Method
FPS
Avg. Frames
Video Perception
Video Reasoning
Avg. Acc.
VideoMME
MVBench
LongVideoBench
MMVU
VideoMMMU
MVP
Prior video reasoning models (reported settings)
Qwen2.5-VL-7B
2.0
139
66.0
67.1
60.9
66.2
54.7
36.5
58.6
Qwen3-VL-8B
2.0
139
72.5
69.4
67.6
69.9
61.0
40.5
63.5
VideoAuto-R1-7B
2.0
139
67.3
71.0
60.5
69.7
58.6
39.4
61.1
Controlled frame-rate evaluation: 0.8B model
Table 2 : Video-QA results under controlled frame budgets. SAVER consistently improves Qwen3.5 across model sizes and frame rates, matching or surpassing dense-frame performance with far fewer frames.
Figure 3 : Qualitative comparison under sparse observations. Dense frames provide finer temporal evidence, whereas sparse frames contain only a few sampled moments. Under sparse input, Qwen3.5 predicts an overly broad interval, while SAVER produces a tighter prediction closer to the ground truth. This shows that dense references help preserve task-relevant temporal evidence under sparse frame budgets.
Table 5
Figure 4 : Effect of sparse view reward weight.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Size
FPS
Avg. Frames
Temporal Grounding
Avg. mIoU
Charades-STA
Frames
ActivityNet
Frames
NExT-GQA
Frames
Controlled frame-rate evaluation: 0.8B model
Qwen3.5
0.8B
2.0
110
43.9
59
38.0
191
26.6
80
36.1
Qwen3.5
0.8B
1.0
63
39.1
29
34.4
121
24.0
39
32.5
Qwen3.5
0.8B
0.5
31
33.4
14
28.4
60
19.0
19
26.9
Qwen3.5
0.8B
0.2
12
24.4
5
23.1
24
12.5
8
20.0
Appendix
Table 5 : Temporal grounding results under controlled frame budgets. For each fps setting, we report grounding performance and the corresponding number of observed frames for three model scales (0.8B, 2B, and 4B). Avg. Frames denotes the average number of observed frames across all grounding benchmarks, and Avg. mIoU denotes the average grounding performance.
Model
Size
FPS
Avg. Frames
Video Perception
Video Reasoning
Avg. Acc.
VideoMME
Frames
MVBench
Frames
LongVideoBench
Frames
MMVU
Frames
VideoMMMU
Frames
MVP
Frames
Controlled frame-rate evaluation: 0.8B model
Qwen3.5
0.8B
2.0
139
55.0
224
49.2
35
54.2
200
50.4
99
32.7
253
25.0
25
44.4
Qwen3.5
0.8B
1.0
118
54.8
197
47.6
17
52.8
190
50.1
51
33.1
240
24.7
12
43.9
Qwen3.5
0.8B
0.5
97
54.0
171
46.4
9
52.1
170
50.2
25
35.3
200
22.2
6
43.4
Qwen3.5
0.8B
0.2
61
51.6
124
44.3
5
50.9
122
49.4
10
34.1
99
21.1
5
41.9
Appendix
Table 6 : Video-QA results under controlled frame budgets. For each fps setting, we report benchmark performance and the corresponding number of observed frames across video perception and video reasoning tasks for three model scales (0.8B, 2B, and 4B). Avg. Frames denotes the average number of observed frames across all benchmarks, and Avg. Acc. denotes the average performance.
Model
Size
Sparse View FPS
TG Avg. mIoU
Video-QA Avg. Acc.
SAVER
2B
0.1
29.6
49.7
SAVER
2B
0.5
30.0
50.8
SAVER
2B
1.0
28.4
48.8
Appendix
Table 7: Effect of sparse view training frame rate. We fix the dense view training frame rate to 2.0 fps and vary the sparse view frame rate used during post-training. Results are evaluated under 0.1-fps inference, where sparse-frame reasoning is most challenging.
Visual tokens ↓
Peak memory (GiB) ↓
E2E latency (s) ↓
GPU inference (s) ↓
Benchmark
2.0 fps
0.1 fps
2.0 fps
0.1 fps
2.0 fps
0.1 fps
2.0 fps
0.1 fps
Charades-STA
3,746
251
5.04
4.24
1.72
0.97
1.49
0.89
NExT-GQA
10,402
515
8.94
4.37
2.12
0.70
1.63
0.60
ActivityNet
8,039
376
6.65
4.30
2.41
0.73
1.62
0.65
MMVU
8,689
474
7.70
4.33
2.16
1.78
1.93
1.67
VideoMME
69,750
3,060
18.72
4.80
10.03
1.41
8.17
1.12
Appendix
Table 8 : Practical efficiency profiling of Qwen3.5-2B at 2.0 and 0.1 fps. Lower frame rates substantially reduce visual-token count, peak GPU memory, and inference latency on average across nine evaluation benchmarks.
Method
ActivityNet TG
LongVideoBench VQA
MMVU VQA
Avg.
Qwen3.5-2B (0.1 fps)
20.7
51.3
57.0
43.0
Qwen3.5-2B + FastVID
31.1
51.0
59.3
47.1
Qwen3.5-2B + WFS-SB
30.0
56.6
60.3
49.0
SAVER (0.1 fps)
37.2
56.0
59.7
51.0
SAVER + WFS-SB
42.1
62.3
60.7
55.0
Appendix
Table 9 : Comparison with inference-time frame/token selection methods. We evaluate Qwen3.5-2B, SAVER, FastVID, and WFS-SB on one temporal-grounding benchmark and two video-QA benchmarks. SAVER and the selection-based approaches target complementary aspects of efficient video reasoning, and combining SAVER with WFS-SB yields the strongest overall performance. The average is computed over the three reported benchmarks.
Method
Visual tokens ↓
Peak memory (GiB) ↓
GPU inference (s) ↓
E2E latency (s) ↓
SAVER
4,933
5.51
1.79
2.27
FastVID
5,292
13.39
2.22
4.42
WFS-SB
5,393
6.96
1.94
6.36
Appendix
Table 10 : Practical inference cost of SAVER and inference-time selection methods. Lower values indicate lower inference cost. SAVER directly operates on sparse inputs, while FastVID and WFS-SB introduce additional inference-time processing for pruning or frame selection.
λref
TG mIoU
QA Avg.
0.0
27.4
48.7
0.5
29.2
50.0
1.0
30.0
50.8
2.0
30.1
50.2
Appendix
Table 11 : Effect of the reference reward weight λref . Temporal-grounding results are averaged over three benchmarks and video-QA results over six benchmarks.
Model
Size
Training Samples
TG Avg. mIoU
Video-QA Avg. Acc.
SAVER
2B
0
18.4
45.9
SAVER
2B
625
25.7
47.6
SAVER
2B
1250
30.0
50.8
SAVER
2B
1875
29.5
49.7
SAVER
2B
2500
28.6
48.2
Appendix
Table 12: Effect of grounding post-training data scale. We vary the number of Time-R1 grounding training examples used during post-training. Results are evaluated under 0.1-fps inference.
Policy
γ
Temporal Grounding
Video-QA
Avg. Frames
Avg. mIoU
Avg. Frames
Avg. Acc.
Qwen3.5 (dense)
–
110
34.9
139
51.0
SAVER (sparse only)
–
7
30.0
38
50.8
Fallback
0.3
8
29.1
39
50.5
Fallback
0.5
11
31.0
43
50.9
Fallback
0.7
47
38.3
54
52.0
Appendix
Table 13 : Confidence-based sparse-to-dense fallback. Inference starts from the 0.1-fps sparse view, and dense 2.0-fps inference is triggered when the sparse prediction confidence is below γ . Increasing γ routes more examples to dense inference and therefore increases the average number of processed frames.
Figure 5 : Temporal grounding success case on Charades-STA. Given the query “a person opens a small cabinet door,” SAVER correctly identifies that the action starts at the beginning of the clip and ends around 4 seconds. The predicted interval closely matches the ground-truth segment, showing that sparse-frame reasoning can succeed when the sampled frames preserve the key temporal evidence.
Figure 6 : Temporal grounding failure case on Charades-STA. Given the query “person throws their blanket inside,” SAVER predicts an earlier interval than the ground truth. Although the reasoning identifies a plausible moment when the person moves near the doorway, the sparse observations do not provide enough evidence to precisely distinguish the target throwing action. This example highlights a common failure mode: sparse-frame inference can still localize an incorrect segment when the key action is ambiguous or insufficiently captured.
Figure 7 : Video-QA success case on VideoMMMU. Given a mechanics question about how the vertical displacement changes when the length increases from 2 inches to 3 inches, SAVER correctly follows the visual solution and derives the proportional change. The model identifies the relevant formula, applies the new length, and selects the correct answer, showing that sparse video reasoning can succeed when the sampled frames retain the necessary visual and textual evidence.
Figure 8 : Video-QA failure case on VideoMMMU. Given a chemistry question about the mass of oxygen corresponding to 4 kg of hydrogen atoms, SAVER applies the standard hydrogen-to-oxygen mass ratio in water (1:8) and selects C (32 kg), whereas the ground-truth answer is D (64 kg), which corresponds to a 1:16 mass relation. Although the arithmetic is internally consistent, the model relies on a default prior assumption instead of the relation presented in the video. This example shows that sparse video reasoning can still fail when the reasoning chain is built on an incorrect premise that is not grounded in the visual evidence.
We present LOVER, a \underline{L}ong to sh\underline{O}rt \underline{V}ideo \underline{E}vidence \underline{R}einforced model for grounded question answering (GQA). LOVER highlights three innovations over existing reinforcement-learning (RL) based video reasoning models: (1) \textbf{Long-to-short Video Evidence Curriculum Learning}, which organizes RL training according to evidence duration and progressively adapts the model from long-range grounding to short-term reasoning; (2) \textbf{GQA Rewards}, which underscore the benefit of IoP reward over IoU for evidence spotting rather than strict temporal span overlap; (3) \textbf{Adaptive Timestamp Rendering}, which adaptively renders timestamps onto video frames using background-aware position and color selection to enhance temporal observability. The three designs are model-agnostic and reciprocal. They effectively improve QA, grounding, and grounded QA performance over different backbones. Notably, LOVER built on Time-R1 achieves new state-of-the-art (SOTA) results among open-source models on popular GQA benchmarks: NExT-GQA and ReXTime. Comprehensive ablation studies further validate the effectiveness of our three innovative components.
VRR-QA evaluates whether video-language systems can infer spatial, temporal, viewpoint, depth, and visibility relations that are not always resolved by a single frame. We present an inference-only system built around adaptive test-time computation. The system first answers each question with a direct video-language model pass, then uses multiple lightweight views to find unstable questions. Only these difficult questions are routed to a high-budget dense evidence module that constructs timestamped frame observations, relation-specific probes, candidate verification, and conservative temporal aggregation. This design separates two problems that are often confused in video question answering: finding plausible alternative answers and deciding when a current answer should actually be changed. On the test split, the final system obtains 90.07 average accuracy and 87.81 macro average accuracy. The report focuses on the final test system and the implementation settings required to reproduce the adaptive dense verifier.
Yuyang Sun, Yongliang Wu, Xingyu Zhu +8
Southeast University · National University of Singapore · Independent Researcher +2
Reinforcement learning has advanced video reasoning in large multi-modal models, yet dominant pipelines either rely on on-policy self-exploration, which plateaus at the model's knowledge boundary, or hybrid replay that mixes policies and demands careful regularization. Dynamic context methods zoom into focused evidence but often require curated pretraining and two-stage tuning, and their context remains bounded by a small model's capability. In contrast, larger models excel at instruction following and multi-modal understanding, can supply richer context to smaller models, and rapidly zoom in on target regions via simple tools. Building on this capability, we introduce an observation-level intervention: a frozen, tool-integrated teacher identifies the missing spatiotemporal dependency and provides a minimal evidence patch (e.g., timestamps, regions etc.) from the original video while the question remains unchanged. The student answers again with the added context, and training updates with a chosen-rollout scheme integrated into Group Relative Policy Optimization (GRPO). We further propose a Robust Improvement Reward (RIR) that aligns optimization with two goals: outcome validity through correct answers and dependency alignment through rationales that reflect the cited evidence. Advantages are group-normalized across the batch, preserving on-policy exploration while directing it along causally meaningful directions with minimal changes to the training stack. Experiments on various related benchmarks show consistent accuracy gains and strong generalization. Code will be available at https://jethrojames.github.io/FFR/.
Haojian Huang, Chuanyu Qin, Yinchuan Li +1
The Hong Kong University of Science and Technology (Guangzhou) · Knowin AI · Institute of Information Engineering, Chinese Academy of Sciences