Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: increasing the number of frames makes autoregressive rollouts expensive, while too few frames may miss temporally localized events and fine-grained visual details. We present \textbf{Frame Differential On-Policy Self-Distillation (FD-OPSD)}, which transfers the useful evidence of dense frame observations to a sparse frame policy during RL training. FD-OPSD compares the policy's token level preferences for the same sampled response under sparse and dense views, and distills the resulting frame differential signal without an external teacher or dense autoregressive rollout. The method preserves sparse-frame rollouts and leaves inference unchanged. Across Qwen2.5-VL-7B and Qwen3-VL-4B on six video reasoning benchmarks, FD-OPSD yields higher overall average performance than the strongest corresponding GRPO, T-GRPO, or Video-KTR baselines across the 16, 32, and 64 frame evaluation settings. These results show that dense visual evidence can be transferred selectively during training through token level self-distillation while retaining sparse frame rollouts and unchanged inference.
Figures & tables
Figure 1: Overview of FD-OPSD for frame differential evidence transfer.
Method
Frames
General Video Understanding
Fine-Grained Video Reasoning
Overall
MVBench
TempCompass
MMVU
VideoMMMU
VideoMME
VSI-Bench
Proprietary MLLMs
GPT-4o OpenAI (2024)
–
64.6
73.8
75.4
61.2
71.9
34.0
63.5
GPT-5 OpenAI (2025)
–
74.1
83.3
82.6
84.6
86.7
55.0
77.7
Gemini-1.5-Pro Team et al. (2024)
–
60.5
67.1
71.2
53.4
75.0
45.4
62.1
Gemini-2.5-Pro Comanici et al. (2025)
–
70.6
84.3
78.4
83.6
84.3
53.5
75.8
Table 1: Main results. Overall is averaged only when all six benchmark results are available.
Component
Variant
General Video Understanding
Fine-Grained Video Reasoning
Overall
MVBench
TempCompass
MMVU
VideoMMMU
VideoMME
VSI-Bench
CASD
FA positive only
64.1
71.2
64.5
50.2
57.1
34.2
56.8
+ positive utility gate
64.6
71.2
64.0
49.2
58.2
35.4
57.1
+ answer correctness gate
64.7
72.5
65.6
49.7
57.1
36.4
57.7
All token distillation
64.6
72.3
64.5
49.7
57.9
35.4
57.4
Fidelity coefficient
λ=0.005
64.8
72.3
66.4
49.0
56.8
33.4
57.1
Table 2: Ablation studies of FD-OPSD . The dense view uniformly samples frames from the video, while the local-dense view adds nearby frames around them.
Figure 2: Training cost comparison between FD-OPSD and different video-RL methods. FD-OPSD uses dense frames only for scoring the same on-policy responses, avoiding additional dense autoregressive rollouts.
Figure 3: Analysis of the frame differential signal throughout training. (a) Category wise prompt level mean absolute raw FA. (b) Video semantic to stopword enrichment across training.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Qwen2.5-VL-7B-Instruct
Qwen3-VL-4B-Instruct
Sparse student frame budget
8
8
Dense teacher frame budget
32
32
Global batch size
16
16
Responses per prompt
8
8
Maximum response length
2,048
2,048
Temperature
1.0
1.0
Appendix
Table 3: Training and evaluation setup for the two model families.
Method
Frames
General Video Understanding
Fine-Grained Video Reasoning
Overall
MVBench
TempCompass
MMVU
VideoMMMU
VideoMME
VSI-Bench
Open Source MLLMs
Qwen2.5-VL-7B-165K-SFT
16
60.9
69.0
60.2
48.4
53.1
30.6
53.7
32
61.6
69.7
62.2
51.3
55.4
32.8
55.5
64
61.7
70.0
61.9
51.1
58.9
31.0
55.8
Video-R1-7B
16
64.2
73.2
64.6
50.2
57.8
31.0
56.8
Appendix
Table 4: Detailed evaluation results beyond the main comparison.
Figure 4: Representative examples of CASD support, correction, and rejection under sparse and dense video views.
Sampling Strategy
Accuracy (%)
Δ vs. Dense-16
Uniform-4
42.8
−5.8
Random-4 (1)
44.2
−4.4
Random-4 (2)
44.8
−3.8
Random-8
46.2
−2.4
Dense-16
48.6
–
Sparse correct / Dense-16 wrong: about 8.0% of questions.
Appendix
Table 5: Effect of frame sampling on answer accuracy.
Figure 5: Video semantic concentration of the frame differential signal. Each panel compares the prompt level mean absolute FA of video semantic tokens and stopwords, and reports video semantic enrichment among high FA positions.
Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level rewards and the lack of fine grained credit assignment over long, temporally grounded reasoning trajectories. While reinforcement learning with verifiable rewards (RLVR) provides reliable supervision, it fails to capture token level contributions, leading to inefficient learning. Conversely, existing self distillation methods offer dense supervision but lack structure and diagnostic specificity, and often interact unstably with reinforcement learning. In this work, we propose VISD, a structured self distillation framework that introduces diagnostically meaningful privileged information for video reasoning. VISD employs a video aware judge model to decompose reasoning quality into multiple dimensions, including answer correctness, logical consistency, and spatio-temporal grounding, and uses this structured feedback to guide a teacher policy for token level supervision. To stably integrate dense supervision with RL, we introduce a direction magnitude decoupling mechanism, where rollout level advantages computed from rewards determine update direction, while structured privileged signals modulate token level update magnitudes. This design enables semantically aligned and fine grained credit assignment, improving both reasoning faithfulness and training efficiency. Additionally, VISD incorporates curriculum scheduling and EMA based teacher stabilization to support robust optimization over long video sequences. Experiments on diverse benchmarks show that VISD consistently outperforms strong baselines, improving answer accuracy and spatio temporal grounding quality. Notably, VISD reaches these gains with nearly 2x faster convergence in optimization steps, highlighting the effectiveness of structured self supervision in improving both performance and sample efficiency for VideoLLMs.
Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of visual information. This dilemma motivates the transfer of the reasoning capabilities of large models to smaller, more efficient ones. On-Policy Distillation (OPD) offers a promising solution by matching output-token distributions along student-generated trajectories. However, video reasoning often depends on evidence accumulated across multiple frames. In this context, output-level supervision only captures information expressed through token predictions and does not directly constrain the latent representations formed during reasoning. To address this limitation, we propose Latent-OPD, which augments OPD with trajectory-level latent distillation. Specifically, our method focuses on the position at the end of each trajectory, where hidden states effectively summarize the accumulated visual evidence and reasoning context. Furthermore, we introduce a progressive teacher-lookahead strategy, which aligns middle-to-late student layers with increasingly deeper teacher layers. Experiments on six video reasoning benchmarks show that Latent-OPD consistently outperforms output-only OPD. Notably, the improvements are particularly pronounced in scenarios with limited frames, long videos, or tasks requiring complex evidence aggregation. These results establish Latent-OPD as a highly effective approach to frame-efficient video reasoning.
Reinforcement learning has improved the reasoning ability of large language models, but applying outcome-only rewards to video multimodal large language models (Video-MLLMs) provides limited guidance on which visual evidence should support the answer. Inspired by multisensory integration, where consistent cues can enhance the salience and reliability of perceptual estimates, we introduce Consensus Frame GRPO (CF-GRPO), a temporal-annotation-free process-level reward framework for evidence-aware video reasoning. CF-GRPO constructs a consensus frame prior from intrinsic video cues, including temporal coverage, scene-transition cues, and query-conditioned visual relevance. It then computes a model-side frame-use score from visual and response representations and optimizes their agreement through the Consensus Frame Reward (CFR). With salience-aware sparse aggregation and distribution sharpening, CFR provides a high-contrast reward signal without requiring human temporal annotations. Experiments show that VideoCFR achieves competitive performance across complex video reasoning benchmarks and improves several metrics over representative Video-MLLM and RL baselines, while the consensus prior provides an interpretable view of the evidence frames emphasized during training. The implementation is available at https://github.com/1Pansy/VideoCFR.
Chengwen Liu, Zhe Huang, Jisheng Dang +3
School of Information Science and Engineering, Lanzhou University, Lanzhou, China · Beijing University of Posts and Telecommunications, Beijing, China · Cloud and AI BU, Huawei, Shenzhen, China +1