While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A further challenge lies in determining whether teacher guidance should refine reward-based updates or provide corrective supervision for failed trajectories. To address these issues, we propose MeSD, a multi-evidence self-distillation framework for VideoLLMs. MeSD constructs three evidence-conditioned teachers with shared parameters, using the ground-truth answer as a common semantic context while separately incorporating temporal and spatial evidence. Given the same student-generated prefixes, MeSD evaluates evidence-specific preferences relative to the Answer Teacher and fuses teacher-common preferences with gated teacher-specific residuals. Furthermore, MeSD introduces Verification-Guided Optimization to classify trajectories as Success, Failure, or Indeterminate. For Success and Indeterminate trajectories, MeSD refines token-level advantage magnitudes while preserving reward-derived signs. For verified failure trajectories that contain the required evidence, MeSD applies failure-conditioned distillation, using reverse-KL correction toward the fused distribution. Experiments on multiple video benchmarks demonstrate consistent gains over reinforcement learning and self-distillation baselines.
Figures & tables
Figure 1: (a) Temporal and spatial evidence yield consistent or conflicting token preferences. (b) MeSD combines reward-anchored credit refinement with corrective distillation on verified failures. (c) MeSD leads on selected metrics across six benchmarks. (d) Comparison of training rewards.
Figure 2: Overview of MeSD. Teachers sharing EMA parameters use answer, temporal, and spatial annotations as privileged information to evaluate student-generated trajectories. Using the Answer Teacher as an anchor, teacher distribution fusion aggregates evidence-specific preferences through a gated consensus–residual decomposition. Trajectory verification determines how the student learns. For failed trajectories, the student matches the fused teacher distribution through reverse-KL distillation. For other trajectories, teacher guidance adjusts the strength of token-level updates, and rewards determine whether tokens are reinforced or suppressed.
Figure 3: Training ablation curves. MeSD achieves higher training rewards than (a) credit refinement alone (w/o FCD) and (b) multi-evidence OPSD (w/o VGO). (c) Varying the teacher cutoff step H and γ yields similar reward trajectories.
Figure 4: Sensitivity of MeSD to γ and the teacher cutoff step H across five benchmarks. Filled bars indicate the default settings ( γ=0.8 , H=1000 ).
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Task type
Samples
Answer
Temporal
Spatial
General video QA (MCQ)
13,000
Option
–
–
General video QA (free-form)
2,000
Text
–
–
Temporal QA
2,280
–
Time intervals
–
Temporal QA (MCQ)
2,904
Option
Time intervals
–
Visual QA
5,000
Object name
–
Object boxes
Joint spatio-temporal QA
12,047
Text
Keyframe timestamps
Object boxes
Appendix
Table 5: Task composition and annotation types of STGR-RL. Answer denotes the semantic answer used in the answer context.
Training and rollout
Self-distillation
Parameter
Value
Parameter
Value
Training epochs
1
EMA update rate
0.01
Optimizer
AdamW
Initial strength λ0
0.5
Learning rate
10−6
Annealing steps
1,000
LR schedule
Cosine
Token-weight clipping
[0.8,1.2]
Prompt batch size
16
Correction strength γ
0.8
Appendix
Table 6: Training configuration of MeSD.
Figure 5: Answer teacher prompt with privileged semantic answers.
Figure 6: Answer+Temporal teacher prompt with semantic answers and keyframe timestamps.
Figure 7: Answer+Spatial teacher prompt with semantic answers and object-level bounding boxes.
Task
Success
Failure
RKL eligibility
Video MCQ
Correct option
Incorrect option
Incorrect option
Temporal localization
tIoU≥0.5
tIoU=0
tIoU=0
Temporal MCQ
Correct option and tIoU≥0.5
Incorrect option or tIoU=0
Incorrect option and tIoU=0
Spatial localization
Not assigned
Not assigned
Not eligible
Other composite or free-form tasks
Not assigned
Not assigned
Not eligible
Appendix
Table 7: Task-specific outcome verification and RKL eligibility under the strict setting.
Method
Steps
Time (h)
A800 GPU-hours
VISD no-feedback
2,327
61.36
981.70
MeSD
2,327
20.96
335.38
Appendix
Table 12: Training efficiency comparison at the same number of training steps.
Model
What
When (tIoU)
Where (sIoU)
Overall
Acc.
Chain1
Chain2
Chain1
Chain2
mAM
mLGM
Gemini-2-Flash [ Google DeepMind, 2025 ]
53.0
24.5
23.8
4.6
2.2
26.9
35.6
GPT-4o [ OpenAI, 2024 ]
60.8
16.7
12.8
6.5
3.0
26.8
38.2
Oryx-1.5-7B [ Liu et al., 2025 ]
20.5
13.5
14.8
10.1
3.5
15.1
13.8
InternVL-2.5-8B [ Chen et al., 2024 ]
44.2
8.7
7.8
0.7
0.1
17.6
24.9
Qwen2.5-VL-7B [ Bai et al., 2025b ]
33.5
15.4
13.8
17.0
2.5
19.3
22.4
Appendix
Table 13: Detailed V-STAR comparison.
Model
WorldSense
Overall
Recognition
Understanding
Reasoning
GPT-4o
42.6
–
–
–
Qwen2.5-VL-7B
35.5
31.8
38.4
38.7
Open-o3-Video-7B
38.9
36.8
39.7
40.5
VISD-7B
39.9
39.4
40.4
40.5
MeSD-7B (Ours)
41.6
38.9
46.0
42.3
Appendix
Table 14: Detailed comparison on WorldSense, VideoMMMU, LRR, and TVGBench.
Model
What
When (tIoU)
Where (sIoU)
Overall
Acc.
Chain1
Chain2
Chain1
Chain2
mAM
mLGM
Component Variants
w/o FCD
56.4
22.0
20.9
23.7
5.0
30.7
41.0
w/o ES
56.5
22.1
21.7
23.4
5.2
30.9
41.3
w/o T&S
56.4
22.3
21.4
23.5
5.1
30.8
41.2
w/o GPA
56.5
22.3
21.7
23.6
5.1
30.9
41.4
Appendix
Table 15: Detailed V-STAR ablations.
Model
WorldSense
Overall
Recognition
Understanding
Reasoning
Component Variants
w/o FCD
38.6
38.2
38.7
39.0
w/o ES
38.3
38.4
38.4
38.1
w/o T&S
37.9
37.3
39.3
37.3
w/o GPA
39.5
37.9
39.5
38.6
Appendix
Table 16: Detailed ablations on WorldSense, VideoMMMU, LRR, and TVGBench.
Figure 8: Training reward for Qwen3-VL models at 4B, 8B, and 32B.
Figure 9: Training reward for MeSD and the w/o ES, w/o T&S, and w/o GPA component variants.
Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level rewards and the lack of fine grained credit assignment over long, temporally grounded reasoning trajectories. While reinforcement learning with verifiable rewards (RLVR) provides reliable supervision, it fails to capture token level contributions, leading to inefficient learning. Conversely, existing self distillation methods offer dense supervision but lack structure and diagnostic specificity, and often interact unstably with reinforcement learning. In this work, we propose VISD, a structured self distillation framework that introduces diagnostically meaningful privileged information for video reasoning. VISD employs a video aware judge model to decompose reasoning quality into multiple dimensions, including answer correctness, logical consistency, and spatio-temporal grounding, and uses this structured feedback to guide a teacher policy for token level supervision. To stably integrate dense supervision with RL, we introduce a direction magnitude decoupling mechanism, where rollout level advantages computed from rewards determine update direction, while structured privileged signals modulate token level update magnitudes. This design enables semantically aligned and fine grained credit assignment, improving both reasoning faithfulness and training efficiency. Additionally, VISD incorporates curriculum scheduling and EMA based teacher stabilization to support robust optimization over long video sequences. Experiments on diverse benchmarks show that VISD consistently outperforms strong baselines, improving answer accuracy and spatio temporal grounding quality. Notably, VISD reaches these gains with nearly 2x faster convergence in optimization steps, highlighting the effectiveness of structured self supervision in improving both performance and sample efficiency for VideoLLMs.
Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, existing RL pipelines rely heavily on human-annotated tasks and solutions, making them costly to scale and fundamentally constrained by human expertise. Self-evolving frameworks have recently emerged as a promising alternative through autonomous Questioner-Solver self-play. Unfortunately, these approaches are primarily designed for static modalities such as text and images, fundamentally failing to capture the temporal dynamics that are central to video reasoning. In this work, we propose EvoVid, a temporal-centric self-evolving framework that enables Video-LLMs to improve directly from raw, unannotated videos. Specifically, we introduce two complementary temporal-centric rewards: a temporal-aware Questioner reward that encourages temporally dependent question generation through temporal perturbation sensitivity, and a temporal-grounded Solver reward that provides automatic temporal supervision via inherent video segment localization. Extensive experiments across four base models and six benchmarks demonstrate consistent improvements over both base models and existing self-evolving baselines, achieving competitive performance with supervised methods. These results highlight temporal-centric self-evolution as an effective and scalable paradigm for video understanding and reasoning.
Shiqi Huang, Ziyue Wang, Zhongrong Zuo +3
School of Electrical and Electronic Engineering, Nanyang Technological University · ByteDance
Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: increasing the number of frames makes autoregressive rollouts expensive, while too few frames may miss temporally localized events and fine-grained visual details. We present \textbf{Frame Differential On-Policy Self-Distillation (FD-OPSD)}, which transfers the useful evidence of dense frame observations to a sparse frame policy during RL training. FD-OPSD compares the policy's token level preferences for the same sampled response under sparse and dense views, and distills the resulting frame differential signal without an external teacher or dense autoregressive rollout. The method preserves sparse-frame rollouts and leaves inference unchanged. Across Qwen2.5-VL-7B and Qwen3-VL-4B on six video reasoning benchmarks, FD-OPSD yields higher overall average performance than the strongest corresponding GRPO, T-GRPO, or Video-KTR baselines across the 16, 32, and 64 frame evaluation settings. These results show that dense visual evidence can be transferred selectively during training through token level self-distillation while retaining sparse frame rollouts and unchanged inference.