On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher and the initial student state, implicitly assuming that selected examples retain positive supervision value throughout optimization. We show that supervision trustworthiness and supervision necessity are distinct yet coupled: the former concerns target credibility, while the latter varies with the student's current task competence; together, they shape supervision value. Building on this coupled view, we introduce Student-Curriculum Coupling (SCC), a closed-loop framework in which a compact Anchor-Frontier curriculum defines the candidate supervision space and the evolving student dynamically determines its active subset. Supervision can therefore be activated, suspended, or reactivated as competence changes, concentrating teacher computation and optimization on current task-level deficits. Across three TVG benchmarks, SCC achieves a 5.1% relative improvement in mean recall over Video-OPD on its original curriculum, while using 60.0% fewer training examples and reducing training time by 50.4%. Ablations support the complementary roles of capability-structured curriculum design and student-dependent supervision in achieving these gains. Together, these results establish SCC as a data- and compute-efficient framework for TVG post-training, delivering stronger temporal grounding by aligning trustworthy supervision with the student's evolving learning needs.
Figures & tables
Figure 1: The persistent-value assumption. All examples in each scheduled batch (shaded) receive OPD supervision as the student evolves, including those that already satisfy the task criterion.
Figure 2: Student–Curriculum Coupling. Shaded regions denote successive scheduled batches within the AF space. Current-student assessment determines active and suspended supervision, supporting capability acquisition and Anchor reactivation. Selective OPD updates the student, whose new state guides supervision in subsequent batches, closing the loop.
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
Models for Evaluation
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
Proprietary Models
GPT-4o ( Hurst et al., 2024 )
60.6
44.5
23.5
55.2
41.4
25.8
69.0
54.8
38.5
GPT-5 ( Singh et al., 2025 )
59.3
42.0
22.0
57.4
44.9
30.4
72.4
60.4
46.4
Gemini-2.0-Flash ( Comanici et al., 2025 )
66.4
53.5
27.1
62.9
54.0
37.7
76.2
66.4
48.3
Gemini-2.5-Flash ( Comanici et al., 2025 )
68.7
56.1
30.6
66.8
57.5
41.3
78.2
69.4
55.0
Table 1: Evaluation on three TVG benchmarks. Published benchmark results are taken from Video-OPD ( Li et al., 2026a ) . Bold values indicate the best performance among the non-proprietary methods shown in the table.
Figure 3: Accuracy–efficiency comparison. QVHighlights mIoU versus training time (left), QVHighlights mIoU across training steps (middle), and cumulative training time (right). Diamonds mark training completion. All times are measured on eight NVIDIA A100 GPUs.
Method
Curriculum
Data and Schedule
Supervision Routing
# Examples
Presentations [-0.2ex] / Rollouts
Teacher / OPD [-0.2ex] Routes
Suspended [-0.2ex] Routes
Route Rate
Video-OPD
TVDF
2,500
2,504
2,504
0
100.0%
Video-OPD
AF
1,000
1,004
1,004
0
100.0%
SCC
AF
1,000
1,004
492
512
49.0%
Table 2: Training data and supervision routing. All runs use 79 scheduled training steps. Presentations / Rollouts counts student rollouts; Teacher / OPD Routes counts rollouts receiving teacher scoring and OPD optimization; Suspended Routes counts rollouts not routed to teacher scoring or OPD after current-student assessment.
Candidate
Student-Dep.
Charades ⋆
ActivityNet ⋆
QVHighlights ⋆
Space
Realization
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
Base Model
–
42.9
61.7
41.5
23.1
30.4
41.2
30.7
20.0
36.9
46.6
38.2
29.5
(a) Candidate-Space Composition
TVDF
✓
52.1
72.7
46.2
32.9
49.3
62.1
47.4
37.4
61.2
73.6
60.2
50.9
Random
✓
51.6
73.0
45.9
31.7
48.8
62.3
46.4
36.4
60.9
73.8
59.8
50.0
Frontier-only
✓
52.0
73.1
46.3
32.5
49.5
63.1
47.4
37.3
62.3
75.3
62.0
51.9
Table 3: Decomposition of Student–Curriculum Coupling. Base Model denotes Qwen3-VL-8B-Instruct before post-training, with scores taken from Video-OPD ( Li et al., 2026a ) . All other results are obtained using the same local evaluation pipeline. The upper block varies the candidate space with student-dependent realization fixed; the lower block fixes AF and compares student-dependent realization with uniform OPD supervision. TVDF uses the original 2,500-example curriculum from Video-OPD ( Li et al., 2026a ) , while Random, Frontier-only, and AF each use 1,000 examples. ✓ denotes student-dependent realization, and × denotes uniform OPD supervision over all scheduled examples. ⋆ marks benchmarks re-annotated with TimeLens. Values are percentages rounded to one decimal place; bold marks the best result for each metric within each block.
Figure 4: Training dynamics of Student–Curriculum Coupling. Current student competence for Anchors, Frontiers, and the full AF curriculum (left); realized OPD updates, suspensions, and Anchor reactivations (middle); and effective-batch OPD loss with cumulative Frontier acquisition (right). The shaded interval marks a representative drop in Anchor competence and the resulting reactivation of supervision.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Bteff←{(xi,yt,i):xi∈Bt,qS,t(xi;yt,i)<τS}
Appendix
Algorithm 1 Student–Curriculum Coupling (SCC)
Source
Anchor
Frontier
Unique
Presentations
CosMo-Cap
49
437
486
489
InternVid-VTime
26
235
261
262
QuerYD
11
106
117
117
DiDeMo
10
89
99
99
HiREST
4
33
37
37
Total
100
900
1,000
1,004
Appendix
Table 4: Source composition of AF. Unique counts refer to selected examples; presentations additionally include four Frontier repeats in the final scheduled step.
Training Steps
Anchor
Frontier
Total
1–20
53
481
534
21–40
31
276
307
41–59
12
104
116
60–79
4
43†
47
Total
100
904†
1,004
Appendix
Table 5: Distribution of AF across the training schedule. Counts denote presentations grouped by evaluation-checkpoint intervals. † includes four Frontier repeats in the final step.
Method
Training Time
Mean R@0.3
Mean R@0.5
Mean R@0.7
Video-OPD ( Li et al., 2026a ) (TVDF)
2:01:46
69.13
50.57
39.53
Video-OPD (AF)
1:11:20
70.69
52.53
40.70
SCC (AF)
1:00:26
71.83
53.72
41.84
Appendix
Table 6: Training time and endpoint recall. Video-OPD (TVDF) results are from Li et al. (2026a) ; other results are obtained locally. Training time on eight NVIDIA A100 GPUs is reported as hh:mm:ss. Recall (%) is averaged across three TimeLens benchmarks before rounding to two decimals. Bold marks the shortest time and highest recall at each threshold.
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
Method
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
TimeLens-8B
53.3
74.6
49.5
33.4
49.3
63.8
48.0
36.4
63.0
77.8
63.4
51.8
Video-OPD (Round 1)
52.0
73.1
45.8
32.4
47.3
60.5
45.6
35.8
61.0
73.8
60.3
50.4
SCC
52.8
74.0
48.1
33.0
50.6
64.8
49.2
38.2
63.7
76.8
63.9
54.3
Appendix
Table 7: Comparison with TimeLens-8B. Baseline results are taken from Video-OPD ( Li et al., 2026a ) ; values are percentages; bold indicates the best result in each column.
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
Method
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
OP-FKD ( Li et al., 2026a )
72.20
47.43
31.28
62.22
46.71
35.76
75.47
62.36
50.49
OP-RKD ( Li et al., 2026a )
69.61
47.01
29.91
61.47
46.67
35.73
74.95
62.04
51.14
GRPO ( Li et al., 2026a )
67.98
47.81
27.27
63.76
49.78
36.31
77.16
65.61
51.53
SCC
73.95
48.14
33.01
64.78
49.18
38.24
76.77
63.85
54.25
Appendix
Table 8: Alternative post-training objectives on AF. All methods use the AF curriculum and are trained and evaluated using our local pipeline. Off-policy forward- and reverse-KL distillation (OP-FKD and OP-RKD), together with GRPO, follow the objective definitions in Video-OPD ( Li et al., 2026a ) . All values are percentages, and bold indicates the best result in each column.
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
Method
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
Video-OPD ( Li et al., 2026a )
73.1
45.8
32.4
60.5
45.6
35.8
73.8
60.3
50.4
Video-OPD + Student-Dep. Realization
72.7
46.2
32.9
62.1
47.4
37.4
73.6
60.2
50.9
Appendix
Table 9: Student-dependent realization on TVDF. Both methods use the TVDF curriculum. Video-OPD results are taken from Li et al. (2026a) ; the student-dependent variant is trained and evaluated locally at checkpoint 79. Values are recall percentages rounded to one decimal place; bold indicates the higher value in each column.
τS
OPD Routes
Suspended Routes
OPD Route Rate
0.5
397
607
39.54%
0.6
438
566
43.63%
0.7
492
512
49.00%
0.8
563
441
56.08%
0.9
646
358
64.34%
Appendix
Table 10: Realized supervision routing under different student-competence criteria. All settings use the same 1,004 scheduled presentations. The shaded row denotes the main setting.
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
τS
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
-
52.26
73.33
47.13
32.80
49.61
63.40
47.84
37.20
62.49
75.34
62.62
52.11
0.5
52.51
73.71
47.81
32.68
50.52
64.20
48.82
38.22
63.17
76.83
62.23
52.56
0.6
52.41
73.68
47.07
32.47
50.30
63.84
48.36
38.18
63.30
76.18
62.69
53.41
0.7
52.79
73.95
48.14
33.01
50.63
64.78
49.18
38.24
63.70
76.77
63.85
54.25
0.8
52.85
74.13
47.79
33.51
50.11
63.44
48.20
38.07
62.85
75.86
62.88
52.56
Appendix
Table 11: TVG performance under different student-competence criteria. All benchmarks use the TimeLens annotations. All settings use AF; “–” denotes uniform OPD supervision, while the remaining rows vary τS . Values are percentages. The shaded row denotes the setting used in the main experiments, and bold indicates the best result in each column.
Teacher
Model
Benchmark
mIoU
R@0.3
R@0.5
R@0.7
Δ mIoU
4B GRPO
Direct teacher
QVHighlights
61.61
76.51
61.65
50.49
–
4B GRPO
SCC student
QVHighlights
64.92
77.87
65.41
55.09
+3.31
4B GRPO
Direct teacher
Charades-STA
51.35
73.15
47.40
30.90
–
4B GRPO
SCC student
Charades-STA
52.39
72.58
47.90
33.66
+1.04
4B GRPO
Direct teacher
ActivityNet
49.83
64.82
48.49
37.02
–
4B GRPO
SCC student
ActivityNet
51.45
65.29
49.91
39.20
+1.62
Appendix
Table 12: Robustness to teacher choice. GRPO teachers and their Qwen3-VL-8B students are evaluated at final checkpoints using TimeLens annotations. Student runs share the same AF configuration and differ only in teacher choice. mIoU and recall are percentages; Δ mIoU denotes the student’s gain over its teacher in percentage points. Bold marks the higher value for each metric within each teacher–student pair.
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
Checkpoint
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
20
48.24
68.66
45.97
28.84
49.39
63.93
49.07
37.31
62.80
77.35
64.50
51.59
40
52.75
74.04
47.78
32.92
50.93
64.67
49.47
38.64
63.46
76.77
63.27
53.28
59
52.73
73.60
47.81
33.24
50.86
65.16
49.31
38.64
63.39
76.31
63.40
53.28
79
52.79
73.95
48.14
33.01
50.63
64.78
49.18
38.24
63.70
76.77
63.85
54.25
Appendix
Table 13: Checkpoint-wise TVG performance of SCC. All checkpoints are taken from the main training run. All values are percentages, and bold indicates the best checkpoint for each metric within each benchmark.
Anchor
Frontier
Overall
A/F ratio
OPD Routes
Suspended
Rate
OPD Routes
Suspended
Rate
OPD Routes
Suspended
Rate
1:9
11
89
11.00%
481
423
53.21%
492
512
49.00%
3:7
41
260
13.62%
379
324
53.91%
420
584
41.83%
1:1
68
433
13.57%
271
232
53.88%
339
665
33.76%
Appendix
Table 14: Sensitivity to the Anchor–Frontier ratio. Each candidate space contains 1,000 unique examples. Ratios denote the numbers of unique Anchor and Frontier examples. Panel (a) reports routing over 1,004 scheduled presentations, including repeated occurrences in the training schedule. Panel (b) reports TVG performance at checkpoint 79. Values are percentages. The shaded rows denote the main 1:9 configuration, and bold values in panel (b) indicate the best result in each column.
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
Candidate-Space
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
1,000
52.79
73.95
48.14
33.01
50.63
64.78
49.18
38.24
63.70
76.77
63.85
54.25
1,500
52.69
74.19
47.58
32.65
50.81
64.40
49.07
38.56
62.40
75.41
61.71
51.79
2,000
52.63
74.07
47.61
32.83
50.07
63.82
48.13
37.91
62.70
75.67
62.43
53.08
2,500
52.00
73.51
47.04
31.91
50.38
64.11
48.82
38.31
62.76
76.38
62.82
52.43
Appendix
Table 15: Sensitivity to candidate-space scale. Candidate-space size counts unique examples. All configurations use an Anchor–Frontier ratio of 1:9 and are evaluated at checkpoint 79. Values are percentages rounded to two decimal places. The shaded row denotes the main configuration, and bold indicates the best result in each column.
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
Configuration
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
mIoU
R@0.3
R@0.5
R@0.7
Video-OPD on AF
52.26
73.33
47.13
32.80
49.61
63.40
47.84
37.20
62.49
75.34
62.62
52.11
Static routing
52.21
72.79
47.19
33.04
49.65
63.36
47.29
37.44
62.12
74.95
61.39
51.53
SCC (effective-batch)
52.82
74.31
48.02
32.74
50.55
64.53
48.78
38.13
62.88
75.54
62.75
53.54
SCC
52.79
73.95
48.14
33.01
50.63
64.78
49.18
38.24
63.70
76.77
63.85
54.25
Appendix
Table 16: Controls on student-dependent realization. All runs use AF and are evaluated at checkpoint 79. Static routing fixes 492 supervised presentations before training. SCC (effective-batch) averages loss over supervised presentations per batch. Values are percentages; bold marks column maxima and shading denotes SCC.
Figure 5: Sensitivity to rollout count. Benchmark mIoU during training with one, three, or five student rollouts per scheduled example. Total training times are 1:00:26, 1:22:24, and 1:33:03, respectively (hours:minutes:seconds).
Method
TempCompass
MVBench
Video-MME
Qwen3-VL-8B-Instruct ( Bai et al., 2025a )
73.16
68.17
67.96
OP-RKD ( Li et al., 2026a )
73.16
68.73
67.52
OP-FKD ( Li et al., 2026a )
73.29
68.15
67.33
GRPO ( Li et al., 2026a )
73.04
68.95
67.63
SCC
73.35
69.08
68.56
Appendix
Table 17: Evaluation on broader video-understanding benchmarks. Values are accuracy percentages, and bold indicates the best result in each column.
Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distillation (OPD) and propose a new training regime for MLLMs termed Counterfactual Attention Policy Distillation (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.
Shaobo Ju, Haiyang Yu, Xuecheng Wu +5
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China. · Fudan University. · MMLab, The Chinese University of Hong Kong. +1
On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods primarily focus on enhancing this teacher-side guidance (e.g., by enriching teacher inputs and refining teacher feedback), yet we find that limited student perception is another critical bottleneck in multimodal OPD. By providing oracle visual facts, the performance of OPD-trained students can still be substantially improved for both weak and strong teachers. To address this bottleneck, we propose S-OPD, a simple multimodal on-policy distillation framework that explicitly strengthens student perceptual learning through two objectives. Specifically, Teacher-calibrated Policy Contrast separates student policies under original and masked images with teacher-based token-level gating, strengthening the student's reliance on visual evidence during reasoning. Policy Agreement aligns student policies under original and noise-perturbed images, further improving perceptual robustness to visual noise. Notably, our method can be seamlessly plugged into existing OPD frameworks, requiring no additional data annotations, model parameters or inference operations. Extensive experiments on eight benchmarks across student scales and distillation paradigms demonstrate consistent performance improvements, with gains of up to 4.25 points on LogicVista. When combined with existing teacher-side supervision methods, our method can yield further gains. Code is available at https://github.com/Sirilaw/S-OPD.
Siyuan Liu, Kanghui Tian, Yue Duan +4
Nanjing University · Fudan University · Nanjing University of Posts and Telecommunications
Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level rewards and the lack of fine grained credit assignment over long, temporally grounded reasoning trajectories. While reinforcement learning with verifiable rewards (RLVR) provides reliable supervision, it fails to capture token level contributions, leading to inefficient learning. Conversely, existing self distillation methods offer dense supervision but lack structure and diagnostic specificity, and often interact unstably with reinforcement learning. In this work, we propose VISD, a structured self distillation framework that introduces diagnostically meaningful privileged information for video reasoning. VISD employs a video aware judge model to decompose reasoning quality into multiple dimensions, including answer correctness, logical consistency, and spatio-temporal grounding, and uses this structured feedback to guide a teacher policy for token level supervision. To stably integrate dense supervision with RL, we introduce a direction magnitude decoupling mechanism, where rollout level advantages computed from rewards determine update direction, while structured privileged signals modulate token level update magnitudes. This design enables semantically aligned and fine grained credit assignment, improving both reasoning faithfulness and training efficiency. Additionally, VISD incorporates curriculum scheduling and EMA based teacher stabilization to support robust optimization over long video sequences. Experiments on diverse benchmarks show that VISD consistently outperforms strong baselines, improving answer accuracy and spatio temporal grounding quality. Notably, VISD reaches these gains with nearly 2x faster convergence in optimization steps, highlighting the effectiveness of structured self supervision in improving both performance and sample efficiency for VideoLLMs.