cs.CVSep 30, 2026

Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding

Authors: Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djurić, Sima Mofakham

Organizations: State University of New York at Stony Brook

Abstract

On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher and the initial student state, implicitly assuming that selected examples retain positive supervision value throughout optimization. We show that supervision trustworthiness and supervision necessity are distinct yet coupled: the former concerns target credibility, while the latter varies with the student's current task competence; together, they shape supervision value. Building on this coupled view, we introduce Student-Curriculum Coupling (SCC), a closed-loop framework in which a compact Anchor-Frontier curriculum defines the candidate supervision space and the evolving student dynamically determines its active subset. Supervision can therefore be activated, suspended, or reactivated as competence changes, concentrating teacher computation and optimization on current task-level deficits. Across three TVG benchmarks, SCC achieves a 5.1% relative improvement in mean recall over Video-OPD on its original curriculum, while using 60.0% fewer training examples and reducing training time by 50.4%. Ablations support the complementary roles of capability-structured curriculum design and student-dependent supervision in achieving these gains. Together, these results establish SCC as a data- and compute-efficient framework for TVG post-training, delivering stronger temporal grounding by aligning trustworthy supervision with the student's evolving learning needs.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Counterfactual Attention Policy Distillation for Temporal Video Grounding

    Sep 28, 2026Shaobo Ju, Haiyang Yu, Xuecheng Wu +5Video Temporal GroundingMultimodal Large Language Models

  2. Is Better Teacher Supervision Enough? Unlocking Student-side Learning in Multimodal On-Policy Distillation

    Sep 30, 2026Siyuan Liu, Kanghui Tian, Yue Duan +4PerceptuallyVista

  3. VISD: Enhancing Video Reasoning via Structured Self-Distillation

    May 7, 2026Hao Lin, Kunyang Lv, Xu Jiang +5Latent Visual ReasoningProcess-Level Supervision