cs.CVOct 5, 2026

MeSD: Multi-Evidence Self-Distillation for VideoLLM

Authors: Weijie Zhu, Han Fang, Hanyu Fu, Yuzhe Zhang, Xin Wei, Zhaoyan Pan, Feiran Liu, Xunjie Jin, +8 more

Organizations: UCAS · Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd · PKU · ZJU · BUAA

Abstract

While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A further challenge lies in determining whether teacher guidance should refine reward-based updates or provide corrective supervision for failed trajectories. To address these issues, we propose MeSD, a multi-evidence self-distillation framework for VideoLLMs. MeSD constructs three evidence-conditioned teachers with shared parameters, using the ground-truth answer as a common semantic context while separately incorporating temporal and spatial evidence. Given the same student-generated prefixes, MeSD evaluates evidence-specific preferences relative to the Answer Teacher and fuses teacher-common preferences with gated teacher-specific residuals. Furthermore, MeSD introduces Verification-Guided Optimization to classify trajectories as Success, Failure, or Indeterminate. For Success and Indeterminate trajectories, MeSD refines token-level advantage magnitudes while preserving reward-derived signs. For verified failure trajectories that contain the required evidence, MeSD applies failure-conditioned distillation, using reverse-KL correction toward the fused distribution. Experiments on multiple video benchmarks demonstrate consistent gains over reinforcement learning and self-distillation baselines.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VISD: Enhancing Video Reasoning via Structured Self-Distillation

    May 7, 2026Hao Lin, Kunyang Lv, Xu Jiang +5Latent Visual ReasoningProcess-Level Supervision

  2. EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models

    May 21, 2026Shiqi Huang, Ziyue Wang, Zhongrong Zuo +3Self-Evolution