cs.CVSep 30, 2026

Frame Differential On-Policy Self-Distillation for Video Reasoning

Authors: Haiying He, Xin Zheng, Shaoli Hu, Shijun Xiao, Xuanhe Liu, Bing Li, Harry Yang

Organizations: HKUST · NKU · SEU · KAUST

Abstract

Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: increasing the number of frames makes autoregressive rollouts expensive, while too few frames may miss temporally localized events and fine-grained visual details. We present \textbf{Frame Differential On-Policy Self-Distillation (FD-OPSD)}, which transfers the useful evidence of dense frame observations to a sparse frame policy during RL training. FD-OPSD compares the policy's token level preferences for the same sampled response under sparse and dense views, and distills the resulting frame differential signal without an external teacher or dense autoregressive rollout. The method preserves sparse-frame rollouts and leaves inference unchanged. Across Qwen2.5-VL-7B and Qwen3-VL-4B on six video reasoning benchmarks, FD-OPSD yields higher overall average performance than the strongest corresponding GRPO, T-GRPO, or Video-KTR baselines across the 16, 32, and 64 frame evaluation settings. These results show that dense visual evidence can be transferred selectively during training through token level self-distillation while retaining sparse frame rollouts and unchanged inference.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VISD: Enhancing Video Reasoning via Structured Self-Distillation

    May 7, 2026Hao Lin, Kunyang Lv, Xu Jiang +5Latent Visual ReasoningProcess-Level Supervision

  2. Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning

    Aug 17, 2026Ao Shen, Yongheng Zhang, Yinghui Li +3Latent Visual ReasoningLarge Multimodal Models

  3. Reasoning as Intersection: Consensus-Frame Alignment for Visual Focus in Video-MLLMs

    Jun 16, 2026Chengwen Liu, Zhe Huang, Jisheng Dang +3Video UnderstandingMultimodal Large Language Models