cs.CVSep 28, 2026

Counterfactual Attention Policy Distillation for Temporal Video Grounding

Authors: Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou

Organizations: Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China. · Fudan University. · MMLab, The Chinese University of Hong Kong. · University of the Chinese Academy of Sciences.

Abstract

Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distillation (OPD) and propose a new training regime for MLLMs termed Counterfactual Attention Policy Distillation (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.

Figures & tables

Explore similar work

CardsList
  1. TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

    Jul 19, 2026Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12Video Temporal GroundingTemporal Grounding

  2. MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

    May 21, 2026Dazhao Du, Liao Duan, Jian Liu +5Video Temporal GroundingTemporal Grounding

  3. TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

    May 2, 2025Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou +11Video Temporal GroundingTemporal Grounding