Counterfactual Attention Policy Distillation for Temporal Video Grounding
Authors: Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou
Organizations: Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China. · Fudan University. · MMLab, The Chinese University of Hong Kong. · University of the Chinese Academy of Sciences.
Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distillation (OPD) and propose a new training regime for MLLMs termed Counterfactual Attention Policy Distillation (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.
Figures & tables
Figure 1: Comparison of OPD, RAL-AD, and CAPD. OPD only transfers output distributions, RAL-AD distills teacher attention, and CAPD uses counterfactual influence to calibrate attention and weight output-level distillation.
Figure 2: Illustration of the CAPD framework. Given the temporally grouped video tokens and query, the student generates an on-policy trajectory and the frozen teacher provides next-token supervision. CAPD masks each temporal group (a), estimates its counterfactual influence from the resulting teacher-output changes (b), and uses this influence to calibrate teacher attention and weight output-level distillation (c). The final objective combines output-level OPD with calibrated attention-policy learning.
Method
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
Avg.
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
Proprietary Models
GPT-4o
60.6
44.5
23.5
55.2
41.4
25.8
69.0
54.8
38.5
45.9
GPT-5
59.3
42.0
22.0
57.4
44.9
30.4
72.4
60.4
46.4
48.4
Gemini-2.0-Flash
66.4
53.5
27.1
62.9
54.0
37.7
76.2
66.4
48.3
54.7
Gemini-2.5-Flash
68.7
56.1
30.6
66.8
57.5
41.3
78.2
69.4
55.0
58.2
Table 1: Comparison of CAPD with proprietary models, open-source models, and post-training frameworks, following Video-OPD ( Li et al. 2026b ) . Bold and underline indicate the best and second-best results among the post-training methods.
Method
Video-MME
LongVideoBench
LVBench
Qwen3-VL-8B-Instruct
70.8
64.0
53.0
+ Vanilla OPD
71.1
64.0
52.6
+ RAL-AD
71.2
63.8
52.4
+ CAPD
71.1
64.0
52.5
Table 2: Comparison of Qwen3-VL-8B-Instruct and its post-trained variants on general video understanding benchmarks.
Figure 3: Analysis of evidence coverage, preference gap, group precision, and selection stability under different numbers of temporal groups.
Method
C-TL
A-TL
QV-TL
Avg.
OPD
51.6
48.4
58.4
52.8
Attention-Only
51.8
48.6
59.3
53.3
Token-Only
52.0
49.4
60.3
53.9
CAPD
52.6
50.7
62.0
55.1
Table 3: Component ablation on TimeLens. C-TL, A-TL, and QV-TL are the means of R@0.3, R@0.5, and R@0.7 on each benchmark; Avg. averages all recalls.
Parameter
Value
C-TL
A-TL
QV-TL
Avg.
Temporal Group ( G )
2
52.2
49.8
60.7
54.2
4
51.9
49.5
60.5
54.0
8
52.6
50.7
62.0
55.1
16
52.0
49.5
60.4
53.9
Attention Weight ( λatt )
0
52.0
49.4
60.3
53.9
0.1
52.4
49.6
60.9
54.3
Table 4: Parameter ablation on TimeLens. C-TL, A-TL, and QV-TL are the mean recalls on each benchmark; Avg. averages all nine recall metrics.
Figure 4: Visualized results of CAPD. The top panel shows the temporal boundary predictions of different methods on a long-form video query, alongside the ground-truth interval, showcasing CAPD’s ability to produce compact and accurate temporal boundaries. The GREEN bar denotes the ground truth, the GRAY segments are OPD’s predictions, the BLUE line represents RAL-AD, and the PURPLE bar indicates CAPD’s output. The bottom-left and bottom-right panels contrast raw teacher attention with counterfactual (CF) influence across temporal groups G0–G7 for two additional queries, demonstrating that while raw attention concentrates on late or visually salient groups, CF influence more precisely identifies the relevant temporal segments.
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either fail to distinguish non-overlapping predictions or require fragile segment matching. TimeLens2 treats temporal evidence as an interval set throughout supervision and optimization. TimeLens2-93K constructs reliable multi-span supervision through caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. Our temporal Wasserstein reward computes exact one-dimensional W1 between uniform distributions over merged interval supports, providing dense, matching-free feedback under unequal cardinalities and equivalent fragmentation; temporal IoU complements it with precise-overlap feedback. Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters. The 2B, 4B, and 8B variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 mIoU points, respectively.
Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12
Nanjing University · Shanghai AI Laboratory · Shanghai Jiao Tong University +3
Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal large language models (MLLMs) understand not only what happens but also when it happens. Although modern MLLMs describe video content fluently, their timestamp predictions remain unreliable, while existing remedies either require costly post-training on temporal annotations or rely on coarse training-free heuristics. In this work, we probe the cross-modal attention of MLLMs and uncover a perception-generation gap. Our key finding is that MLLMs often know the target interval during prefill, but lose this signal when generating the final answer. In the prefill stage, a sparse set of attention heads, which we call Temporal Grounding Heads (TG-Heads), concentrates query-to-video attention on the ground-truth interval. During autoregressive decoding, however, the answer tokens shift attention away from this interval toward visually salient but query-irrelevant segments. This observation motivates an inference-time read-then-regenerate framework. We first convert TG-Head prefill attention into a debiased frame-level relevance signal and extract the high-attention interval it highlights. We then re-invoke the MLLM with visual context restricted to this interval using video cropping. Without parameter updates or architectural changes, our framework consistently improves MiMo-VL, Qwen3-VL, and Molmo2 on three benchmarks, with gains of up to +3.5 mIoU. The project website can be found at https://ddz16.github.io/mllmsknowwhen.github.io/.
Dazhao Du, Liao Duan, Jian Liu +5
Hong Kong University of Science and Technology · Tencent · Xi’an Jiaotong University
Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models (VLMs). We propose TEMPURA (Temporal Event Masked Prediction and Understanding for Reasoning in Action), a two-stage training framework that enhances the video temporal understanding of VLMs. Inspired by infilling techniques in language modeling, TEMPURA first performs masked event prediction, learning to reconstruct missing events and generate step-by-step causal explanations from dense event annotations. It then learns video segmentation and dense captioning, decomposing videos into non-overlapping events with detailed, timestamp-aligned descriptions. We train TEMPURA on VER, our large-scale dataset of 500K videos annotated with temporally aligned event descriptions and structured reasoning steps. Experiments on video temporal grounding and highlight detection benchmarks show that TEMPURA substantially improves strong base VLMs across model families and scales, confirming that combining event-level reasoning with fine-grained temporal segmentation is an effective recipe for video temporal understanding.
Jen-Hao Cheng, Yi-Hao Peng, Huapeng Zhou +11
University of Washington · Carnegie Mellon University · National Yang Ming Chiao Tung University +1