Counterfactual Attention Policy Distillation for Temporal Video Grounding
Authors: Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou
Organizations: Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China. · Fudan University. · MMLab, The Chinese University of Hong Kong. · University of the Chinese Academy of Sciences.
Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspective of On-policy distillation (OPD) and propose a new training regime for MLLMs termed Counterfactual Attention Policy Distillation (CAPD). In particular, OPD is a viable solution for MLLMs via providing dense teacher supervision on student-generated trajectories. But its next-token based teacher-student distillation is hard to identify the specific video segments supporting each predicted timestamp, which is critical for temporal grounding. In this case, CAPD measures how masking each temporal group changes the teacher's output distribution. The resulting counterfactual influence calibrates the teacher's attention and weights token-level distillation, allowing the student to learn the temporal evidence that affects boundary prediction. To validate CAPD, we trained it on Qwen3-VL-8B-Instruct using only 2,500 samples for one epoch, and evaluated it on the TimeLens and multiple general video benchmarks. Experimental results show that CAPD improves average recall by 12.0% relative to GRPO on TimeLens while preserving general video understanding, achieving comparable accuracy to the base model.
Figures & tables
Figure 1: Comparison of OPD, RAL-AD, and CAPD. OPD only transfers output distributions, RAL-AD distills teacher attention, and CAPD uses counterfactual influence to calibrate attention and weight output-level distillation.
Figure 2: Illustration of the CAPD framework. Given the temporally grouped video tokens and query, the student generates an on-policy trajectory and the frozen teacher provides next-token supervision. CAPD masks each temporal group (a), estimates its counterfactual influence from the resulting teacher-output changes (b), and uses this influence to calibrate teacher attention and weight output-level distillation (c). The final objective combines output-level OPD with calibrated attention-policy learning.
Method
Charades-TimeLens
ActivityNet-TimeLens
QVHighlights-TimeLens
Avg.
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
R@0.3
R@0.5
R@0.7
Proprietary Models
GPT-4o
60.6
44.5
23.5
55.2
41.4
25.8
69.0
54.8
38.5
45.9
GPT-5
59.3
42.0
22.0
57.4
44.9
30.4
72.4
60.4
46.4
48.4
Gemini-2.0-Flash
66.4
53.5
27.1
62.9
54.0
37.7
76.2
66.4
48.3
54.7
Gemini-2.5-Flash
68.7
56.1
30.6
66.8
57.5
41.3
78.2
69.4
55.0
58.2
Table 1: Comparison of CAPD with proprietary models, open-source models, and post-training frameworks, following Video-OPD ( Li et al. 2026b ) . Bold and underline indicate the best and second-best results among the post-training methods.
Method
Video-MME
LongVideoBench
LVBench
Qwen3-VL-8B-Instruct
70.8
64.0
53.0
+ Vanilla OPD
71.1
64.0
52.6
+ RAL-AD
71.2
63.8
52.4
+ CAPD
71.1
64.0
52.5
Table 2: Comparison of Qwen3-VL-8B-Instruct and its post-trained variants on general video understanding benchmarks.
Figure 3: Analysis of evidence coverage, preference gap, group precision, and selection stability under different numbers of temporal groups.
Method
C-TL
A-TL
QV-TL
Avg.
OPD
51.6
48.4
58.4
52.8
Attention-Only
51.8
48.6
59.3
53.3
Token-Only
52.0
49.4
60.3
53.9
CAPD
52.6
50.7
62.0
55.1
Table 3: Component ablation on TimeLens. C-TL, A-TL, and QV-TL are the means of R@0.3, R@0.5, and R@0.7 on each benchmark; Avg. averages all recalls.
Parameter
Value
C-TL
A-TL
QV-TL
Avg.
Temporal Group ( G )
2
52.2
49.8
60.7
54.2
4
51.9
49.5
60.5
54.0
8
52.6
50.7
62.0
55.1
16
52.0
49.5
60.4
53.9
Attention Weight ( λatt )
0
52.0
49.4
60.3
53.9
0.1
52.4
49.6
60.9
54.3
Table 4: Parameter ablation on TimeLens. C-TL, A-TL, and QV-TL are the mean recalls on each benchmark; Avg. averages all nine recall metrics.
Figure 4: Visualized results of CAPD. The top panel shows the temporal boundary predictions of different methods on a long-form video query, alongside the ground-truth interval, showcasing CAPD’s ability to produce compact and accurate temporal boundaries. The GREEN bar denotes the ground truth, the GRAY segments are OPD’s predictions, the BLUE line represents RAL-AD, and the PURPLE bar indicates CAPD’s output. The bottom-left and bottom-right panels contrast raw teacher attention with counterfactual (CF) influence across temporal groups G0–G7 for two additional queries, demonstrating that while raw attention concentrates on late or visually salient groups, CF influence more precisely identifies the relevant temporal segments.