Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector τl, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts τl at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.
Figures & tables
Figure 1: A representative failure case of temporal reasoning. Logit lens [ 17 ] shows the higher-probability answer between Yes and No at each layer. The forward video maintains the correct answer throughout, while the reversed video shifts toward the correct answer at Lintermediate before subsequent layers overturn it, converging both inputs to the same prediction by Llast .
Figure 2: Layer-wise temporal divergence analysis across three VideoLLMs. (a) τ^l measured on frame-order-reversed pairs under temporal, spatial, and video-irrelevant questions. The temporal divergence profile, peaking in the mid-to-late layers and diminishing toward the output, emerges only under temporal questions, while both non-temporal conditions remain substantially attenuated. (b) Change in ground-truth probability pgt in %p when the last token is blocked from attending to preceding positions per layer. The largest drop aligns with the τ^l peak, confirming these layers are functionally critical for temporal predictions. Red shaded regions mark the peak layers.
Figure 3: Overview of Temporal Activation Injection (TAI). Given a video V , its time-reversed counterpart V~ , and question prompt Q , the reversed input is processed up to Lsrc and early-exited, bypassing all subsequent layers. τsteer is computed at Lsrc from the difference between forward and reversed last-token hidden states. The τ^ profile (top) provides the per-layer weights wl that scale the injection into each subsequent layer of the forward pass, producing the TAI output.
Models
Training- free
TempCompass
TVBench
Act
Attr
Dir
Ord
Spd
AVG
AC
AL
AS
ES
MD
OC
OS
ST
UA
AVG
Qwen2.5-VL-7B
-
94.9
77.2
59.1
75.8
59.2
73.4
25.9
37.5
62.9
42.0
32.3
55.4
40.0
82.2
31.7
44.6
+TCD
✓
94.7
80.0
60.0
77.6
58.8
74.2
26.1
40.6
63.8
42.5
33.6
54.1
39.1
82.2
30.5
45.0
+DINO-HEAL
✓
95.5
74.9
58.2
74.8
59.0
72.7
26.9
40.0
64.1
42.0
31.5
53.4
38.7
83.2
32.9
45.0
+ArrowRL
✗
95.3
82.8
60.6
76.4
59.7
75.0
32.5
38.1
66.8
43.5
34.9
54.7
31.6
80.5
46.3
46.9
+TAI (Ours)
✓
93.2
83.8
61.4
79.7
60.1
75.6
26.7
40.0
67.0
43.0
34.5
58.8
39.1
86.5
31.7
46.6
Table 1: Temporal reasoning results by category on TempCompass and TVBench. Bold and underline mark the best and second best per group, and blue highlights our method. All baselines are reproduced under our evaluation protocol, see Appendix K .
Table 5
Figure 4: TempCompass examples per category on Qwen2.5-VL-7B. Attribute Change, Direction, and Order produce high τ^Lsrc as their answers change under frame reversal, while Action and Speed produce low τ^Lsrc as they remain invariant. τ^Lsrc scales with reversal sensitivity without supervision.
Method
Action
Attr Change
Direction
Order
Speed
Baseline
94.9
77.2
59.1
75.8
59.2
TAI ( +β )
− 1.7
+ 6.6
+ 2.3
+ 3.9
+ 0.9
Anti-TAI ( −β )
− 0.4
− 36.3
− 12.6
− 27.4
− 0.7
τ^Lsrc†
0.059
0.169
0.074
0.138
0.057
Table 4: Steering specificity on TempCompass with Qwen2.5-VL-7B. Anti-TAI reverses injection direction and selectively degrades reversal-sensitive categories while leaving invariant ones largely unchanged. Red marks reversal-sensitive and orange moderately sensitive categories. † Per-category mean of τ^Lsrc computed from the temporal divergence profile in Sec. 3 .
Schedule
Accuracy
Baseline
73.4
Uniform ( wl=1 )
75.4
Reversed profile
75.1
Ours
75.6
Table 5: Injection schedule ablation on TempCompass.
Figure 5: Configuration sensitivity of Qwen2.5-VL-7B on TempCompass. (a) Frame count, (b) injection strength β , (c) source layer Lsrc , (d) extraction window size k .
Method
AoTBench
TempCompass
TVBench
Baseline
54.5
73.4
44.6
+ TAI
+3.2
+2.2
+2.0
ArrowRL
58.2
75.0
46.9
+ TAI
+2.1
+1.5
+0.9
Table 6: Orthogonality of TAI with other methods.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A: Temporal divergence decomposition across token positions and sub-layer computations. (a) Token-group τ^l profiles. (b) τ^l comparison between MLP block and layer output. (c, d) Per-head temporal sensitivity before (c) and after (d) the attention output projection.
Figure B: τ^l profile stability on Qwen2.5-VL-7B as the number of video pairs varies. (a) Overlaid profiles at varying n . (b) Pearson correlation with the n=50 reference profile.
Profile Source
Runs
Peak
TempCompass
TVBench
AoTBench
TempCompass
1
L20
75.58
46.58
57.70
TVBench
9
L20
-0.15
0.00
0.00
AoTBench
9
L20
-0.11
-0.19
+0.28
Appendix
Table A: Cross-domain τ^l profile estimation on Qwen2.5-VL-7B.
Model
LM Backbone
#Layers
Lsrc
Baseline
+TAI
Δ
Molmo2-O-7B
OLMo
32
23
75.8
77.6
+1.8
Gemma4-12B
Native
48
46
70.6
72.4
+1.8
Appendix
Table B: TempCompass results of TAI on Molmo2-O-7B and Gemma4-12B, with Lsrc set at the τ^l peak of each model. TAI improves over the baseline on both models.
Figure C: τ^l profiles on additional model families. Red shaded regions mark the peak layers.
Table 16
Models
Training -Free
Act. Antonym
Act. Count
Act. Local.
Act. Pred.
Act. Seq.
Char. Order
Counterfact.
Ego. Nav.
Fine Act.
Mov. Attr.
Mov. Count
Mov. Dir.
Obj. Exist.
Obj. Inter.
Obj. Shuf.
Scene Trans.
State Chg.
Unexp. Act.
AVG
Qwen2.5-VL-7B
-
74.5
46.5
41.0
58.0
67.0
73.5
70.5
32.0
46.0
91.0
62.5
52.5
86.9
67.5
44.0
91.0
54.5
74.0
62.9
+TCD
✓
76.0
46.0
42.0
61.0
70.2
73.5
66.5
31.0
46.0
92.0
63.5
52.0
88.9
66.0
43.0
89.0
54.5
74.0
63.0
+DINO-HEAL
✓
73.5
47.0
43.0
60.5
67.6
74.0
71.0
33.5
47.0
92.0
65.0
49.5
85.4
65.5
42.5
90.0
53.5
75.0
63.1
+ArrowRL
✗
74.5
42.5
41.0
60.0
70.2
68.0
64.0
31.5
45.5
92.0
67.0
52.0
85.4
66.0
31.5
88.0
54.0
68.0
61.1
+TAI (Ours)
✓
77.5
46.0
42.0
60.5
70.7
73.5
70.0
33.5
46.5
92.5
66.5
53.0
89.4
64.5
41.0
91.0
55.0
75.0
63.7
Qwen3-VL-8B
-
81.5
41.0
41.0
70.5
76.1
76.5
70.5
38.0
49.5
92.0
68.5
64.5
84.3
71.0
41.5
94.0
72.5
81.0
67.4
Appendix
Table E: General video understanding results on MVBench subtasks. Bold marks the best and underline the second-best per group; blue highlights our method. All methods use 16-frame input.
Figure D: Free-form generation under different injection positions. Responses to a sunrise video and its reversal. Correct and incorrect directions are marked in green and red .
Figure E: τ^Lsrc in the Direction category varies with the strength of the temporal cue. Each question is (a) uncertain, (b) ambiguous, and (c) clear to answer, and τ^Lsrc corresponds to this cue.
Subtask
ReverseFilm
UCF101
Rtime t2v
Rtime v2t
AoTBench QA
Reversal
Invariant
Sensitive
τ^Lsrc
0.042
0.024
0.033
0.190
0.172
Appendix
Table F: Per-subtask τ^Lsrc on AoTBench with Qwen2.5-VL-7B. Reversal-invariant subtasks produce near-zero τ^Lsrc , and TAI applies correspondingly minimal injection.
Benchmark
Median duration
Δt at n=16
Δt at n=32
TempCompass
10.0 s
0.63 s
0.31 s
MVBench
13.0 s
0.81 s
0.41 s
TVBench
20.0 s
1.25 s
0.63 s
AoTBench
20.0 s
1.25 s
0.63 s
Appendix
Table G: Inter-frame sampling intervals at nframes∈{16,32} on each benchmark, computed at the median video duration. At n=32 , the interval falls below the duration of typical events probed by these benchmarks (turning, jumping, pouring, etc.), producing redundant within-event sampling.
Models
Micro
Macro
Qwen2.5-VL-7B
54.5
53.1
+ArrowRL
58.2
56.3
+TAI (Ours)
57.7
55.3
Qwen3-VL-8B
56.6
54.6
+TAI (Ours)
59.9
57.0
InternVL2.5-8B
54.6
53.3
Appendix
Table H: Micro- and macro-averaged AoTBench accuracy.
Method
Act
Attr
Dir
Ord
Spd
AVG
Δ
Baseline
94.9
77.2
59.1
75.8
59.2
73.4
—
+ SEASON
93.2
72.2
53.7
71.8
58.3
70.1
-3.3
+ VTD
92.2
70.9
53.9
68.8
56.7
68.8
-4.6
Appendix
Table I: Reproduction of SEASON and VTD on Qwen2.5-VL-7B under our 16-frame protocol, averaged across all TempCompass question formats. Both methods degrade below the baseline, motivating their exclusion from our main comparison.
Method
Max. Fwd Passes
Peak Mem (GB)
Time (s/sample)
Baseline
1.0×
18.66
1.14±0.22
DINO-HEAL
1.0× + DINOv2
19.23
1.48±0.23
TCD
2.0×
18.70
1.56±0.20
VTD
2.0× + distort.
21.82
1.93±0.23
SEASON
3.0× + diag.
24.46
5.13±0.36
TAI (Ours)
∼1.7×
18.70
1.52±0.12
Appendix
Table J: Computational overhead comparison of inference-time methods on Qwen2.5-VL-7B with 16-frame input on a single A6000 GPU. Time averaged over 100 samples.
Figure F: TVBench example videos with Scene Transition, Action Sequence, Object Count, and Unexpected Action. Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
Figure G: TVBench example videos with Action Localization, Moving Direction, and Object Shuffle. Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
Figure H: AoTBench example videos with AoTBench QA and Rtime v2t . Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
Figure I: AoTBench example videos with UCF101, ReverseFilm, and Rtime t2v . Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
Figure J: MVBench example videos with Temporal-Relevant, Temporal-Irrelevant, and Hybrid groups. Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
School of Artificial Intelligence, University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · OPPO AI Center, OPPO Inc.