Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector τl, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts τl at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.
Figures & tables
Figure 1: A representative failure case of temporal reasoning. Logit lens [ 17 ] shows the higher-probability answer between Yes and No at each layer. The forward video maintains the correct answer throughout, while the reversed video shifts toward the correct answer at Lintermediate before subsequent layers overturn it, converging both inputs to the same prediction by Llast .
Figure 2: Layer-wise temporal divergence analysis across three VideoLLMs. (a) τ^l measured on frame-order-reversed pairs under temporal, spatial, and video-irrelevant questions. The temporal divergence profile, peaking in the mid-to-late layers and diminishing toward the output, emerges only under temporal questions, while both non-temporal conditions remain substantially attenuated. (b) Change in ground-truth probability pgt in %p when the last token is blocked from attending to preceding positions per layer. The largest drop aligns with the τ^l peak, confirming these layers are functionally critical for temporal predictions. Red shaded regions mark the peak layers.
Figure 3: Overview of Temporal Activation Injection (TAI). Given a video V , its time-reversed counterpart V~ , and question prompt Q , the reversed input is processed up to Lsrc and early-exited, bypassing all subsequent layers. τsteer is computed at Lsrc from the difference between forward and reversed last-token hidden states. The τ^ profile (top) provides the per-layer weights wl that scale the injection into each subsequent layer of the forward pass, producing the TAI output.
Models
Training- free
TempCompass
TVBench
Act
Attr
Dir
Ord
Spd
AVG
AC
AL
AS
ES
MD
OC
OS
ST
UA
AVG
Qwen2.5-VL-7B
-
94.9
77.2
59.1
75.8
59.2
73.4
25.9
37.5
62.9
42.0
32.3
55.4
40.0
82.2
31.7
44.6
+TCD
✓
94.7
80.0
60.0
77.6
58.8
74.2
26.1
40.6
63.8
42.5
33.6
54.1
39.1
82.2
30.5
45.0
+DINO-HEAL
✓
95.5
74.9
58.2
74.8
59.0
72.7
26.9
40.0
64.1
42.0
31.5
53.4
38.7
83.2
32.9
45.0
+ArrowRL
✗
95.3
82.8
60.6
76.4
59.7
75.0
32.5
38.1
66.8
43.5
34.9
54.7
31.6
80.5
46.3
46.9
+TAI (Ours)
✓
93.2
83.8
61.4
79.7
60.1
75.6
26.7
40.0
67.0
43.0
34.5
58.8
39.1
86.5
31.7
46.6
Table 1: Temporal reasoning results by category on TempCompass and TVBench. Bold and underline mark the best and second best per group, and blue highlights our method. All baselines are reproduced under our evaluation protocol, see Appendix K .
Table 5
Figure 4: TempCompass examples per category on Qwen2.5-VL-7B. Attribute Change, Direction, and Order produce high τ^Lsrc as their answers change under frame reversal, while Action and Speed produce low τ^Lsrc as they remain invariant. τ^Lsrc scales with reversal sensitivity without supervision.
Method
Action
Attr Change
Direction
Order
Speed
Baseline
94.9
77.2
59.1
75.8
59.2
TAI ( +β )
− 1.7
+ 6.6
+ 2.3
+ 3.9
+ 0.9
Anti-TAI ( −β )
− 0.4
− 36.3
− 12.6
− 27.4
− 0.7
τ^Lsrc†
0.059
0.169
0.074
0.138
0.057
Table 4: Steering specificity on TempCompass with Qwen2.5-VL-7B. Anti-TAI reverses injection direction and selectively degrades reversal-sensitive categories while leaving invariant ones largely unchanged. Red marks reversal-sensitive and orange moderately sensitive categories. † Per-category mean of τ^Lsrc computed from the temporal divergence profile in Sec. 3 .
Schedule
Accuracy
Baseline
73.4
Uniform ( wl=1 )
75.4
Reversed profile
75.1
Ours
75.6
Table 5: Injection schedule ablation on TempCompass.
Figure 5: Configuration sensitivity of Qwen2.5-VL-7B on TempCompass. (a) Frame count, (b) injection strength β , (c) source layer Lsrc , (d) extraction window size k .
Method
AoTBench
TempCompass
TVBench
Baseline
54.5
73.4
44.6
+ TAI
+3.2
+2.2
+2.0
ArrowRL
58.2
75.0
46.9
+ TAI
+2.1
+1.5
+0.9
Table 6: Orthogonality of TAI with other methods.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A: Temporal divergence decomposition across token positions and sub-layer computations. (a) Token-group τ^l profiles. (b) τ^l comparison between MLP block and layer output. (c, d) Per-head temporal sensitivity before (c) and after (d) the attention output projection.
Figure B: τ^l profile stability on Qwen2.5-VL-7B as the number of video pairs varies. (a) Overlaid profiles at varying n . (b) Pearson correlation with the n=50 reference profile.
Profile Source
Runs
Peak
TempCompass
TVBench
AoTBench
TempCompass
1
L20
75.58
46.58
57.70
TVBench
9
L20
-0.15
0.00
0.00
AoTBench
9
L20
-0.11
-0.19
+0.28
Appendix
Table A: Cross-domain τ^l profile estimation on Qwen2.5-VL-7B.
Model
LM Backbone
#Layers
Lsrc
Baseline
+TAI
Δ
Molmo2-O-7B
OLMo
32
23
75.8
77.6
+1.8
Gemma4-12B
Native
48
46
70.6
72.4
+1.8
Appendix
Table B: TempCompass results of TAI on Molmo2-O-7B and Gemma4-12B, with Lsrc set at the τ^l peak of each model. TAI improves over the baseline on both models.
Figure C: τ^l profiles on additional model families. Red shaded regions mark the peak layers.
Table 16
Models
Training -Free
Act. Antonym
Act. Count
Act. Local.
Act. Pred.
Act. Seq.
Char. Order
Counterfact.
Ego. Nav.
Fine Act.
Mov. Attr.
Mov. Count
Mov. Dir.
Obj. Exist.
Obj. Inter.
Obj. Shuf.
Scene Trans.
State Chg.
Unexp. Act.
AVG
Qwen2.5-VL-7B
-
74.5
46.5
41.0
58.0
67.0
73.5
70.5
32.0
46.0
91.0
62.5
52.5
86.9
67.5
44.0
91.0
54.5
74.0
62.9
+TCD
✓
76.0
46.0
42.0
61.0
70.2
73.5
66.5
31.0
46.0
92.0
63.5
52.0
88.9
66.0
43.0
89.0
54.5
74.0
63.0
+DINO-HEAL
✓
73.5
47.0
43.0
60.5
67.6
74.0
71.0
33.5
47.0
92.0
65.0
49.5
85.4
65.5
42.5
90.0
53.5
75.0
63.1
+ArrowRL
✗
74.5
42.5
41.0
60.0
70.2
68.0
64.0
31.5
45.5
92.0
67.0
52.0
85.4
66.0
31.5
88.0
54.0
68.0
61.1
+TAI (Ours)
✓
77.5
46.0
42.0
60.5
70.7
73.5
70.0
33.5
46.5
92.5
66.5
53.0
89.4
64.5
41.0
91.0
55.0
75.0
63.7
Qwen3-VL-8B
-
81.5
41.0
41.0
70.5
76.1
76.5
70.5
38.0
49.5
92.0
68.5
64.5
84.3
71.0
41.5
94.0
72.5
81.0
67.4
Appendix
Table E: General video understanding results on MVBench subtasks. Bold marks the best and underline the second-best per group; blue highlights our method. All methods use 16-frame input.
Figure D: Free-form generation under different injection positions. Responses to a sunrise video and its reversal. Correct and incorrect directions are marked in green and red .
Figure E: τ^Lsrc in the Direction category varies with the strength of the temporal cue. Each question is (a) uncertain, (b) ambiguous, and (c) clear to answer, and τ^Lsrc corresponds to this cue.
Subtask
ReverseFilm
UCF101
Rtime t2v
Rtime v2t
AoTBench QA
Reversal
Invariant
Sensitive
τ^Lsrc
0.042
0.024
0.033
0.190
0.172
Appendix
Table F: Per-subtask τ^Lsrc on AoTBench with Qwen2.5-VL-7B. Reversal-invariant subtasks produce near-zero τ^Lsrc , and TAI applies correspondingly minimal injection.
Benchmark
Median duration
Δt at n=16
Δt at n=32
TempCompass
10.0 s
0.63 s
0.31 s
MVBench
13.0 s
0.81 s
0.41 s
TVBench
20.0 s
1.25 s
0.63 s
AoTBench
20.0 s
1.25 s
0.63 s
Appendix
Table G: Inter-frame sampling intervals at nframes∈{16,32} on each benchmark, computed at the median video duration. At n=32 , the interval falls below the duration of typical events probed by these benchmarks (turning, jumping, pouring, etc.), producing redundant within-event sampling.
Models
Micro
Macro
Qwen2.5-VL-7B
54.5
53.1
+ArrowRL
58.2
56.3
+TAI (Ours)
57.7
55.3
Qwen3-VL-8B
56.6
54.6
+TAI (Ours)
59.9
57.0
InternVL2.5-8B
54.6
53.3
Appendix
Table H: Micro- and macro-averaged AoTBench accuracy.
Method
Act
Attr
Dir
Ord
Spd
AVG
Δ
Baseline
94.9
77.2
59.1
75.8
59.2
73.4
—
+ SEASON
93.2
72.2
53.7
71.8
58.3
70.1
-3.3
+ VTD
92.2
70.9
53.9
68.8
56.7
68.8
-4.6
Appendix
Table I: Reproduction of SEASON and VTD on Qwen2.5-VL-7B under our 16-frame protocol, averaged across all TempCompass question formats. Both methods degrade below the baseline, motivating their exclusion from our main comparison.
Method
Max. Fwd Passes
Peak Mem (GB)
Time (s/sample)
Baseline
1.0×
18.66
1.14±0.22
DINO-HEAL
1.0× + DINOv2
19.23
1.48±0.23
TCD
2.0×
18.70
1.56±0.20
VTD
2.0× + distort.
21.82
1.93±0.23
SEASON
3.0× + diag.
24.46
5.13±0.36
TAI (Ours)
∼1.7×
18.70
1.52±0.12
Appendix
Table J: Computational overhead comparison of inference-time methods on Qwen2.5-VL-7B with 16-frame input on a single A6000 GPU. Time averaged over 100 samples.
Figure F: TVBench example videos with Scene Transition, Action Sequence, Object Count, and Unexpected Action. Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
Figure G: TVBench example videos with Action Localization, Moving Direction, and Object Shuffle. Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
Figure H: AoTBench example videos with AoTBench QA and Rtime v2t . Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
Figure I: AoTBench example videos with UCF101, ReverseFilm, and Rtime t2v . Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
Figure J: MVBench example videos with Temporal-Relevant, Temporal-Irrelevant, and Hybrid groups. Each row shows the input video, the question prompt, the ground-truth answer, and τ^Lsrc .
The Arrow-of-Time (AoT) task, determining whether a video plays forward or backward by recognizing temporal irreversibility, is one humans solve with near-perfect accuracy, yet frontier Video Large Language Models (Video-LLMs) perform only modestly above chance. This gap raises a key question: do visual backbones fail to encode temporal information, or does information bottleneck lie elsewhere in the Video-LLM architecture? We address this question by isolating the vision encoder from the Video-LLM and tracing temporal information across the encoder, projector, and LLM. We find that video-centric encoders with explicit temporal modeling encode strong temporal signals, whereas frame-centric encoders do not. However, when video-centric representations are passed through a standard Video-LLM architecture, performance often collapses, revealing a bottleneck of temporal information flow. We identify projector design as a key factor: Q-Former disrupts temporal information, while a time-preserved MLP projection substantially improves the LLM's access to such information. Our layer-wise analysis further shows temporal representation dynamics across encoder layers. Guided by these findings, we build a Video-LLM with temporal-aware video-centric encoder, time-preserved projector, and AoT supervision, surpassing human performance on AoTPPB with 98.1% accuracy, and improving broader temporal reasoning tasks by up to 6.0 points on VITATECS-Direction and 1.3 points on TVBench. Our results show that temporal reasoning in Video-LLMs requires both effective temporal encoding and reliable transfer of this information to the LLM.
Peitao Han, Fei Cheng, Lis K. Pereira +2
The University of Osaka · Center for Information and Neural Networks · National Institute of Information and Communications Technology +2
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.
Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu +4
The University of Hong Kong · National University of Singapore · Peking University
Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Video-LLMs) have shown promising reasoning abilities when aligned with reinforcement learning, yet existing approaches typically rely on outcome-based rewards that supervise only the final prediction. Such supervision provides limited guidance on how models should discover the relevant temporal evidence during intermediate reasoning. In this work, we propose TimeThink, a reinforcement learning framework that explicitly guides temporal evidence discovery in Video-LLMs. Our key idea is to treat temporal clue steps as the fundamental optimization primitive of video reasoning, where each reasoning step references a candidate time interval in the video. We introduce a step-wise temporal process reward that provides localized credit assignment for these clues and a joint process--outcome optimization objective that balances reasoning fidelity with task correctness. To enable scalable training, we construct TimeThink-RFT-20K, a dataset with automatically derived temporal evidence segments. Extensive experiments across video reasoning, temporal grounding, and general video understanding benchmarks show that TimeThink consistently improves both temporal localization and reasoning performance, achieving state-of-the-art results among open-source video RL models.
Handong Li, Longteng Guo, Zikang Liu +8
School of Artificial Intelligence, University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · OPPO AI Center, OPPO Inc.