cs.CVOct 1, 2026

Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

Authors: Youngwoo Shin, Yusung Ro, Minseo Kim, Junmo Kim

Organizations: Korea Advanced Institute of Science and Technology (KAIST)

Abstract

Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector τlτ_l, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts τlτ_l at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs

    May 8, 2026Peitao Han, Fei Cheng, Lis K. Pereira +2Fine-Grained Video UnderstandingSpatial Encoding

  2. TempCloze: Can Video-LLMs Identify the Missing Middle?

    Sep 1, 2026Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu +4Egocentric VideoMedium

  3. TimeThink: Reasoning with Time for Video LLMs

    Jul 6, 2026Handong Li, Longteng Guo, Zikang Liu +8Offline Reinforcement Learning