Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive frame is still tokenized on its own: the tokens are a function of the current primitives, not of a carried reference. We argue that a more natural function is of both---the current primitives and a carried reference. A clip and its time reversal share the same frames and differ only in the order of changes---an axis that symmetric pooling discards by construction, and that is non-empty in the frozen vision features VideoLMs use---and the codec recurrence already composes those changes in order against a reference state. We introduce RESUME, a stateful codec representation: an anchor I-frame initializes a compact latent state, each subsequent predictive frame is consumed as an update to that state, and a shared readout exposes VideoLM-compatible tokens from the accumulated state. Codec prediction is thereby kept at the representation level and handed to the language model as a trajectory, not as a set of independent token groups. At the same per-predictive-frame token budget as prior codec-aware methods, a predictive frame enters the language model as a readout of what the front-end already knows, not as an encoding of the current primitives alone. Across ten benchmarks, the gains concentrate on temporal reasoning: on all three temporal benchmarks RESUME improves over both the RGB-frame baseline LLaVA-Video-7B (by 2.8, 5.1, and 3.9 points on TempCompass, TOMATO, and MVBench) and the codec-based baseline CoPE-7B, while staying competitive on general and long-form QA. Frozen-transition tests further show anchor dependence, order sensitivity, and useful rollout behavior beyond the training horizon.
Figures & tables
Figure 1: Three ways to turn a group of pictures into visual tokens. (a) Dense RGB encoding treats every frame as a full image. (b) Deployed codec-aware tokenization reads motion vectors and residuals, but emits an independent token group per predictive frame; temporal dependence is left to language-model attention. (c) RESUME initializes a compact latent state from the anchor I-frame and updates that same state with each codec observation; the tokens of a predictive frame are a readout of the accumulated state. The two codec columns tokenize each predictive frame; only (c) maintains a reference-dependent state in the visual front-end.
Figure 2: RESUME as a codec-driven transition system. An anchor I-frame initializes a compact latent state z0 and, in parallel, keeps its own token path into the language model. Each predictive frame contributes one fused observation ot from motion vectors and residuals, written into the carried state. A shared head reads Yt from zt at the same per-predictive-frame token budget used by prior codec-aware front-ends. The state is reset at the next I-frame. A dedicated projector adapts state readouts to the language embedding space; the original multimodal projector is left untouched for I-frame tokens.
Table 1: Question answering across general, temporal, and long-form benchmarks. RESUME uses at most 64 I-frames, and each of the four predictive updates per I-frame produces one 8 -token readout. Prior numbers are as reported by CoPE-VideoLM (cited in the table); a dash means that source does not report the benchmark. Video-MME is without subtitles. ActivityNet-QA is scored by Claude Opus 4.8 ( Anthropic, 2026 ) for RESUME and by a language model for the rows above.
Figure 3: The reversal axis separates static content, net change, and temporal order. (a) We compare frame-symmetric (P0), difference-symmetric (P1), and order-sensitive (P2) probes on a clip and its time reversal. (b) The probes are evaluated on order questions and on two frozen vision towers.
Figure 4: The frozen transition carries state across predictive updates. (a) Anchor sensitivity, (b) order sensitivity, and (c) rollout behavior are compared with state controls.
Video-MME
LVBench
NI
I-frames only
+ readouts
I-frames only
+ readouts
8
53.3
54.1 +0.8
36.7
37.3 +0.6
16
57.3
58.6 +1.3
39.1
40.0 +0.9
Table 2: Ablation of the codec-driven state readouts under a matched I-frame budget. Each pair uses exactly the same NI I-frames; “+ readouts” augments them with the 256 state readouts. Accuracy (%) on Video-MME (without subtitles) and LVBench.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Codec primitives inside one group of pictures. The I-frame is an independently coded RGB image. Each predictive frame is stored as block-wise motion vectors τ (where existing content moves) and residuals δ (what motion compensation cannot explain). RESUME consumes (τ,δ) as a single observation that updates a state initialized from the I-frame, rather than encoding each predictive frame as an independent token group. Motion vectors are the bitstream quiver overlaid on a faded reconstruction of the same frame.
Figure 6: Video length covered at one frame per second, against the visual-token budget. Markers distinguish dense frame encoding from three dense-readout settings ( 4 , 8 , or 16 eight-token readouts per I-frame). The star is the published Gemini 2.5 Pro point.
Input
TTFT (s)
E2EL (s)
32 I-frames + 32 readout groups
0.461
1.795
16 I-frames + 48 readout groups
0.570
1.895
8 I-frames + 56 readout groups
0.631
1.954
64 I-frames (same weights)
0.619
1.938
64 frames, LLaVA-Video-7B
0.686
2.094
Appendix
Table 3: Inference latency for a 64 -second clip at one frame per second, generating 64 text tokens. Time to first token (TTFT) is prefill; E2EL is the time to emit 64 tokens.
Training data
Total
NextQA
PerceptionTest
Video-MME
LLaVA-Video ( Zhang et al., 2024c )
LLaVA-Hound
0.25 M
64.4
51.4
54.1
+ LLaVA-Video-178K
1.58 M
80.1
57.1
63.2
+ 3 QA datasets
1.64 M
80.1
69.0
61.9
+ LLaVA-OV (images)
2.74 M
83.2
67.9
63.4
LLaVA-Video-178K (sampled)
1.08 M
73.2
55.9
59.6
Appendix
Table 4: Training-data scale and distribution. LLaVA-Video rows are as reported in ( Zhang et al., 2024c ) across incremental training stages; RESUME fine-tunes Stage 2 on LLaVA-Video-178K only. The three QA datasets are the training splits of PerceptionTest, NextQA, and ActivityNet-QA.