RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling
Organizations: Beijing University of Chemical Technology · Baidu, Inc.
Abstract
Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes between sampled frames. Codec-aware front-ends read the motion vectors and residuals that encoding produced, but in their deployed form each predictive frame is still tokenized on its own: the tokens are a function of the current primitives, not of a carried reference. We argue that a more natural function is of both---the current primitives and a carried reference. A clip and its time reversal share the same frames and differ only in the order of changes---an axis that symmetric pooling discards by construction, and that is non-empty in the frozen vision features VideoLMs use---and the codec recurrence already composes those changes in order against a reference state. We introduce RESUME, a stateful codec representation: an anchor I-frame initializes a compact latent state, each subsequent predictive frame is consumed as an update to that state, and a shared readout exposes VideoLM-compatible tokens from the accumulated state. Codec prediction is thereby kept at the representation level and handed to the language model as a trajectory, not as a set of independent token groups. At the same per-predictive-frame token budget as prior codec-aware methods, a predictive frame enters the language model as a readout of what the front-end already knows, not as an encoding of the current primitives alone. Across ten benchmarks, the gains concentrate on temporal reasoning: on all three temporal benchmarks RESUME improves over both the RGB-frame baseline LLaVA-Video-7B (by 2.8, 5.1, and 3.9 points on TempCompass, TOMATO, and MVBench) and the codec-based baseline CoPE-7B, while staying competitive on general and long-form QA. Frozen-transition tests further show anchor dependence, order sensitivity, and useful rollout behavior beyond the training horizon.
Figures & tables
| (a) General QA PT NQA AQA VMME Proprietary GPT-5 [-1pt] ( OpenAI, 2025 ) – 86.3 – 83.3 Gemini 3 Pro [-1pt] ( Google, 2025 ) – 84.3 – 88.6 Gemini 2.5 Pro [-1pt] ( Gemini Team, 2025 ) – 85.3 – 87.8 Claude 4.5 [-1pt] ( Anthropic, 2025 ) – 79.2 – 74.2 Open-source VILA-40B [-1pt] ( Lin et al., 2024 ) 54.0 67.9 58.0 60.1 IXC-2.5-7B [-1pt] ( Zhang et al., 2024a ) 34.4 71.0 52.8 55.8 LLaVA-OV-7B [-1pt] ( Li et al., 2024a ) 57.1 79.4 56.6 58.2 Oryx-7B [-1pt] ( Liu et al., 2024c ) 68.6 81.9 – 58.3 LLaVA-Video-7B [-1pt] ( Zhang et al., 2024c ) 67.9 83.2 56.5 63.3 CoPE-7B [-1pt] ( Sarkar et al., 2026 ) 70.3 82.1 60.3 61.9 RESUME 71.2 81.9 60.9 61.9 | (b) Temporal TC TOM MVB Proprietary GPT-5 80.4 53.0 74.1 Gemini 3 Pro 82.8 48.3 70.4 Gemini 2.5 Pro 81.9 48.6 70.6 Claude 4.5 72.8 39.6 62.1 Open-source IXC-2.5-7B 67.1 – 69.1 LLaVA-OV-7B 64.8 25.5 56.7 VideoLLaMA2 [-1pt] ( Cheng et al., 2024 ) – 18.5 54.6 InternVL2-8B [-1pt] ( Chen et al., 2024b ) 65.3 21.7 65.8 VideoChat2-7B [-1pt] ( Li et al., 2023b ) 45.5 – 51.1 LLaVA-Video-7B 66.6 24.9 58.6 CoPE-7B 68.9 28.3 61.9 RESUME 69.4 30.0 62.5 | (c) Long-form VTT VMMU LVB Proprietary GPT-5 – – 68.8 Gemini 3 Pro – – 78.0 Gemini 2.5 Pro – – 78.4 Claude 4.5 – – 50.5 Open-source LongVA-7B [-1pt] ( Zhang et al., 2024b ) – 23.9 – LLaVA-OV-7B 44.0 33.9 38.1 InternVL2-8B – 37.4 – LLaVA-Video-7B 41.8 36.1 44.2 CoPE-7B 45.5 38.2 46.4 RESUME 45.7 38.2 43.2 |
| Video-MME | LVBench | |||
|---|---|---|---|---|
| I-frames only | + readouts | I-frames only | + readouts | |
| +0.8 | +0.6 | |||
| +1.3 | +0.9 | |||
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Input | TTFT (s) | E2EL (s) |
|---|---|---|
| I-frames + readout groups | ||
| I-frames + readout groups | ||
| I-frames + readout groups | ||
| I-frames (same weights) | ||
| frames, LLaVA-Video-7B |
| Training data | Total | NextQA | PerceptionTest | Video-MME |
|---|---|---|---|---|
| LLaVA-Video ( Zhang et al., 2024c ) | ||||
| LLaVA-Hound | M | |||
| + LLaVA-Video-178K | M | |||
| + 3 QA datasets | M | |||
| + LLaVA-OV (images) | M | |||
| LLaVA-Video-178K (sampled) | M | |||