The design of a video encoder determines when frames begin to interact and which frame-specific visual evidence remains accessible to the language model. Native video pathways couple neighboring frames during visual encoding, whereas image pathways preserve independently computed frame representations but incur a much larger visual-token cost when all image tokens are forwarded. We ask a basic question: whether a compact video encoder can instead be built on image representations. To answer this question, we separate three operations that are often coupled: per-frame representation, cross-frame token allocation, and temporal interaction. A frozen image encoder first produces frame-specific candidates. A question-aware selector then allocates a fixed token budget across frames using relevance, diversity, and cross-frame correspondence, after which a lightweight learned refiner reads neighboring-frame context and writes residual updates only to the retained anchors. This preserves source positions and keeps the visual output at the fixed budget. Across 13 benchmarks and three vision-language backbones, the resulting pathway matches full-image aggregate performance while using only about 28%-35% of its visual tokens. Specifically, on Qwen3-VL-8B, it achieves a 13-benchmark macro-average of 62.75 with 1,535 visual tokens, compared with 62.58 for the full Image pathway at 4,424 tokens and 59.49 for native Conv3D at 2,212 tokens. On Qwen3-VL-32B, it reaches a 13-benchmark macro-average of 66.28, compared with 66.09 for Image, while providing a 2.16x end-to-end speedup. These results show that compact video encoding does not require early temporal mixing: frame-specific evidence can be preserved first, allocated jointly, and temporally contextualized after selection.
Figures & tables
Figure 1: Image-first video encoding. (a) Independent encoding preserves frame-specific candidates; coordinated selection allocates the token budget; and optional refinement adds temporal context without extra output tokens. (b–c) Results compare complete pathways.
Figure 2: Three video-input pathways. (a) Conv3D mixes neighboring frames before selection. (b) Full Image preserves frame-specific tokens but forwards the dense sequence. (c) Ours selects B source anchors before temporal refinement , preserving B→B . Highlighted patches indicate source support only.
Figure 3: Selection before temporal refinement. Left: native encoding mixes neighboring frames before selection. Right: framewise candidates remain separate through budgeted selection , followed by sparse residual interaction . The displayed (IB+ASS)ZS uses selected-token context; the refiner may also read unselected neighbors. Matrix blocks indicate dependencies, not linearity.
Method
Visual tokens ↓
Motion Bench
Video- MME
Video- MME-v2
NExT-QA
Perception Test
Perception Comp
Avg. (6) ↑
Qwen2-VL-7B
Video/Conv3D
2,780
53.61
57.78
22.50
80.33
59.17
18.39
48.63
Image
5,561
51.62
62.78
22.03
81.55
58.65
28.25
50.81
Uniform
1,536
50.11
51.65
20.35
76.88
55.42
17.65
45.34
FlashVID
1,536
51.27
53.64
21.67
78.89
57.39
19.13
47.00
VidCom 2
1,529
51.84
54.78
22.37
79.77
58.02
19.24
47.67
Table 1: Six benchmarks. Avg.(6) is the unweighted mean. Bold and underlined values mark best and second-best budget-matched results; Image and Video/Conv3D are full-input references.
Figure 4: General video understanding across three backbones. Scores are normalized to the corresponding Image pathway (100%). The shared radial range is 60%–110% and does not start at zero. FastV, SparseVLM, and DyCoke are evaluated only on Qwen3-VL-8B. Exact scores are reported in Appendix B .
Method
Visual tokens ↓
Selection (ms) ↓
TTFT (ms) ↓
E2E (ms) ↓
Throughput (q/s) ↑
Memory (GB) ↓
Avg.(13) ↑
Δ Image (pp) ↑
Speedup ↑
Video/Conv3D
2212.0
–
454.0
508.5
1.97
67.97
64.20
-1.89
1.79 ×
Image
4424.0
–
849.1
911.2
1.10
69.04
66.09
0.00
1.00 ×
FlashVID
1507.4
43.59
297.7
330.5
3.03
67.52
54.32
-11.77
2.76 ×
VidCom 2
1515.3
3.28
348.4
403.3
2.48
67.88
63.86
-2.23
2.26 ×
FastVID
1515.0
15.38
356.7
413.8
2.42
67.81
63.58
-2.51
2.20 ×
VisionZip
1532.0
57.16
409.2
459.4
2.18
67.82
63.29
-2.80
1.98 ×
Table 2: Inference efficiency on Qwen3-VL-32B-Instruct. Speedup is relative to Image using end-to-end latency.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Method
EGO
VINO text
VideoAds
LVBench
MVBench
STAR
IntentQA
Avg.(13) ↑
Qwen2-VL-7B
Video/Conv3D
63.00
65.00
51.82
39.35
63.19
76.06
94.85
57.31
Image
70.00
65.50
58.64
41.94
60.97
75.07
93.44
59.26
Uniform
61.35
58.74
46.52
36.15
60.18
71.44
91.22
53.67
FlashVID
62.18
59.83
49.46
38.08
62.33
73.37
93.14
55.41
VidCom 2
63.05
60.48
49.12
38.36
63.22
74.11
94.42
56.06
Appendix
Table 3: Additional seven-benchmark results. Avg.(13) is the unweighted mean over all 13 benchmarks. Bold and underlined values denote the best and second-best results within each backbone.
Method
Train
Visual tokens ↓
Avg.(6) ↑
Avg.(13) ↑
Δ prior TF ↑
Qwen2-VL-7B
Full Image
–
5,561
50.81
59.26
–
VisionZip
TF
1,532
48.15
56.88
0.00
Ours: selection
TF
1,536
50.97
59.37
+2.49
Ours: + refiner
Learned
1,536
51.11
59.43
+2.55
Qwen3-VL-8B
Appendix
Table 4: Training-fair comparison. TF denotes training-free compression; Selection disables the learned refiner. Δ is measured against the strongest prior TF baseline.
Dynamic Allocation
Cross-Frame Coordination
Answer Loss
KD Loss
VideoAds
Video-MME
LVBench
P.Test
Avg.(13)
✗
✗
–
–
55.45
63.85
39.80
67.85
60.12
✓
✗
–
–
60.30
65.40
42.15
68.30
61.74
✓
✓
–
–
62.55
66.25
42.26
68.65
62.64
✓
✓
✓
✗
61.85
66.80
42.85
68.80
62.67
✓
✓
✗
✓
61.95
67.55
43.60
69.10
62.71
✓
✓
✓
✓
62.10
68.05
44.05
69.25
62.75
Appendix
Table 5: Component and refiner-objective ablations on Qwen3-VL-8B. All variants use 1,535 visual tokens; dashes indicate that no learned refiner is used.
Dimension
Setting
Candidate frames
Token budget
Actual tokens ↓
Avg.(13) ↑
Δ
Selection
Independent
32
1,536
1,534.9
58.50
−0.93
Shuffled correspondence
32
1,536
1,534.9
57.53
−1.90
Coupled (Ours)
32
1,536
1,534.9
59.43
0.00
Candidate frames
16 frames
16
1,536
1,517.7
57.97
−1.46
24 frames
24
1,536
1,530.1
58.85
−0.58
32 frames (Ours)
32
1,536
1,534.9
59.43
0.00
Appendix
Table 6: System-level ablations on Qwen2-VL-7B. We vary the selection strategy, number of candidate frames, and visual-token budget. Δ is measured relative to the default configuration.
Figure 5: Inference latency on Qwen3-VL-32B-Instruct. Selection is included in TTFT. Ours reduces end-to-end latency from 911.2 ms to 421.8 ms relative to Image.
Figure 6: Multi-budget accuracy–latency trade-off. We sweep 0.75K, 1.5K, and 3K visual-token budgets on three backbones. Each curve connects the three operating points of one method.
Figure 7: Representative qualitative successes (I). Four examples illustrate different forms of evidence preservation. In (a), the answer depends on a subtle downward head movement. In (b) and (d), the decision requires preserving interactions among multiple objects across time. In (c), camera motion is visible from the relative displacement of the cup group.
Figure 8: Representative qualitative successes (II). The examples cover late evidence, repeated outcomes, object counting, and directional temporal reasoning. In (a), the decisive event occurs near the end of the clip when the cat approaches and licks the hand. In (b), the answer requires comparing repeated trials and identifying the one with a different outcome. In (c), the model must avoid double-counting repeated views of objects. In (d), temporal order and spatial roles determine the throw direction. Ours predicts the ground-truth option in all four cases.
Figure 9: Representative failure case: long-range subject binding. The question asks what the farther parrot does after spreading its wing near the end of the clip. The full Image pathway predicts the correct answer ( E : continuing to clean itself), whereas Ours, Video/Conv3D, and all compact baselines shown here predict C . The example suggests that aggressive visual compression can make it difficult to preserve the identity of the queried subject over a long temporal span when several visually similar instances are present.
Input: video V , question q , budget B , frozen image encoder, frozen allocator, optional refiner Gθ .
1
Encode every frame independently ; collect (Z,P,D) .
2
Inherit quotas bt from the frozen allocator, with 0≤bt≤Nt and ∑tbt=B .
3
Build question-weighted DPP kernels {Lt} and pair graph (E,w) .
4
S←∅ .
5
while ∣S∣<B do
6
U←FeasibleMoves(S,{bt},E) .
Appendix
Algorithm 1 Image-first encoding with a fixed output budget.
Question weight
Motion weight
Saliency weight
Diversity weight
Frame floor
0.8
0.6
0.4
0.4
12
Preview frames
Preview boost
Uncertainty weight
Temporal NMS gap
DPP ratio
5
0.75
0.45
2
0.5
Appendix
Table 7: Frozen SITE allocator configuration.
Parameter
Value
Parameter
Value
Relevance weight η
0.8
Scene-similarity threshold
0.65
Minimum raw quality
10−6
Confidence margin scale
0.05
Diagonal-jitter coefficient
10−3
Spatial-change scale
0.25
Residual floor
10−14
Feature-change multiplier
2.0
Cosine threshold
0.65
Pair-reward weight γ
2.0
Margin threshold
0.005
Context size K
8
Appendix
Table 8: Selection and refinement implementation constants.
Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each temporal segment is equipped with learnable abstract tokens that learn a compact segment-level representation, while fine-grained patch tokens remain available throughout the encoder. Alternating intra-segment and abstract-communication layers preserve video-level context through the abstract-token channel. A lightweight selector further exposes either abstract tokens alone or abstract tokens augmented with a runtime-selected subset of patch tokens, yielding a compact visual interface that reduces the visual context and prefill burden of downstream MLLMs while retaining fine-grained evidence when needed. Pretrained with contrastive objectives on 565M image--text pairs and 6.4M videos, CoVisco shows competitive performance on video-oriented embedding and multimodal understanding benchmarks. In the evaluated four-segment, 64-frame setting, abstract-only inference uses only 400 visual tokens while achieving video-understanding performance close to, and on some benchmarks exceeding, OneVision-Encoder. Selected patch tokens further improve fine-grained video reasoning. Project URL: https://github.com/ernie-research/CoVisco.git
Yulong Liu, Xiaotian Han, Junyuan Shang +6
ERNIE Team, Baidu Inc. · The Hong Kong University of Science and Technology · Institute of Automation, Chinese Academy of Sciences (CASIA)
The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long-form video lies in managing the explosion of visual-token context length. Existing strategies predominantly focus on "post-hoc" token reduction -- reducing visual tokens after feature extraction to alleviate the LLM's computational overhead. While these methods effectively reduce the number of visual tokens, we observe that the primary latency bottleneck then shifts from the LLM to the expensive per-frame processing of the vision encoder. To address this, we introduce LiteFrame, a strong, yet highly efficient video encoder backbone for Video LLMs. To train LiteFrame, we propose Compressed Token Distillation (CTD), a novel training framework that teaches a compact student vision encoder to directly predict information-dense, spatio-temporally compressed representations produced by a large teacher vision model, effectively bypassing redundant computation. When coupled with further Language Model Adaptation (LMA), this approach results in a new latency-accuracy Pareto frontier -- compared with InternVL3-8B, LiteFrame provides a 35% reduction in end-to-end latency while processing 8× more frames and improves average video understanding accuracy across multiple benchmarks. Our results demonstrate a new potential path to unlocking longer-form video understanding under fixed compute budgets.
Video is temporally redundant: adjacent frames usually share most objects, background, and layout. Yet existing video multimodal large language models (video MLLMs) usually encode each sampled frame as an independent RGB image, causing visual tokens to repeat content already present in earlier frames. This suggests a more direct video interface: send a full reference frame only when the scene cannot be predicted well from prior context, and otherwise transmit a compact description of inter-frame changes. We call this interface a \emph{predictive visual code}, and instantiate it for video MLLMs as \textbf{AdaCodec}. AdaCodec spends full visual tokens on a reference frame only when its conditional predictive cost is high; otherwise, it encodes inter-frame changes, including motion and prediction residuals, as compact P-tokens. Across all eleven benchmarks, AdaCodec improves over the Qwen3-VL-8B per-frame RGB baseline at a matched visual-token budget. Even at 1/7 the budget, AdaCodec with 32k tokens surpasses the 224k baseline on all long-video benchmarks; on five general-video benchmarks, it raises the average score while substantially cutting time-to-first-token from 9.26s to 1.62s.
Haowen Hou, Zhen Huang, Zheming Liang +8
Shanghai Jiao Tong University · Shanghai Innovation Institute