The design of a video encoder determines when frames begin to interact and which frame-specific visual evidence remains accessible to the language model. Native video pathways couple neighboring frames during visual encoding, whereas image pathways preserve independently computed frame representations but incur a much larger visual-token cost when all image tokens are forwarded. We ask a basic question: whether a compact video encoder can instead be built on image representations. To answer this question, we separate three operations that are often coupled: per-frame representation, cross-frame token allocation, and temporal interaction. A frozen image encoder first produces frame-specific candidates. A question-aware selector then allocates a fixed token budget across frames using relevance, diversity, and cross-frame correspondence, after which a lightweight learned refiner reads neighboring-frame context and writes residual updates only to the retained anchors. This preserves source positions and keeps the visual output at the fixed budget. Across 13 benchmarks and three vision-language backbones, the resulting pathway matches full-image aggregate performance while using only about 28%-35% of its visual tokens. Specifically, on Qwen3-VL-8B, it achieves a 13-benchmark macro-average of 62.75 with 1,535 visual tokens, compared with 62.58 for the full Image pathway at 4,424 tokens and 59.49 for native Conv3D at 2,212 tokens. On Qwen3-VL-32B, it reaches a 13-benchmark macro-average of 66.28, compared with 66.09 for Image, while providing a 2.16x end-to-end speedup. These results show that compact video encoding does not require early temporal mixing: frame-specific evidence can be preserved first, allocated jointly, and temporally contextualized after selection.
Figures & tables
Figure 1: Image-first video encoding. (a) Independent encoding preserves frame-specific candidates; coordinated selection allocates the token budget; and optional refinement adds temporal context without extra output tokens. (b–c) Results compare complete pathways.
Figure 2: Three video-input pathways. (a) Conv3D mixes neighboring frames before selection. (b) Full Image preserves frame-specific tokens but forwards the dense sequence. (c) Ours selects B source anchors before temporal refinement , preserving B→B . Highlighted patches indicate source support only.
Figure 3: Selection before temporal refinement. Left: native encoding mixes neighboring frames before selection. Right: framewise candidates remain separate through budgeted selection , followed by sparse residual interaction . The displayed (IB+ASS)ZS uses selected-token context; the refiner may also read unselected neighbors. Matrix blocks indicate dependencies, not linearity.
Method
Visual tokens ↓
Motion Bench
Video- MME
Video- MME-v2
NExT-QA
Perception Test
Perception Comp
Avg. (6) ↑
Qwen2-VL-7B
Video/Conv3D
2,780
53.61
57.78
22.50
80.33
59.17
18.39
48.63
Image
5,561
51.62
62.78
22.03
81.55
58.65
28.25
50.81
Uniform
1,536
50.11
51.65
20.35
76.88
55.42
17.65
45.34
FlashVID
1,536
51.27
53.64
21.67
78.89
57.39
19.13
47.00
VidCom 2
1,529
51.84
54.78
22.37
79.77
58.02
19.24
47.67
Table 1: Six benchmarks. Avg.(6) is the unweighted mean. Bold and underlined values mark best and second-best budget-matched results; Image and Video/Conv3D are full-input references.
Figure 4: General video understanding across three backbones. Scores are normalized to the corresponding Image pathway (100%). The shared radial range is 60%–110% and does not start at zero. FastV, SparseVLM, and DyCoke are evaluated only on Qwen3-VL-8B. Exact scores are reported in Appendix B .
Method
Visual tokens ↓
Selection (ms) ↓
TTFT (ms) ↓
E2E (ms) ↓
Throughput (q/s) ↑
Memory (GB) ↓
Avg.(13) ↑
Δ Image (pp) ↑
Speedup ↑
Video/Conv3D
2212.0
–
454.0
508.5
1.97
67.97
64.20
-1.89
1.79 ×
Image
4424.0
–
849.1
911.2
1.10
69.04
66.09
0.00
1.00 ×
FlashVID
1507.4
43.59
297.7
330.5
3.03
67.52
54.32
-11.77
2.76 ×
VidCom 2
1515.3
3.28
348.4
403.3
2.48
67.88
63.86
-2.23
2.26 ×
FastVID
1515.0
15.38
356.7
413.8
2.42
67.81
63.58
-2.51
2.20 ×
VisionZip
1532.0
57.16
409.2
459.4
2.18
67.82
63.29
-2.80
1.98 ×
Table 2: Inference efficiency on Qwen3-VL-32B-Instruct. Speedup is relative to Image using end-to-end latency.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Method
EGO
VINO text
VideoAds
LVBench
MVBench
STAR
IntentQA
Avg.(13) ↑
Qwen2-VL-7B
Video/Conv3D
63.00
65.00
51.82
39.35
63.19
76.06
94.85
57.31
Image
70.00
65.50
58.64
41.94
60.97
75.07
93.44
59.26
Uniform
61.35
58.74
46.52
36.15
60.18
71.44
91.22
53.67
FlashVID
62.18
59.83
49.46
38.08
62.33
73.37
93.14
55.41
VidCom 2
63.05
60.48
49.12
38.36
63.22
74.11
94.42
56.06
Appendix
Table 3: Additional seven-benchmark results. Avg.(13) is the unweighted mean over all 13 benchmarks. Bold and underlined values denote the best and second-best results within each backbone.
Method
Train
Visual tokens ↓
Avg.(6) ↑
Avg.(13) ↑
Δ prior TF ↑
Qwen2-VL-7B
Full Image
–
5,561
50.81
59.26
–
VisionZip
TF
1,532
48.15
56.88
0.00
Ours: selection
TF
1,536
50.97
59.37
+2.49
Ours: + refiner
Learned
1,536
51.11
59.43
+2.55
Qwen3-VL-8B
Appendix
Table 4: Training-fair comparison. TF denotes training-free compression; Selection disables the learned refiner. Δ is measured against the strongest prior TF baseline.
Dynamic Allocation
Cross-Frame Coordination
Answer Loss
KD Loss
VideoAds
Video-MME
LVBench
P.Test
Avg.(13)
✗
✗
–
–
55.45
63.85
39.80
67.85
60.12
✓
✗
–
–
60.30
65.40
42.15
68.30
61.74
✓
✓
–
–
62.55
66.25
42.26
68.65
62.64
✓
✓
✓
✗
61.85
66.80
42.85
68.80
62.67
✓
✓
✗
✓
61.95
67.55
43.60
69.10
62.71
✓
✓
✓
✓
62.10
68.05
44.05
69.25
62.75
Appendix
Table 5: Component and refiner-objective ablations on Qwen3-VL-8B. All variants use 1,535 visual tokens; dashes indicate that no learned refiner is used.
Dimension
Setting
Candidate frames
Token budget
Actual tokens ↓
Avg.(13) ↑
Δ
Selection
Independent
32
1,536
1,534.9
58.50
−0.93
Shuffled correspondence
32
1,536
1,534.9
57.53
−1.90
Coupled (Ours)
32
1,536
1,534.9
59.43
0.00
Candidate frames
16 frames
16
1,536
1,517.7
57.97
−1.46
24 frames
24
1,536
1,530.1
58.85
−0.58
32 frames (Ours)
32
1,536
1,534.9
59.43
0.00
Appendix
Table 6: System-level ablations on Qwen2-VL-7B. We vary the selection strategy, number of candidate frames, and visual-token budget. Δ is measured relative to the default configuration.
Figure 5: Inference latency on Qwen3-VL-32B-Instruct. Selection is included in TTFT. Ours reduces end-to-end latency from 911.2 ms to 421.8 ms relative to Image.
Figure 6: Multi-budget accuracy–latency trade-off. We sweep 0.75K, 1.5K, and 3K visual-token budgets on three backbones. Each curve connects the three operating points of one method.
Figure 7: Representative qualitative successes (I). Four examples illustrate different forms of evidence preservation. In (a), the answer depends on a subtle downward head movement. In (b) and (d), the decision requires preserving interactions among multiple objects across time. In (c), camera motion is visible from the relative displacement of the cup group.
Figure 8: Representative qualitative successes (II). The examples cover late evidence, repeated outcomes, object counting, and directional temporal reasoning. In (a), the decisive event occurs near the end of the clip when the cat approaches and licks the hand. In (b), the answer requires comparing repeated trials and identifying the one with a different outcome. In (c), the model must avoid double-counting repeated views of objects. In (d), temporal order and spatial roles determine the throw direction. Ours predicts the ground-truth option in all four cases.
Figure 9: Representative failure case: long-range subject binding. The question asks what the farther parrot does after spreading its wing near the end of the clip. The full Image pathway predicts the correct answer ( E : continuing to clean itself), whereas Ours, Video/Conv3D, and all compact baselines shown here predict C . The example suggests that aggressive visual compression can make it difficult to preserve the identity of the queried subject over a long temporal span when several visually similar instances are present.
Input: video V , question q , budget B , frozen image encoder, frozen allocator, optional refiner Gθ .
1
Encode every frame independently ; collect (Z,P,D) .
2
Inherit quotas bt from the frozen allocator, with 0≤bt≤Nt and ∑tbt=B .
3
Build question-weighted DPP kernels {Lt} and pair graph (E,w) .
4
S←∅ .
5
while ∣S∣<B do
6
U←FeasibleMoves(S,{bt},E) .
Appendix
Algorithm 1 Image-first encoding with a fixed output budget.
Question weight
Motion weight
Saliency weight
Diversity weight
Frame floor
0.8
0.6
0.4
0.4
12
Preview frames
Preview boost
Uncertainty weight
Temporal NMS gap
DPP ratio
5
0.75
0.45
2
0.5
Appendix
Table 7: Frozen SITE allocator configuration.
Parameter
Value
Parameter
Value
Relevance weight η
0.8
Scene-similarity threshold
0.65
Minimum raw quality
10−6
Confidence margin scale
0.05
Diagonal-jitter coefficient
10−3
Spatial-change scale
0.25
Residual floor
10−14
Feature-change multiplier
2.0
Cosine threshold
0.65
Pair-reward weight γ
2.0
Margin threshold
0.005
Context size K
8
Appendix
Table 8: Selection and refinement implementation constants.