Organizations: ERNIE Team, Baidu Inc. · The Hong Kong University of Science and Technology · Institute of Automation, Chinese Academy of Sciences (CASIA)
Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each temporal segment is equipped with learnable abstract tokens that learn a compact segment-level representation, while fine-grained patch tokens remain available throughout the encoder. Alternating intra-segment and abstract-communication layers preserve video-level context through the abstract-token channel. A lightweight selector further exposes either abstract tokens alone or abstract tokens augmented with a runtime-selected subset of patch tokens, yielding a compact visual interface that reduces the visual context and prefill burden of downstream MLLMs while retaining fine-grained evidence when needed. Pretrained with contrastive objectives on 565M image--text pairs and 6.4M videos, CoVisco shows competitive performance on video-oriented embedding and multimodal understanding benchmarks. In the evaluated four-segment, 64-frame setting, abstract-only inference uses only 400 visual tokens while achieving video-understanding performance close to, and on some benchmarks exceeding, OneVision-Encoder. Selected patch tokens further improve fine-grained video reasoning. Project URL: https://github.com/ernie-research/CoVisco.git
Figures & tables
Figure 1 : Framework of CoVisco. (1) The input video is divided into temporal segments and converted into visual tokens using codec-based, uniform, or collage sampling. (2) Inside the ViT, alternating attention blocks perform full attention within each segment and abstract-mediated communication across segments. The central attention maps make this distinction explicit. Each segment contains learnable abstract tokens and fine-grained patch tokens. (3) The ViT outputs both token types: abstract tokens are pooled and supervised by image–text, image–image, and video–text contrastive objectives, whereas fine-grained tokens are preserved throughout the encoder and remain available to a lightweight selector. At inference, the two outputs support abstract-only, abstract-plus-top- K , and full-token modes.
ImageNet-1k
COCO
Flickr
XM3600
ViT
Res.
Seq.
Model
val
v2
ReaL
ObjNet
T → I
I → T
T → I
I → T
T → I
I → T
L/14
224
256
OpenCLIP [ 7 ]
74.0
61.1
–
66.4
46.1
62.1
75.0
88.7
–
–
CLIP [ 21 ]
75.5
69.0
–
69.9
36.5
56.3
65.2
85.2
–
–
MetaCLIP [ 29 ]
79.2
72.6
–
74.6
55.7
–
83.3
–
–
–
CLIPA-v2 [ 15 ]
79.7
72.8
–
71.1
46.3
64.1
73.0
89.1
–
–
EVA-CLIP [ 23 ]
79.8
72.9
–
75.3
47.5
63.7
77.3
89.7
–
–
Table 1 : Zero-shot image classification and image-text retrieval reference. Results are adapted from the SigLIP 2 zero-shot comparison [ 25 ] ; only the ViT-L/14 and ViT-L/16 blocks are retained. Best values within each block are in bold formatting.
Model
Vision Tower
Image
Video
VisDoc
CLS
CLS
RET
VDRv1
VDRv2
VR
# of Datasets →
10
5
5
10
4
6
VLM2Vec [ 10 ]
2B
58.7
33.4
20.6
49.8
13.5
51.8
VLM2Vec-V2 [ 20 ]
2B
62.9
39.3
28.8
75.5
44.9
79.4
GME [ 33 ]
2B
54.4
34.9
25.6
86.1
54.0
82.5
Ops-MM-embedding-v1 †
2B
68.1
53.6
41.8
76.4
53.2
77.6
Table 2 : Results on the MMEB-V2 benchmark [ 20 ] . CLS: classification, RET: retrieval, VDR: ViDoRe, VR: VisRAG. † : link to the model’s homepage.
Model / setting
MVBench
MLVU-dev
NExT-QA (MC)
VideoMME
Perception Test
TOMATO
LongVideoBench-Val-V
OV-Encoder (Codec,504,1.0) [ 24 ]
52.4
46.3
75.6
53.4
60.3
22.2
50.4
OV-Encoder-Frame (Frame,504,1.0) [ 24 ]
49.8
49.4
71.9
49.3
56.7
21.8
45.5
SigLIP2 (Frame,512,1.0) [ 25 ]
47.2
48.4
70.6
46.8
56.0
22.3
45.2
CoVisco (Codec,224,0.0)
55.8
57.8
72.8
52.9
58.7
25.1
48.0
CoVisco (Codec,224,0.4)
56.7
58.5
74.0
55.7
59.7
26.0
50.5
CoVisco (Codec,504,0.0)
53.2
56.7
69.9
51.3
56.1
25.1
47.3
Table 3 : Video benchmark comparison under a fixed language backbone. All models are evaluated using Qwen3-4B-Instruct-2507. OneVision-Encoder uses a 10,368-token visual budget in both its codec and uniform-frame settings; these visual tokens enter its vision encoder and are forwarded to the language model. In the codec setting, the tokens are selected from a 64-frame input using codec scores, whereas the frame setting uses 8 uniformly sampled frames at 504×504 . SigLIP2 uses 8 uniformly sampled frames at 512×512 . For CoVisco, Codec denotes codec-guided input selection and Frame denotes uniform frame sampling. The notation is (sampling mode, input resolution, fine-grained token retention ratio for LLM); “–” indicates that the corresponding setting was not evaluated. Bold values indicate the best performance among all rows shown for each benchmark.
Model / setting
AI2D
ChartQA
DocVQA
InfoVQA
MMBench-EN
OCRBench
OCRBench v2
MMStar
RealWorldQA
OV-Encoder (Codec,native,1.0) [ 24 ]
75.7
76.5
78.4
43.1
77.2
605
26.3
52.1
60.8
OV-Encoder-Frame (Frame,native,1.0) [ 24 ]
76.5
77.8
79.5
45.5
78.5
630
26.1
54.3
61.2
SigLIP2 (Frame,native,1.0) [ 25 ]
78.6
76.4
75.0
42.0
79.6
621
26.1
55.0
62.1
CoVisco (Frame,native,0.0)
66.8
17.4
20.5
20.7
68.2
261
21.2
43.8
45.5
CoVisco (Frame,native,0.2)
67.9
51.8
64.0
32.1
70.9
474
24.0
45.7
52.8
CoVisco (Frame,native,0.8)
69.8
68.8
75.1
42.2
71.8
567
24.0
48.8
54.1
Table 4 : Image benchmark comparison under a fixed language backbone. All models are evaluated using Qwen3-4B-Instruct-2507. Codec denotes codec-guided input selection, whereas Frame denotes uniform frame sampling. Each setting is written as (sampling mode, input resolution, fine-grained token retention ratio for LLM); “–” indicates that the corresponding setting was not evaluated. Bold values indicate the best performance among all rows shown for each benchmark.
Figure 2 : Qualitative visualization of the token selector on images and videos. Yellow boxes mark fine-grained patch tokens retained by the selector. The selected regions typically focus on objects, text, and localized changes, illustrating how selected patch tokens complement the abstract-token summaries.
Benchmark
Resolution
Frames
Segments
Sampling
Token strategy
Accuracy (%)
MLVU
224
64
4
Frame
Abstract-only
61.2
MLVU
224
64
4
Codec
Abstract-only
57.8
MLVU
224
128
8
Frame
Abstract-only
62.6
MLVU
224
128
16
Frame
Abstract-only
62.3
MLVU
224
256
16
Frame
Abstract-only
61.3
MLVU
224
256
16
Codec
Abstract-only
59.9
Table 5 : Long-video understanding under different segment configurations. The language model is SFT-tuned with four segments; evaluation changes the number of segments at inference time without retraining. All configurations use the abstract-only mode. The abstract-token budget is S×Q with Q=100 .
Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates this bottleneck through dual-axis compression: (i) an intra-frame compressor that merges semantically similar tokens within each frame via optimal-transport inspired matching, and (ii) an inter-frame compressor that identifies and merges temporally redundant frames. The key design insight is hierarchical ordering: by first compressing spatial dimensions, VETO drastically reduces the cost of subsequent global temporal matching, bypassing the efficiency wall of single-axis approaches, with an advantage that grows with modern fully-fused attention infrastructure. Empirically, VETO achieves up to 45% faster inference (e.g., on LLaVA-OneVision-7B) while preserving or improving accuracy. Under extreme token starvation (10% budget), VETO outperforms VFlowOpt (54.9%), VisionZip (52.6%), and FastV (47.9%) with 55.7% accuracy. We demonstrate universal applicability across LLaVA-OneVision, InternVL-2.5, and LongVA, with zero-shot accuracy preserved or improved in all cases.
The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long-form video lies in managing the explosion of visual-token context length. Existing strategies predominantly focus on "post-hoc" token reduction -- reducing visual tokens after feature extraction to alleviate the LLM's computational overhead. While these methods effectively reduce the number of visual tokens, we observe that the primary latency bottleneck then shifts from the LLM to the expensive per-frame processing of the vision encoder. To address this, we introduce LiteFrame, a strong, yet highly efficient video encoder backbone for Video LLMs. To train LiteFrame, we propose Compressed Token Distillation (CTD), a novel training framework that teaches a compact student vision encoder to directly predict information-dense, spatio-temporally compressed representations produced by a large teacher vision model, effectively bypassing redundant computation. When coupled with further Language Model Adaptation (LMA), this approach results in a new latency-accuracy Pareto frontier -- compared with InternVL3-8B, LiteFrame provides a 35% reduction in end-to-end latency while processing 8× more frames and improves average video understanding accuracy across multiple benchmarks. Our results demonstrate a new potential path to unlocking longer-form video understanding under fixed compute budgets.
Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compressed representation does not explicitly encode how much visual information each representative carries or where it lies in the original image. This limitation arises even after a single reduction and becomes more pronounced when compression is repeated across decoder layers. To address this issue, we propose CoViST, a training-free framework that represents a compressed image as a composable visual state. Specifically, the state combines representative features with original positions, effective contribution weights, and reusable selection metadata. CoViST constructs this state through coverage-guided selection and conservation-based contribution composition, and explicitly incorporates its contribution and positional information into decoder attention. Each component of the state retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder. Experimental results on seven LLaVA-1.5-7B benchmarks show that CoViST-Fixed retains 99.9%, 99.5%, and 98.1% of uncompressed performance at 192, 128, and 64 tokens, respectively, and CoViST-Pro retains 99.8%, 99.9%, and 99.1% at the corresponding layer-average budgets, outperforming state-of-the-art methods under their respective budget settings. Code will be released publicly.
Qi Zhang, Xiandong Meng, Ronggang Wang +1
Peng Cheng Laboratory Shenzhen, China · Peking University Shenzhen Graduate School Peking University Beijing, China