Organizations: ERNIE Team, Baidu Inc. · The Hong Kong University of Science and Technology · Institute of Automation, Chinese Academy of Sciences (CASIA)
Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each temporal segment is equipped with learnable abstract tokens that learn a compact segment-level representation, while fine-grained patch tokens remain available throughout the encoder. Alternating intra-segment and abstract-communication layers preserve video-level context through the abstract-token channel. A lightweight selector further exposes either abstract tokens alone or abstract tokens augmented with a runtime-selected subset of patch tokens, yielding a compact visual interface that reduces the visual context and prefill burden of downstream MLLMs while retaining fine-grained evidence when needed. Pretrained with contrastive objectives on 565M image--text pairs and 6.4M videos, CoVisco shows competitive performance on video-oriented embedding and multimodal understanding benchmarks. In the evaluated four-segment, 64-frame setting, abstract-only inference uses only 400 visual tokens while achieving video-understanding performance close to, and on some benchmarks exceeding, OneVision-Encoder. Selected patch tokens further improve fine-grained video reasoning. Project URL: https://github.com/ernie-research/CoVisco.git
Figures & tables
Figure 1 : Framework of CoVisco. (1) The input video is divided into temporal segments and converted into visual tokens using codec-based, uniform, or collage sampling. (2) Inside the ViT, alternating attention blocks perform full attention within each segment and abstract-mediated communication across segments. The central attention maps make this distinction explicit. Each segment contains learnable abstract tokens and fine-grained patch tokens. (3) The ViT outputs both token types: abstract tokens are pooled and supervised by image–text, image–image, and video–text contrastive objectives, whereas fine-grained tokens are preserved throughout the encoder and remain available to a lightweight selector. At inference, the two outputs support abstract-only, abstract-plus-top- K , and full-token modes.
ImageNet-1k
COCO
Flickr
XM3600
ViT
Res.
Seq.
Model
val
v2
ReaL
ObjNet
T → I
I → T
T → I
I → T
T → I
I → T
L/14
224
256
OpenCLIP [ 7 ]
74.0
61.1
–
66.4
46.1
62.1
75.0
88.7
–
–
CLIP [ 21 ]
75.5
69.0
–
69.9
36.5
56.3
65.2
85.2
–
–
MetaCLIP [ 29 ]
79.2
72.6
–
74.6
55.7
–
83.3
–
–
–
CLIPA-v2 [ 15 ]
79.7
72.8
–
71.1
46.3
64.1
73.0
89.1
–
–
EVA-CLIP [ 23 ]
79.8
72.9
–
75.3
47.5
63.7
77.3
89.7
–
–
Table 1 : Zero-shot image classification and image-text retrieval reference. Results are adapted from the SigLIP 2 zero-shot comparison [ 25 ] ; only the ViT-L/14 and ViT-L/16 blocks are retained. Best values within each block are in bold formatting.
Model
Vision Tower
Image
Video
VisDoc
CLS
CLS
RET
VDRv1
VDRv2
VR
# of Datasets →
10
5
5
10
4
6
VLM2Vec [ 10 ]
2B
58.7
33.4
20.6
49.8
13.5
51.8
VLM2Vec-V2 [ 20 ]
2B
62.9
39.3
28.8
75.5
44.9
79.4
GME [ 33 ]
2B
54.4
34.9
25.6
86.1
54.0
82.5
Ops-MM-embedding-v1 †
2B
68.1
53.6
41.8
76.4
53.2
77.6
Table 2 : Results on the MMEB-V2 benchmark [ 20 ] . CLS: classification, RET: retrieval, VDR: ViDoRe, VR: VisRAG. † : link to the model’s homepage.
Model / setting
MVBench
MLVU-dev
NExT-QA (MC)
VideoMME
Perception Test
TOMATO
LongVideoBench-Val-V
OV-Encoder (Codec,504,1.0) [ 24 ]
52.4
46.3
75.6
53.4
60.3
22.2
50.4
OV-Encoder-Frame (Frame,504,1.0) [ 24 ]
49.8
49.4
71.9
49.3
56.7
21.8
45.5
SigLIP2 (Frame,512,1.0) [ 25 ]
47.2
48.4
70.6
46.8
56.0
22.3
45.2
CoVisco (Codec,224,0.0)
55.8
57.8
72.8
52.9
58.7
25.1
48.0
CoVisco (Codec,224,0.4)
56.7
58.5
74.0
55.7
59.7
26.0
50.5
CoVisco (Codec,504,0.0)
53.2
56.7
69.9
51.3
56.1
25.1
47.3
Table 3 : Video benchmark comparison under a fixed language backbone. All models are evaluated using Qwen3-4B-Instruct-2507. OneVision-Encoder uses a 10,368-token visual budget in both its codec and uniform-frame settings; these visual tokens enter its vision encoder and are forwarded to the language model. In the codec setting, the tokens are selected from a 64-frame input using codec scores, whereas the frame setting uses 8 uniformly sampled frames at 504×504 . SigLIP2 uses 8 uniformly sampled frames at 512×512 . For CoVisco, Codec denotes codec-guided input selection and Frame denotes uniform frame sampling. The notation is (sampling mode, input resolution, fine-grained token retention ratio for LLM); “–” indicates that the corresponding setting was not evaluated. Bold values indicate the best performance among all rows shown for each benchmark.
Model / setting
AI2D
ChartQA
DocVQA
InfoVQA
MMBench-EN
OCRBench
OCRBench v2
MMStar
RealWorldQA
OV-Encoder (Codec,native,1.0) [ 24 ]
75.7
76.5
78.4
43.1
77.2
605
26.3
52.1
60.8
OV-Encoder-Frame (Frame,native,1.0) [ 24 ]
76.5
77.8
79.5
45.5
78.5
630
26.1
54.3
61.2
SigLIP2 (Frame,native,1.0) [ 25 ]
78.6
76.4
75.0
42.0
79.6
621
26.1
55.0
62.1
CoVisco (Frame,native,0.0)
66.8
17.4
20.5
20.7
68.2
261
21.2
43.8
45.5
CoVisco (Frame,native,0.2)
67.9
51.8
64.0
32.1
70.9
474
24.0
45.7
52.8
CoVisco (Frame,native,0.8)
69.8
68.8
75.1
42.2
71.8
567
24.0
48.8
54.1
Table 4 : Image benchmark comparison under a fixed language backbone. All models are evaluated using Qwen3-4B-Instruct-2507. Codec denotes codec-guided input selection, whereas Frame denotes uniform frame sampling. Each setting is written as (sampling mode, input resolution, fine-grained token retention ratio for LLM); “–” indicates that the corresponding setting was not evaluated. Bold values indicate the best performance among all rows shown for each benchmark.
Figure 2 : Qualitative visualization of the token selector on images and videos. Yellow boxes mark fine-grained patch tokens retained by the selector. The selected regions typically focus on objects, text, and localized changes, illustrating how selected patch tokens complement the abstract-token summaries.
Benchmark
Resolution
Frames
Segments
Sampling
Token strategy
Accuracy (%)
MLVU
224
64
4
Frame
Abstract-only
61.2
MLVU
224
64
4
Codec
Abstract-only
57.8
MLVU
224
128
8
Frame
Abstract-only
62.6
MLVU
224
128
16
Frame
Abstract-only
62.3
MLVU
224
256
16
Frame
Abstract-only
61.3
MLVU
224
256
16
Codec
Abstract-only
59.9
Table 5 : Long-video understanding under different segment configurations. The language model is SFT-tuned with four segments; evaluation changes the number of segments at inference time without retraining. All configurations use the abstract-only mode. The abstract-token budget is S×Q with Q=100 .