cs.CVSep 30, 2026

CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

Authors: Yulong Liu, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Guibo Zhu, Sirui Han, +1 more

Organizations: ERNIE Team, Baidu Inc. · The Hong Kong University of Science and Technology · Institute of Automation, Chinese Academy of Sciences (CASIA)

Abstract

Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each temporal segment is equipped with learnable abstract tokens that learn a compact segment-level representation, while fine-grained patch tokens remain available throughout the encoder. Alternating intra-segment and abstract-communication layers preserve video-level context through the abstract-token channel. A lightweight selector further exposes either abstract tokens alone or abstract tokens augmented with a runtime-selected subset of patch tokens, yielding a compact visual interface that reduces the visual context and prefill burden of downstream MLLMs while retaining fine-grained evidence when needed. Pretrained with contrastive objectives on 565M image--text pairs and 6.4M videos, CoVisco shows competitive performance on video-oriented embedding and multimodal understanding benchmarks. In the evaluated four-segment, 64-frame setting, abstract-only inference uses only 400 visual tokens while achieving video-understanding performance close to, and on some benchmarks exceeding, OneVision-Encoder. Selected patch tokens further improve fine-grained video reasoning. Project URL: https://github.com/ernie-research/CoVisco.git

Figures & tables

Explore similar work

CardsList
  1. VETO: Video Efficient Token Optimization for Vision Language Models

    Oct 1, 2026Gueter Josmy Faure, Hao Ping Wang, Min-Hung Chen +1Token CompressionVideo-Language Models

  2. LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs

    May 17, 2026Jihwan Kim, Nikhil Parthasarathy, Danfeng Qin +5Vision EncodersContext Length

  3. CoViST: Visual Token Compression via Composable States

    Sep 27, 2026Qi Zhang, Xiandong Meng, Ronggang Wang +1Token CompressionData Compression Methods