cs.AISep 27, 2026

CoViST: Visual Token Compression via Composable States

Authors: Qi Zhang, Xiandong Meng, Ronggang Wang, Siwei Ma

Organizations: Peng Cheng Laboratory Shenzhen, China · Peking University Shenzhen Graduate School Peking University Beijing, China

Abstract

Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compressed representation does not explicitly encode how much visual information each representative carries or where it lies in the original image. This limitation arises even after a single reduction and becomes more pronounced when compression is repeated across decoder layers. To address this issue, we propose CoViST, a training-free framework that represents a compressed image as a composable visual state. Specifically, the state combines representative features with original positions, effective contribution weights, and reusable selection metadata. CoViST constructs this state through coverage-guided selection and conservation-based contribution composition, and explicitly incorporates its contribution and positional information into decoder attention. Each component of the state retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder. Experimental results on seven LLaVA-1.5-7B benchmarks show that CoViST-Fixed retains 99.9%, 99.5%, and 98.1% of uncompressed performance at 192, 128, and 64 tokens, respectively, and CoViST-Pro retains 99.8%, 99.9%, and 99.1% at the corresponding layer-average budgets, outperforming state-of-the-art methods under their respective budget settings. Code will be released publicly.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression

    Jul 14, 2026Yupeng Zheng, Kai Zou, Bin Liu +1Token CompressionData Compression Methods

  2. Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression

    Aug 3, 2026Long Qian, Jiaqi Wei, Bingke Zhu +2Large Language Model CompressionLearned Image Compression

  3. CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding

    Sep 30, 2026Yulong Liu, Xiaotian Han, Junyuan Shang +6Vision EncodersToken Compression