CoViST: Visual Token Compression via Composable States
Organizations: Peng Cheng Laboratory Shenzhen, China · Peking University Shenzhen Graduate School Peking University Beijing, China
Abstract
Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compressed representation does not explicitly encode how much visual information each representative carries or where it lies in the original image. This limitation arises even after a single reduction and becomes more pronounced when compression is repeated across decoder layers. To address this issue, we propose CoViST, a training-free framework that represents a compressed image as a composable visual state. Specifically, the state combines representative features with original positions, effective contribution weights, and reusable selection metadata. CoViST constructs this state through coverage-guided selection and conservation-based contribution composition, and explicitly incorporates its contribution and positional information into decoder attention. Each component of the state retains its interpretation under successive reductions, enabling the same formulation to support both fixed compression before prefill and progressive compression within the decoder. Experimental results on seven LLaVA-1.5-7B benchmarks show that CoViST-Fixed retains 99.9%, 99.5%, and 98.1% of uncompressed performance at 192, 128, and 64 tokens, respectively, and CoViST-Pro retains 99.8%, 99.9%, and 99.1% at the corresponding layer-average budgets, outperforming state-of-the-art methods under their respective budget settings. Code will be released publicly.
Figures & tables
| Method | Budget | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{v2}}} | (%) | |
|---|---|---|---|---|---|---|---|---|---|---|
| Vanilla, | – | 62.0 | 64.0 | 1867 | 85.8 | 69.5 | 58.2 | 76.7 | 100.0 | |
| DivPrune (CVPR 2025) | fixed | 58.9 | 63.1 | 1723 | 86.5 | 69.0 | 55.7 | 76.1 | 96.9 | |
| PyramidDrop (CVPR 2025) | avg. | 57.3 | 63.3 | 1797 | 84.8 | 69.2 | 56.5 | 76.4 | 97.1 | |
| VisionZip (CVPR 2025) | fixed | 59.3 | 63.0 | 1783 | 85.3 | 68.9 | 57.3 | 76.8 | 97.7 | |
| MMTok (ICLR 2026) | fixed | 60.1 | 63.4 | 1774 | 86.4 | 68.8 | 57.7 | 77.1 | 98.2 | |
| ZOO-Prune (CVPR 2026) | fixed | 60.0 | 62.9 | 1782 | 87.2 | 69.2 | 57.3 | 77.3 | 98.3 |
| Method | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{v2}}} | (%) |
| Vanilla | 63.3 | 68.7 | 1824 | 86.0 | 72.8 | 61.3 | 78.3 | 100.0 |
| ( ) | ||||||||
| VisionZip | 59.1 | 66.9 | 1754 | 85.1 | 73.5 | 59.5 | 78.1 | 97.6 |
| DivPrune | 59.4 | 66.6 | 1782 | 86.8 | 72.9 | 58.5 | 78.0 | 97.8 |
| CDPruner | 60.4 | 67.2 | 1776 | 86.6 | 72.4 | 58.7 | 78.4 | 97.8 |
| ZOO-Prune | 60.0 | 66.7 | 1762 | 86.7 | 73.1 | 59.1 | 78.7 | 98.1 |
| Method | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{v2}}} | (%) |
| Vanilla | 63.3 | 68.7 | 1824 | 86.0 | 72.8 | 61.3 | 78.3 | 100.0 |
| ( ) | ||||||||
| VisionZip | 59.1 | 66.9 | 1754 | 85.1 | 73.5 | 59.5 | 78.1 | 97.6 |
| DivPrune | 59.4 | 66.6 | 1782 | 86.8 | 72.9 | 58.5 | 78.0 | 97.8 |
| CDPruner | 60.4 | 67.2 | 1776 | 86.6 | 72.4 | 58.7 | 78.4 | 97.8 |
| ZOO-Prune | 60.0 | 66.7 | 1762 | 86.7 | 73.1 | 59.1 | 78.7 | 98.1 |
| Method | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{v2}}} | MMB {}^{\scalebox{0.8}{\scriptstyle\mathrm{CN}}} | (%) |
| Vanilla | 59.8 | 83.4 | 2320 | 86.6 | 88.1 | 76.7 | 81.6 | 80.1 | 100.0 |
| ( ) | |||||||||
| VisionZip | 56.6 | 78.9 | 2317 | 85.8 | 80.5 | – | 80.7 | – | 95.8 |
| PruneSID | 59.8 | 80.9 | 2218 | 85.9 | 87.6 | – | 80.4 | – | 97.6 |
| CoViST-Pro | 58.0 | 82.1 | 2315 | 86.4 | 86.2 | 73.1 | 80.1 | 79.6 | 98.2 |
| CoViST-Fixed | 59.2 | 83.0 | 2308 | 85.5 | 87.6 | 76.7 | 80.9 | 80.2 | 99.4 |
| Method | Budget | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{v2}}} | (%) | |
|---|---|---|---|---|---|---|---|---|---|---|
| Vanilla | – | 64.2 | 64.3 | 1817 | 86.9 | 68.0 | 61.1 | 79.9 | 100.0 | |
| DivPrune (CVPR 2025) | fixed | 61.6 | 65.4 | 1773 | 85.5 | 67.8 | 55.4 | 78.9 | 96.1 | |
| PyramidDrop (CVPR 2025) | avg. | 62.9 | 66.5 | 1733 | 86.4 | 69.4 | 58.3 | – | 97.3 | |
| ApET (CVPR 2026) | avg. † | 63.0 | 65.3 | 1815 | 87.2 | – | 57.9 | 79.2 | 97.5 | |
| ZOO-Prune (CVPR 2026) | fixed | 62.2 | 65.2 | 1816 | 86.8 | 68.0 | 58.0 | 79.6 | 97.5 | |
| VisionZip (CVPR 2025) | fixed | 61.3 | 66.3 | 1787 | 87.7 | 68.1 | 60.2 | – | 97.8 |
| Method | TTFT (ms) | Generation (ms) | FLOPs (T) | KV (MiB) | ||
|---|---|---|---|---|---|---|
| Vanilla | – | 263.94 | 289.64 | 43.408 | 1494.82 | |
| CoViST-Fixed | 640 | 178.96 | 207.66 | ( ) | 10.004 | 374.81 |
| CoViST-Fixed | 320 | 125.85 | 153.32 | ( ) | 5.661 | 214.81 |
| CoViST-Fixed | 160 | 102.72 | 130.25 | ( ) | 3.530 | 134.81 |
| CoViST-Pro | 640 | 225.41 | 253.89 | ( ) | 10.058 | 374.81 |
| CoViST-Pro | 320 | 153.56 | 183.57 | ( ) | 5.675 | 214.81 |
| Method | TTFT (ms) | Generation (ms) | FLOPs (T) | KV (MiB) | ||
|---|---|---|---|---|---|---|
| Vanilla | – | 263.94 | 289.64 | 43.408 | 1494.82 | |
| CoViST-Fixed | 640 | 178.96 | 207.66 | ( ) | 10.004 | 374.81 |
| CoViST-Fixed | 320 | 125.85 | 153.32 | ( ) | 5.661 | 214.81 |
| CoViST-Fixed | 160 | 102.72 | 130.25 | ( ) | 3.530 | 134.81 |
| CoViST-Pro | 640 | 225.41 | 253.89 | ( ) | 10.058 | 374.81 |
| CoViST-Pro | 320 | 153.56 | 183.57 | ( ) | 5.675 | 214.81 |
| Configuration | ||||
| (a) | CSI representatives at persistent positions, | 99.28 | 98.81 | 96.35 |
| (b) | + composed under the conservation constraint | 99.74 | 99.18 | 97.10 |
| (c) | + coverage-gap reserve | 99.84 | 99.27 | 97.05 |
| (d) | + closed-form head gains | 99.91 | 99.50 | 97.49 |
| (e) | + composed with confidence gating | 99.95 | 99.47 | 97.59 |
| (f) | + on the full original span (main) | 99.96 | 99.49 | 98.10 |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Value | Role |
|---|---|---|
| 1.5 | evidence exponent, Equation 4 | |
| 0.40 / 0.30 / 0.25 at / 128 / 192, concentration-adaptive, at most 0.60 | uniform component | |
| , | 0.30, 0.10 | instruction residual bound and selectivity scale |
| (first-layer route) | 0.15, without instruction anchors | used when no CLIP text encoder is available |
| 16 regions with a 10% uniform regional weight | region balancing | |
| core fraction | of the budget | protected attention peaks |
| Method | Target | Layers 0–13 | Layers 14–23 | Layers 24–31 | Mean |
|---|---|---|---|---|---|
| CoViST-Pro | 192 | 288 | 173 | 48 | 192.0625 |
| CoViST-Pro | 128 | 192 | 115 | 32 | 127.9375 |
| CoViST-Pro | 64 | 96 | 58 | 16 | 64.1250 |
| CoViST-Fixed |
| Method | GQA | MMB | MME-PC | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | (%) | |
|---|---|---|---|---|---|---|---|---|
| Vanilla | 576 | 62.0 | 64.0 | 1867 | 85.8 | 69.5 | 58.2 | 100.0 |
| CoViST-Fixed | 192 | 61.3 | 64.3 | 1901 | 86.2 | 69.1 | 57.6 | 100.0 |
| CoViST-Fixed | 128 | 61.3 | 64.1 | 1867 | 86.3 | 69.0 | 57.1 | 99.5 |
| CoViST-Fixed | 64 | 60.2 | 63.1 | 1819 | 85.8 | 68.5 | 56.4 | 98.1 |
| CoViST-Pro | 192 | 61.6 | 64.5 | 1855 | 86.2 | 69.3 | 57.9 | 99.8 |
| CoViST-Pro | 128 | 61.3 | 64.7 | 1890 | 86.1 | 69.1 | 57.7 | 100.0 |
| Setting | GQA | MMB | MME-PC | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | (%) | |
| LLaVA-1.5-13B | ||||||||
| Vanilla | 576 | 63.3 | 68.7 | 1824 | 86.0 | 72.8 | 61.3 | 100.0 |
| ( ) | 192 | 62.9 | 68.5 | 1821 | 87.0 | 73.1 | 60.6 | 99.9 |
| ( ) | 128 | 62.7 | 68.0 | 1811 | 87.1 | 73.7 | 60.1 | 99.7 |
| ( ) | 64 | 61.4 | 67.1 | 1775 | 86.3 | 73.0 | 59.5 | 98.3 |
| LLaVA-NeXT-7B | ||||||||
| Setting | MME-P | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} (no OCR) | SEED-I | MMMU |
|---|---|---|---|---|
| LLaVA-NeXT-7B | ||||
| Vanilla | 1496 | 64.6 | 69.9 | 36.3 |
| ( ) | 1531 | 63.1 | 69.0 | 35.2 |
| ( ) | 1482 | 59.4 | 67.7 | 35.3 |
| ( ) | 1449 | 53.0 | 65.2 | 34.8 |
| Qwen2.5-VL-7B | ||||
| Content | Attention | Perception | Cognition | Total | |
|---|---|---|---|---|---|
| 192 | Yes | Yes | 1537 | 364 | 1901 |
| No | Yes | 1534 | 364 | 1898 | |
| Yes | No | 1485 | 349 | 1834 | |
| No | No | 1483 | 351 | 1834 | |
| 288 | Yes | Yes | 1497 | 351 | 1848 |
| No | Yes | 1496 | 351 | 1847 |
| Span | |||
|---|---|---|---|
| 288 | 98.49 | 98.06 | 97.59 |
| 432 | 99.11 | 99.24 | 98.17 |
| 576 (main) | 99.96 | 99.49 | 98.10 |
| Configuration | |||
|---|---|---|---|
| Reference | 100.0 | 99.5 | 98.1 |
| Head gains | 99.9 | 99.5 | 97.8 |
| Head gains | 99.8 | 99.6 | 98.3 |
| Coverage reserve | 99.9 | 99.5 | 98.3 |
| Coverage reserve | 99.9 | 99.5 | 97.9 |
| Evidence exponent | 99.6 | 99.6 | 98.1 |
| First-stage gain | |||
|---|---|---|---|
| 99.9 | 99.9 | 99.2 | |
| 99.8 | 100.0 | 99.2 | |
| 99.7 | 99.9 | 99.3 |
| Method | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | (%) | |
|---|---|---|---|---|---|---|---|---|
| (a) Fixed counts: MME-PC and | ||||||||
| Vanilla, | 62.0 | 64.0 | 1867 | 85.8 | 69.5 | 58.2 | 100.0 | |
| HoloV (NeurIPS 2025) | 58.6 | 62.6 | 1779 | 85.0 | 67.3 | 55.8 | 96.5 | |
| VisionZip R (CVPR 2025) | 59.2 | 62.5 | 1749 | 85.2 | 68.7 | 55.8 | 96.7 | |
| DivPrune (CVPR 2025) | 58.9 | 63.1 | 1723 | 86.5 | 69.0 | 55.7 | 96.8 | |
| ToMe (ICLR 2023) | 59.5 | 62.6 | 1727 | 86.9 | 69.0 | 55.8 | 97.0 | |
| Method | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | (%) |
|---|---|---|---|---|---|---|---|
| (a) Layer schedules: MME-PC and | |||||||
| Vanilla, | 62.0 | 64.0 | 1867 | 85.8 | 69.5 | 58.2 | 100.0 |
| ( ) | |||||||
| SparseVLM (ICML 2025) | 57.6 | 62.5 | 1721 | 83.6 | 69.1 | 56.1 | 95.9 |
| PyramidDrop (CVPR 2025) | 57.3 | 63.3 | 1797 | 84.8 | 69.2 | 56.5 | 97.0 |
| ApET ‡ (CVPR 2026) | 60.2 | 63.4 | 1808 | 86.3 | 68.5 | 54.4 | 97.5 |
| Method | Budget | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | (%) |
|---|---|---|---|---|---|---|---|---|
| PruneSID-Dyn S (ICLR 2026) PS | 192 | 60.2 | 63.8 | 1797 | 87.1 | 69.1 | 56.9 | 98.5 |
| PruneSID-Dyn S (ICLR 2026) PS | 128 | 58.9 | 62.6 | 1760 | 86.9 | 68.8 | 55.1 | 96.9 |
| PruneSID-Dyn S (ICLR 2026) PS | 64 | 57.2 | 59.7 | 1734 | 84.1 | 68.1 | 54.2 | 94.5 |
| OccamToken S (arXiv 2026) OC | 128 | 60.9 | 63.9 | 1825 | 86.3 | 69.1 | 58.0 | 99.1 |
| OccamToken S (arXiv 2026) OC | 64 | 59.3 | 63.2 | 1801 | 86.2 | 69.0 | 57.4 | 98.1 |
| OccamToken S (arXiv 2026) OC | 32 | 59.1 | 62.9 | 1780 | 86.2 | 69.1 | 56.3 | 97.5 |
| Method | Budget | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | (%) |
|---|---|---|---|---|---|---|---|---|
| TOPS ‡ (arXiv 2026) TP | 128 | 60.5 | 62.5 | 1483 | 86.8 | 68.2 | 57.0 | 98.3 |
| TOPS ‡ (arXiv 2026) TP | 64 | 58.7 | 60.9 | 1443 | 86.5 | 68.6 | 56.2 | 96.8 |
| TOPS ‡ (arXiv 2026) TP | 32 | 56.7 | 59.5 | 1385 | 83.5 | 68.8 | 54.9 | 94.3 |
| STAR-Pro (arXiv 2026) SR | 128 | 60.8 | 62.4 | 1456 | 87.3 | 68.9 | 57.2 | 98.4 |
| STAR-Pro (arXiv 2026) SR | 64 | 59.4 | 61.0 | 1421 | 87.5 | 68.6 | 56.3 | 97.0 |
| STAR-Pro (arXiv 2026) SR | 32 | 57.2 | 60.1 | 1358 | 86.6 | 68.7 | 54.6 | 94.8 |
| Code | Reporting paper | Table | Base |
|---|---|---|---|
| R | RESTORE ( Cho et al., 2026 ) | 1 | B1 |
| D | DeSAP ( Ma et al., 2026b ) | 1 | B2 |
| C | CRISP ( Li et al., 2026d ) | 1 | B2 |
| V | VisionTrim ( Yu et al., 2026 ) | 1 | B2 |
| T | CDPruner ( Zhang et al., 2025b ) | 1 | B3 |
| S | SCoRe ( Xu et al., 2026 ) | 1 | B4 |
| Base | MME type | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} |
|---|---|---|---|---|---|---|---|
| B1 | PC | 61.9 | 64.6 | 1862 | 85.9 | 69.5 | 58.2 |
| B2 | PC | 61.9 | 64.7 | 1862 | 85.9 | 69.5 | 58.2 |
| B3 | P | 61.9 | 64.7 | 1507 | 85.9 | 69.5 | 58.2 |
| B4 | P | 62.0 | 64.3 | 1511 | 85.9 | 66.8 | 58.2 |
| B5 | P | 61.9 | 64.7 | 1513 | 85.9 | 69.6 | 58.2 |
| B6 | P | 61.9 | 64.0 | 1509 | 85.8 | 69.6 | 58.3 |
| LLaVA-NeXT-7B: six tasks | ||||||
|---|---|---|---|---|---|---|
| Source | GQA | MMB | MME | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} |
| ZOO-Prune (MME-PC) | 64.2 | 67.9 | 1842 | 86.4 | 70.2 | 61.3 |
| AnchorPrune (MME-P) | 64.2 | 67.2 | 1529 | 86.4 | 70.2 | 61.3 |
| (a) MME-PC and six-task retention | ||||||||
| Method | GQA | MMB | MME-PC | POPE | SQA | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | (%) | |
| Vanilla | 64.2 | 64.3 | 1817 | 86.9 | 68.0 | 61.1 | 100.0 | |
| Full (source) | 64.2 | 67.9 | 1842 | 86.4 | 70.2 | 61.3 | 100.0 | |
| DivPrune (CVPR 2025) | 61.6 | 65.4 | 1773 | 85.5 | 67.8 | 55.4 | 95.7 | |
| ZOO-Prune (CVPR 2026) | 62.2 | 65.2 | 1816 | 86.8 | 68.0 | 58.0 | 97.2 | |
| VisionZip (CVPR 2025) | 61.3 | 66.3 | 1787 | 86.3 | 68.1 | 60.2 | 97.5 | |
| Qwen2.5-VL-7B: MMBench, MME-PC, and TextVQA | |||||
| Method | MMB | MME-PC | VQA {}^{\scalebox{0.8}{\scriptstyle\mathrm{T}}} | (%) | |
| Vanilla | 83.4 | 2320 | 82.4 | 100.0 | |
| HiPrune E (ACL Findings 2026) | 80.3 | 2177 | 75.8 | 93.3 | |
| EADP E (ECCV 2026) | 81.6 | 2213 | 78.6 | 95.5 | |
| DivPrune C (CVPR 2025) | 81.6 | 2279 | 81.8 | 98.0 | |
| CoViST-Fixed | 83.0 | 2308 | 81.5 | 99.3 | |
| (a) General visual reasoning | |||||
|---|---|---|---|---|---|
| Method | GQA | MMB | POPE | SQA | (%) |
| Vanilla | 59.8 | 83.4 | 86.6 | 88.1 | 100.0 |
| Full (source) | 60.8 | 83.8 | 86.3 | 88.2 | 100.0 |
| ( ) | |||||
| VisionZip (CVPR 2025) | 57.0 | 78.6 | 83.2 | 84.5 | 94.9 |
| CoIn (CVPR 2026) | 58.7 | 79.4 | 85.1 | 84.6 | 96.5 |
| Method | Weight | Bound | Head | Position | Inherit | Content |
|---|---|---|---|---|---|---|
| ToMe | patch count | none | shared | n.a. (ViT) | within ViT | average |
| RESTORE | merge count | none | shared | original, distance factor | n.r. | base method |
| ERA | saliency-weighted count | none | shared | n.r. | no | recycled into anchors |
| CaRe | none | – | – | n.r. | no | confidence-gated |
| HiDrop | none | – | – | persistent IDs | no | none (trained) |
| CoViST | evidence-conserved transfer | closed-form gains | original, distance factor | yes | confidence-gated, norm-preserving |
| (a) Measured latency and output rate | |||||||
| Method | Mean visual tokens | TTFT (ms) | LLM prefill (ms) | Generation (ms) | Generation speedup | Output (token/s) | |
| Vanilla | All | 2928 | 263.94 | 245.66 | 289.64 | 6.91 | |
| CoViST-Fixed | 640 | 688 | 178.96 | 150.34 | 207.66 | 9.63 | |
| CoViST-Fixed | 320 | 368 | 125.85 | 97.71 | 153.32 | 13.04 | |
| CoViST-Fixed | 160 | 208 | 102.72 | 74.41 | 130.25 | 15.36 | |
| CoViST-Pro | 640 | 688 | 225.41 | 197.16 | 253.89 | 7.88 | |