cs.CVOct 6, 2026

Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs

Authors: Donghyun Han, Jangho Park, Yuseok Bae

Organizations: ETRI, South Korea · Kyung Hee University, South Korea

Abstract

Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline. This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved? We introduce Foveated Compression, which encodes a full-resolution image once and represents it with a mixture of native- and compressed-resolution visual tokens. A behaviorally self-distilled Foveated Merger compresses local visual tokens while preserving compatibility with their native counterparts, and a lightweight Foveated Selector chooses one of nine spatial cells to retain at native resolution using exhaustive budget-matched intervention supervision. At 11.11% visual tokens, uniform Foveated Compression shows no significant paired difference from iso-token downsampling. At 20.99%, the learned selector significantly outperforms random and fixed allocation, but remains below strong whole-image resizing, showing that localized fidelity is not universally preferable. A budget-matched region-choice oracle reaches 82.73 macro accuracy versus 69.61 for the learned selector, revealing substantial headroom within the same spatial action space. Matched probing further shows that signals predicting when compression breaks the answer are substantially more accessible after language-model computation than to the lightweight prefill-free selector. These results expose complementary bottlenecks in region selection and compressed-region fidelity.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AsymVLM: Asymmetric Token Pruning for Efficient Vision-Language Model Inference

    May 28, 2026Yilin Feng, Ahmed Burak Gulhan, Mahmut Taylan KandemirVisual Token PruningLarge Language Model Compression

  2. Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression

    Aug 3, 2026Long Qian, Jiaqi Wei, Bingke Zhu +2Large Language Model CompressionLearned Image Compression