Foveated Compression: Selective High-Resolution Preservation for Token-Efficient VLMs
Organizations: ETRI, South Korea · Kyung Hee University, South Korea
Abstract
Visual tokens are a major source of inference cost in vision-language models, yet simple image downsampling remains a surprisingly strong compression baseline. This raises a complementary question: under a fixed token budget, where should visual fidelity be preserved? We introduce Foveated Compression, which encodes a full-resolution image once and represents it with a mixture of native- and compressed-resolution visual tokens. A behaviorally self-distilled Foveated Merger compresses local visual tokens while preserving compatibility with their native counterparts, and a lightweight Foveated Selector chooses one of nine spatial cells to retain at native resolution using exhaustive budget-matched intervention supervision. At 11.11% visual tokens, uniform Foveated Compression shows no significant paired difference from iso-token downsampling. At 20.99%, the learned selector significantly outperforms random and fixed allocation, but remains below strong whole-image resizing, showing that localized fidelity is not universally preferable. A budget-matched region-choice oracle reaches 82.73 macro accuracy versus 69.61 for the learned selector, revealing substantial headroom within the same spatial action space. Matched probing further shows that signals predicting when compression breaks the answer are substantially more accessible after language-model computation than to the lightweight prefill-free selector. These results expose complementary bottlenecks in region selection and compressed-region fidelity.
Figures & tables
| Method | Tokens | GQA | MMB-E | MMB-C | MME | POPE | MMStar | OCR | ChartQA | Macro |
|---|---|---|---|---|---|---|---|---|---|---|
| Native | 100.00% | 60.97 | 86.86 | 85.61 | 2356.7 | 87.74 | 60.93 | 80.20 | 84.28 | 78.85 |
| Uniform Foveated Compression | 11.11% | 58.28 | 82.51 | 82.01 | 2130.5 | 84.34 | 50.20 | 49.30 | 34.24 | 64.62 |
| Iso-Token Downsample | 11.11% | 55.88 | 82.77 | 81.75 | 2240.8 | 82.69 | 51.07 | 57.30 | 27.44 | 64.86 |
| Random | 20.99% | 58.40 | 84.18 | 82.86 | 2190.7 | 84.34 | 52.33 | 55.10 | 48.36 | 67.98 |
| Fixed | 20.99% | 58.90 | 83.69 | 82.70 | 2170.3 | 84.54 | 51.87 | 53.70 | 54.92 | 68.48 |
| Foveated Selector | 20.99% | 58.24 | 84.59 | 82.67 | 2186.7 | 84.92 | 52.13 | 58.40 | 57.84 | 69.61 |
| Variant | Tokens | GQA | MMB-E | MMB-C | MME | POPE | MMStar | OCR | ChartQA | Macro |
|---|---|---|---|---|---|---|---|---|---|---|
| Mean-Pooling Compression | 11.11% | 51.76 | 79.65 | 78.33 | 2013.4 | 80.47 | 44.73 | 41.40 | 24.12 | 59.05 |
| Behaviorally Trained Merger | 11.11% | 58.28 | 82.51 | 82.01 | 2130.5 | 84.34 | 50.20 | 49.30 | 34.24 | 64.62 |
| Single-Winner Supervision | 20.99% | 57.47 | 83.51 | 81.75 | 2203.5 | 83.80 | 52.47 | 56.20 | 54.40 | 68.54 |
| Tie-Aware Supervision | 20.99% | 58.38 | 85.08 | 83.04 | 2189.5 | 84.86 | 52.40 | 59.20 | 57.04 | 69.77 |
| Question Content Removed | 20.99% | 58.42 | 84.59 | 82.65 | 2205.7 | 85.01 | 53.00 | 58.90 | 57.08 | 69.80 |
| Full Selector | 20.99% | 58.38 | 85.08 | 83.04 | 2189.5 | 84.86 | 52.40 | 59.20 | 57.04 | 69.77 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Merger | Selector |
| Trainable parameters | 32.77M | 10.8M |
| Learning rate | ||
| Epochs | 1 | 4 |
| Warmup steps | 200 | 100 |
| Optimization steps | 2,378 | – |
| Gradient accumulation | 4 | – |
| Condition | Visual-token ratio |
|---|---|
| Native | 100.00% |
| Uniform Foveated Compression | 11.11% |
| Iso-Token Downsample | 11.11% |
| Random / Fixed / Learned | 20.99% |
| Matched-Budget Downsample | 19.82% |
| Downsample | 24.57% |
| Statistic | Value |
|---|---|
| Total selector labels | 18,048 |
| Spatially informative labels | 7,678 |
| Informative fraction | 42.54% |
| Mean winner-set size | 3.66 |
| Random winner membership | 40.65% |
| Best training-only fixed cell | 45.40% |
| Diagnostic | Value |
|---|---|
| Evaluation examples | 37,610 |
| Overall native-correct retention | 86.1% |
| Lowest retention (ChartQA) | 38.8% |
| Highest retention (POPE) | 94.3% |
| All-nine-tie examples | 78.5% |
| Spatially informative examples | 21.5% |
| Comparison | Macro | 95% CI |
|---|---|---|
| Uniform Foveated Compression Iso-Token Downsample | ||
| Foveated Selector Native | ||
| Foveated Selector Random | ||
| Foveated Selector Fixed | ||
| Foveated Selector Matched-Budget Downsample | ||
| Downsample Foveated Selector |
| Signal | AUC | 95% CI |
|---|---|---|
| Prefill-free raw signal | 0.5239 | |
| First-token entropy | 0.7651 | |
| Prefill-free OOF probe | 0.5476 | |
| Decoder-side OOF probe | 0.8072 |