MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference
Organizations: State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · University of Electronic Science and Technology of China · Nanyang Technological University · The University of Hong Kong · New York University
Abstract
Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visual tokens results in high computational costs. While many methods have been proposed to reduce the number of visual tokens, most of them rely on heuristics and are prone to discarding substantial visual information during pruning, leading to degradation in model performance. In this work, by using a semantic erasure model, we derive a general mutual information coverage objective from task log-loss and propose MiCo, a training-free two-stage pruning method. MiCo first uses visual signals to select a representative candidate pool before visual tokens enter the language model, then performs task-aware subset selection within it. At each stage, suitable observable proxies instantiate the derived objective as a monotone submodular coverage function, which MiCo greedily optimizes under the token budget. MiCo is evaluated on diverse MLLMs ranging from 7B to 13B parameters across a broad range of image and video benchmarks spanning general visual reasoning, fine-grained OCR and grounding, hallucination detection, and long-video understanding. MiCo consistently achieves the best performance across nearly all evaluated models under all pruning ratios. On LLaVA-NEXT-13B, MiCo uses only 5.6% visual tokens, retains 97.5% of baseline performance, and achieves a 3.8-fold inference speedup. Our experiments demonstrate the effectiveness of MiCo and our mutual information coverage objective for visual token pruning.
Figures & tables
| Method | LLaVA | Qwen-VL | InternVL | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | 1.5-7B | 1.5-13B | NeXT-7B | 2.5-VL-7B | 3-VL-8B | 3.5-9B | 3-8B | ||||||||||
| Budget | 128 | 64 | 32 | 128 | 64 | 32 | 640 | 320 | 160 | 256 | 128 | 256 | 128 | 256 | 128 | 256 | 128 |
| FastV | 92.6 | 75.5 | – | 94.1 | 85.0 | – | 95.4 | 80.4 | – | 93.0 | 77.9 | 84.3 | 63.7 | 84.0 | 62.8 | 95.5 | 81.5 |
| DivPrune | 96.8 | 94.4 | 91.3 | 96.7 | 93.2 | 91.2 | 99.0 | 97.2 | 94.2 | 94.5 | 90.3 | 93.6 | 88.6 | 95.9 | 91.8 | 93.3 | 87.7 |
| VisPruner | 96.6 | 93.3 | 87.8 | 96.6 | 93.6 | 88.5 | 98.7 | 95.6 | 91.0 | 91.8 | 88.0 | 94.2 | 86.8 | 97.0 | 90.1 | 89.4 | 81.1 |
| VScan | 97.8 | 96.0 | 91.8 | 97.3 | 96.0 | 88.7 | 99.5 | 96.2 | 90.3 | 94.8 | 90.6 | 95.4 | 90.1 | – | – | 92.9 | 87.3 |
| Qwen2.5-VL-7B | Qwen3-VL-8B | |||||
| Method | RefCOCO | RefCOCO+ | RefCOCOg | RefCOCO | RefCOCO+ | RefCOCOg |
| Vanilla | 86.8 (100.0%) | 79.4 (100.0%) | 85.5 (100.0%) | 92.4 (100.0%) | 88.7 (100.0%) | 90.4 (100.0%) |
| Retain 256 tokens (80.2% pruned) | ||||||
| FastV | 33.0 (38.0%) | 28.3 (35.6%) | 35.5 (41.5%) | 67.2 (72.7%) | 59.0 (66.5%) | 68.1 (75.3%) |
| VisPruner | 22.3 (25.7%) | 17.8 (22.4%) | 21.7 (25.4%) | 77.8 (84.2%) | 71.7 (80.8%) | 77.0 (85.2%) |
| ApET | 44.0 (50.7%) | 37.1 (46.7%) | 42.6 (49.8%) | 51.4 (55.6%) | 47.1 (53.1%) | 49.7 (55.0%) |
| Qwen2.5-VL-7B | Qwen3-VL-8B | |||||||||||
| Method | Text | Chart | Doc | Text | Chart | Doc | Text | Chart | Doc | Text | Chart | Doc |
| Vanilla | 85.06 | 87.28 | 94.86 | 85.06 | 87.28 | 94.86 | 84.28 | 83.04 | 95.72 | 84.28 | 83.04 | 95.72 |
| FastV | 78.44 | 68.28 | 38.22 | 46.45 | 33.20 | 14.54 | 62.18 | 28.64 | 14.79 | 32.88 | 16.24 | 9.83 |
| VisPruner | 64.02 | 56.08 | 27.66 | 49.40 | 36.68 | 19.69 | 71.96 | 63.52 | 39.63 | 54.58 | 39.64 | 23.00 |
| ApET | 69.17 | 50.60 | 23.12 | 45.82 | 31.16 | 14.35 | 59.54 | 54.52 | 39.25 | 41.27 | 35.44 | 25.59 |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Energy | Reliability | Task relevance |
|---|---|---|---|
| Stage 1 | |||
| Stage 2 |
| Backbone | Decoder blocks | Refinement | Stage-1 keep | Stage-2 keep |
|---|---|---|---|---|
| LLaVA-1.5/NeXT-7B | 32 | 7 | ||
| LLaVA-1.5-13B | 40 | 8 | ||
| LLaVA-NeXT-13B | 40 | 14 | ||
| Qwen2.5-VL-7B | 28 | 2 | ||
| InternVL3-8B | 28 | 4 | ||
| Qwen3-VL-8B | 36 | 6 |
| Method | LLaVA-NeXT-7B | LLaVA-NeXT-13B | ||||
|---|---|---|---|---|---|---|
| Latency (ms) (Speedup) | Peak memory (GiB) (Reduction) | Logical FLOPs (T) (Reduction) | Latency (ms) (Speedup) | Peak memory (GiB) (Reduction) | Logical FLOPs (T) (Reduction) | |
| Vanilla | 290.0 (1.00 ) | 15.8 (0.0%) | 45.6 (0.0%) | 484.1 (1.00 ) | 28.4 (0.0%) | 85.1 (0.0%) |
| Budget tokens per image | ||||||
| FastV | 113.0 (2.57 ) | 16.0 (-1.7%) | 11.8 (74.1%) | 154.9 (3.13 ) | 28.1 (1.1%) | 20.9 (75.4%) |
| SparseVLM | 137.1 (2.11 ) | 16.0 (-1.7%) | 11.8 (74.2%) | 176.3 (2.75 ) | 28.1 (1.1%) | 21.0 (75.3%) |
| DivPrune | 108.7 (2.67 ) | 14.2 (9.7%) | 11.7 (74.4%) | 154.2 (3.14 ) | 25.9 (8.8%) | 20.8 (75.6%) |
| Component (GFLOPs) | |||
|---|---|---|---|
| Vision encoder | 1836.289 (36.63%) | 1836.289 (25.57%) | 1836.289 (15.78%) |
| Multimodal projector | 120.879 (2.41%) | 120.879 (1.68%) | 120.879 (1.04%) |
| Decoder layers and final normalization | 3002.246 (59.90%) | 5136.922 (71.54%) | 9520.805 (81.81%) |
| Vocabulary output head | 48.350 (0.96%) | 78.496 (1.09%) | 139.052 (1.19%) |
| Stage 1: scores, similarity and selection | 3.729 (0.07%) | 4.047 (0.06%) | 4.684 (0.04%) |
| Stage 2: extra attention / relevance | 0.113 (0.002%) | 0.206 (0.003%) | 0.391 (0.003%) |
| Component (GFLOPs) | |||
|---|---|---|---|
| Vision encoder | 41228.062 (69.53%) | 41228.062 (54.73%) | 41228.062 (37.19%) |
| Multimodal projector | 1585.032 (2.67%) | 1585.032 (2.10%) | 1585.032 (1.43%) |
| Visual pooling / plumbing | 0.155 (0.000%) | 0.155 (0.000%) | 0.155 (0.000%) |
| Decoder layers and final normalization | 15444.005 (26.05%) | 30518.109 (40.51%) | 63934.284 (57.68%) |
| Vocabulary output head | 929.368 (1.57%) | 1738.100 (2.31%) | 3357.739 (3.03%) |
| Pruning operators (Stage 1 and Stage 2, jointly counted) | 109.016 (0.18%) | 261.046 (0.35%) | 745.496 (0.67%) |
| Method / relevance | AI2D | POPE | HallB | MME | MMB EN | MMB CN | MMStar | SQA IMG | Acc | Rel | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| MMTok (ICLR’26) | 128 | 76.3 | 84.2 | 43.3 | 2120.7 | 79.8 | 77.5 | 53.7 | 81.4 | 75.3 | 90.8% |
| uniform | 128 | 78.2 | 85.2 | 42.6 | 2149.09 | 80.2 | 78.9 | 56.3 | 83.8 | 76.6 | 92.3% |
| text-vision-sim | 128 | 79.1 | 85.5 | 45.9 | 2142.01 | 79.7 | 79 | 56.5 | 84.3 | 77.1 | 93.0% |
| 1-(text-vision-sim) | 128 | 79.4 | 86.2 | 45.9 | 2170.42 | 80.3 | 78.7 | 55.5 | 83.4 | 77.2 | 93.1% |
| last-token-attn | 128 | 80.5 | 85.6 | 48.3 | 2235.97 | 80.6 | 79 | 56.1 | 84.3 | 78.3 | 94.4% |
| MiCo (all-text-attn) | 128 | 80.4 | 85.5 | 48.6 | 2252.37 | 80.8 | 79.1 | 56.5 | 84.6 | 78.5 | 94.7% |
| Model | Refinement-layer configuration |
|---|---|
| LLaVA-1.5-7B | |
| LLaVA-1.5-13B | |
| LLaVA-NeXT-7B | |
| LLaVA-NeXT-13B | |
| Qwen3-VL-8B | |
| InternVL3-8B |
| Method | GQA | SQA-IMG | TextVQA | POPE | MME | MMB-EN | MMB-CN | SEED | AI2D | MMMU | Acc | Rel |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LLaVA-1.5-7B | ||||||||||||
| Upper Bound — 576 tokens (no pruning) | ||||||||||||
| Vanilla | 61.9 | 69.5 | 58.2 | 85.9 | 1508.8 | 64.7 | 58.1 | 66.0 | 55.5 | 35.0 | 63.0 | 100.0% |
| Retain 128 tokens ( 77.8%) | ||||||||||||
| FastV (ECCV’24) | 54.0 | 69.2 | 56.4 | 68.2 | 1376.5 | 63.0 | 55.9 | 59.7 | 53.6 | 34.3 | 58.3 | 92.6% |
| SparseVLM (ICML’25) | 57.3 | 69.0 | 56.3 | 83.1 | 1401.3 | 62.6 | 56.9 | 61.7 | 54.7 | 35.9 | 60.8 | 96.4% |
| Method | AI2D | POPE | HallB | MME | MMB-EN | MMB-CN | MMStar | SQA-IMG | Acc | Rel |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-VL-7B | ||||||||||
| Vanilla | 84.9 | 87.7 | 55.9 | 2302.0 | 84.8 | 82.9 | 65.5 | 86.8 | 83.0 | 100.0% |
| Retain 256 tokens (80.2% pruned) | ||||||||||
| FastV (ECCV’24) | 78.4 | 83.0 | 49.1 | 2169.0 | 80.5 | 78.8 | 55.5 | 83.6 | 77.2 | 93.0% |
| SparseVLM (ICML’25) | 77.6 | 82.9 | 46.4 | 2207.5 | 80.9 | 79.4 | 55.5 | 87.7 | 77.6 | 93.5% |
| DivPrune (CVPR’25) | 81.2 | 85.3 | 46.6 | 2167.0 | 81.8 | 80.9 | 57.9 | 84.8 | 78.4 | 94.5% |
| GQA | SQA IMG | TextVQA | POPE | MME | MMB EN | MMB CN | SEED | AI2D | MMMU | Acc | Rel | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LLaVA-1.5-7B | ||||||||||||
| Vanilla | 61.9 | 69.5 | 58.2 | 85.9 | 1508.8 | 64.7 | 58.1 | 66.0 | 55.5 | 35.0 | 63.0 | 100.0% |
| 1 | 56.1 | 69.0 | 54.3 | 80.4 | 1295.0 | 58.9 | 52.2 | 57.5 | 52.7 | 33.3 | 57.9 | 91.9% |
| 2 | 55.6 | 68.6 | 54.7 | 79.7 | 1298.2 | 58.4 | 52.2 | 57.8 | 53.3 | 34.4 | 58.0 | 92.0% |
| 3 | 55.6 | 68.8 | 54.9 | 79.0 | 1336.5 | 59.7 | 53.0 | 57.9 | 53.0 | 34.2 | 58.3 | 92.5% |
| 4 | 55.5 | 68.8 | 54.6 | 78.7 | 1323.3 | 59.5 | 52.7 | 57.9 | 53.0 | 35.2 | 58.2 | 92.4% |
| GQA | SQA IMG | TextVQA | POPE | MME | MMB EN | MMB CN | SEED | AI2D | MMMU | Acc | Rel | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LLaVA-NeXT-7B | ||||||||||||
| Vanilla | 62.5 | 67.5 | 60.3 | 86.8 | 1511.8 | 65.8 | 57.3 | 69.7 | 64.7 | 35.2 | 64.5 | 100.0% |
| 1 | 59.4 | 68.0 | 54.3 | 82.3 | 1379.2 | 62.4 | 54.3 | 63.4 | 62.8 | 36.8 | 61.3 | 94.9% |
| 2 | 59.5 | 67.3 | 55.9 | 81.9 | 1388.0 | 62.6 | 54.9 | 63.4 | 63.0 | 37.0 | 61.5 | 95.3% |
| 3 | 58.8 | 67.7 | 56.3 | 81.0 | 1442.7 | 62.8 | 55.6 | 63.5 | 62.6 | 35.3 | 61.6 | 95.4% |
| 4 | 59.1 | 68.5 | 56.1 | 81.5 | 1419.2 | 62.3 | 55.1 | 63.5 | 62.9 | 35.3 | 61.5 | 95.3% |
| Method | Chart QA | GQA | MMB CN | MMB EN | MME | MMStar | OCR Bench | POPE | Acc | Rel |
|---|---|---|---|---|---|---|---|---|---|---|
| Reference image-area reduction: 75.00% | ||||||||||
| Vanilla | 88.5 (94.3) | 65.0 (88.2) | 67.4 (94.7) | 76.5 (97.9) | 70.4 (93.9) | 75.4 (86.7) | 83.8 (90.4) | 81.0 (96.7) | 76.0 (92.8) | 100.0 (100.0) |
| FastV | 57.8 ( 83.7 ) | 54.7 (77.3) | 58.7 (93.5) | 65.9 (95.9) | 62.0 (91.8) | 55.2 (76.2) | 43.0 (64.8) | 65.8 (91.5) | 57.9 (84.3) | 76.1 (90.8) |
| SparseVLM | 26.2 (51.5) | 50.4 (78.8) | 59.6 (93.7) | 71.0 (95.4) | 60.6 (92.5) | 53.0 (73.0) | 19.0 (44.3) | 70.2 (90.8) | 51.2 (77.5) | 67.4 (83.5) |
| DivPrune | 57.4 (78.0) | 60.3 ( 86.1 ) | 66.5 (93.9) | 68.7 (96.6) | 66.2 (92.4) | 59.6 ( 79.8 ) | 53.1 ( 72.3 ) | 75.4 ( 95.7 ) | 63.4 ( 86.8 ) | 83.4 ( 93.5 ) |
| VisionZip | 41.4 (65.9) | 57.6 (85.6) | 62.2 (93.7) | 70.0 (96.5) | 62.0 (92.5) | 55.7 (77.3) | 35.2 (60.4) | 64.6 (93.2) | 56.1 (83.1) | 73.8 (89.5) |
| Method | Chart QA | GQA | MMB CN | MMB EN | MME | MMStar | OCR Bench | POPE | Acc | Rel |
|---|---|---|---|---|---|---|---|---|---|---|
| Reference image-area reduction: 75.00% | ||||||||||
| Vanilla | 88.5 (94.3) | 65.0 (88.2) | 67.4 (94.7) | 76.5 (97.9) | 70.4 (93.9) | 75.4 (86.7) | 83.8 (90.4) | 81.0 (96.7) | 76.0 (92.8) | 100.0 (100.0) |
| FastV | 17.7 (40.1) | 37.3 (60.8) | 43.5 (82.3) | 49.3 (82.5) | 50.7 (77.8) | 34.4 (55.9) | 15.1 (26.9) | 49.6 (69.8) | 37.2 (62.0) | 48.9 (66.8) |
| SparseVLM | 10.0 (24.5) | 40.8 (68.5) | 50.9 (89.5) | 58.5 (89.7) | 57.7 (84.7) | 39.3 (61.4) | 11.7 (23.2) | 54.0 (80.3) | 40.4 (65.2) | 53.1 (70.2) |
| DivPrune | 35.4 ( 62.2 ) | 54.1 (82.9) | 59.1 (92.8) | 63.6 (94.7) | 67.6 (90.1) | 52.5 (74.0) | 31.3 ( 58.3 ) | 66.7 ( 94.2 ) | 53.8 ( 81.2 ) | 70.7 ( 87.4 ) |
| VisionZip | 24.3 (45.4) | 50.9 (81.3) | 53.9 (92.5) | 62.7 (94.8) | 62.0 (90.2) | 46.4 (70.7) | 17.3 (49.4) | 55.6 (91.1) | 46.6 (76.9) | 61.4 (82.8) |
| Qwen2.5-VL-7B | Qwen3-VL-8B | |||||
| Method | val | testA | testB | val | testA | testB |
| Vanilla | 66.8 (100.0%) | 47.3 (100.0%) | 51.9 (100.0%) | 49.1 (100.0%) | 47.6 (100.0%) | 47.8 (100.0%) |
| Retain 256 tokens (80.2% pruned) | ||||||
| FastV (ECCV’24) | 61.9 (92.8%) | 27.4 (57.9%) | 34.6 (66.7%) | 55.3 (112.7%) | 38.0 (79.9%) | 39.6 (82.8%) |
| VisPruner (ICCV’25) | 64.0 (95.8%) | 35.8 (75.8%) | 40.8 (78.5%) | 52.0 (106.0%) | 41.7 (87.5%) | 42.9 (89.6%) |
| ApET (CVPR’26) | 62.3 (93.3%) | 27.6 (58.3%) | 33.0 (63.6%) | 50.0 (101.9%) | 31.4 (66.0%) | 35.1 (73.4%) |
| LLaVA-NeXT-7B | LLaVA-NeXT-13B | Qwen2.5-VL-7B | InternVL3-8B | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method / | 640 | 320 | 160 | 640 | 320 | 160 | 256 | 128 | 256 | 128 |
| Vanilla | 50.7 | 51.1 | 87.7 | 88.1 | ||||||
| FastV | 37.0 | 11.9 | – | 38.9 | 18.9 | – | 70.6 | 39.6 | 55.0 | 25.0 |
| DivPrune | 38.2 | 30.1 | 23.8 | 39.5 | 31.7 | 26.2 | 75.1 | 66.7 | 44.9 | 32.5 |
| VisPruner | 43.6 | 34.2 | 25.8 | 47.8 | 37.9 | 31.6 | 71.5 | 65.5 | 26.9 | 6.9 |
| MMTok | 40.8 | 30.4 | 24.9 | 42.3 | 34.1 | 26.3 | 71.1 | 62.7 | 51.7 | 35.3 |