Beyond Selection: Token Parameterization for Extreme Visual Token Compression
Organizations: Zhejiang University
Abstract
Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training complexity. We revisit compression through a token parameterization lens, separating (i) basis transformation and structured truncation (retained subspace/compressibility) from (ii) coordinate organization (optimization and cross-modal alignment). This view yields two coupled objectives, compressibility and learnability, which we formalize as unified functionals. Guided by these objectives, we design Braco, a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling. Experiments show that Braco forms the favorable empirical accuracy-efficiency frontier under -- compression and remains competitive at , reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%--86.7% relative to the uncompressed upper bound. Against prior methods, Braco matches or improves accuracy while achieving up to approximately 36% end-to-end speedup and using / lower compressor latency/FLOPs.
Figures & tables
| Method | GQA | MMB | MMB | MME | POPE | SQA | VQA-T | MMVet | Acc. (%) | FLOPs (T) | Lat. (ms) |
| Upper Bound, 576 Tokens ( 1 ) | |||||||||||
| Vanilla | 62.9 | 65.5 | 60.7 | 1785 | 85.7 | 69.7 | 58.0 | 32.8 | 100.0 | 8.67 | 67.25 |
| 25 Retained Tokens ( 23 ) | |||||||||||
| PruMerge (ICCV25) | 49.8 | 55.8 | 47.0 | 1515 | 58.5 | 68.7 | 50.4 | 20.2 | 80.2 | 1.37 | 44.10 |
| DivPrune (CVPR25) | 54.0 | 59.8 | 50.8 | 1510 | 75.2 | 68.5 | 49.5 | 26.0 | 87.0 | 1.40 | 41.00 |
| MQT-LLaVA (NIPS24) | 57.1 | 61.4 | 53.1 | 1689 | 79.9 | 69.4 | 50.2 | 27.7 | 91.3 | 1.37 | 44.20 |
| Method | Boundary | Acc. (%) | Lat. (ms) | FLOPs (G) | Mem. (MB) |
| QueCC (ICLR25) | pre-proj. | 93.9 | 17.864 | 109.504 | 14.875 |
| MQT-LLaVA (NIPS24) | pre-proj. | 90.2 | 5.154 | 2.674 | 11.336 |
| PruMerge (ICCV25) | pre-proj. | 76.2 | 3.300 | 0.724 | 13.749 |
| Braco (ours) | pre-proj. | 94.0 | 1.073 | 1.389 | 8.731 |
| TokenPacker (IJCV25) | post-proj. | 93.7 | 1.004 | 15.271 | 15.125 |
| DivPrune (CVPR25) | post-proj. | 83.2 | 1.061 | 29.603 | 23.008 |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Design axis | Failure mode diagnosed | Braco choice | End-to-end confirmation |
| Retained subspace | Tiny structured budgets may discard task-relevant visual directions. | Use a DCT low-frequency backbone. | DCT has higher energy/readability in Figure 3 ; replacing it with spatial/Haar tokens drops Acc. in Figure 4(c) ; functional rankings are validated in Section J.6 . |
| Coordinate organization | Equal-information coordinates can differ in optimization and alignment. | Use a budget-dependent coordinate rule. | Figure 2(c) isolates this effect within the same retained subspace; wrong coordinates or random rotations reduce Acc. in Figure 4(c) . |
| Token identity | Transform coefficients no longer carry ordinary patch-position semantics. | Add basis-coordinate positional embeddings. | Replacing the basis-coordinate embedding with learned or 2D sine-cosine alternatives reduces Acc. in Figure 4(b) . |
| Local evidence | A pure low-pass backbone can miss sparse localized evidence. | Allocate a small spatial residual budget. | The hybrid c2s5 interface outperforms backbone-only and residual-only variants in Figures 4(a) and 6 . |
| Compressor overhead | Learned resamplers can recover accuracy while adding module cost. | Keep the interface prompt-independent and lightweight. | Braco matches QueCC-level Acc. with lower compressor latency/FLOPs in Tables 2 and 7 . |
| Budget | Config. | Backbone | Residual | Coordinate org. | PE |
| 4 | c1s3 | 1 | 3 | vanilla | yes |
| 9 | c2s5 | 4 | 5 | vanilla | yes |
| 16 | c3s7 | 9 | 7 | vanilla | yes |
| 25 | c4s9 | 16 | 9 | idct | yes |
| Allocation | Step 1 | Step 2 | Step 3 | Step 4 |
| c3s7 | 56.623 | 38.928 | 0 | 1236.293 |
| c4s9 | 56.623 | 38.928 | 0.262 | 1240.663 |
| Method | Acc. (%) | FLOPs (G) | Latency (ms) | Peak memory (MB) |
| c2s5 | 93.2 | 1.382 | 1.067 | 8.729 |
| c2s0 | 85.7 | 0.152 | 0.288 | 4.504 |
| c0s5 | 88.2 | 1.382 | 1.091 | 8.729 |
| c3s0 | 89.8 | 0.152 | 0.289 | 4.504 |
| c0s9 | 92.0 | 1.396 | 1.077 | 8.734 |
| Method | #Tokens | FLOPs (G) | Latency (ms) | Peak memory (MB) |
| Braco | 4 | 1.375 | 1.060 | 8.727 |
| Braco | 9 | 1.382 | 1.067 | 8.729 |
| Braco | 25 | 1.396 | 1.197 | 8.734 |
| Method | LLM | Input | Retain | GQA | MMB | MMB | MME | POPE | SQA | VQA | MMVet | Acc. | FLOPs | Lat. |
| tokens | tokens | (%) | (T) | (ms) | ||||||||||
| Vanilla | Vicuna-7B | 576 | 576 | 62.9 | 65.5 | 60.7 | 1785 | 85.7 | 69.7 | 58.0 | 32.8 | 100.0 | 8.67 | 67.25 |
| QueCC | Vicuna-7B | 576 | 16 (36 ) | 59.0 | 63.1 | 54.6 | 1668 | 83.5 | 70.6 | 52.8 | 28.8 | 93.9 | 1.36 | 63.22 |
| Braco | Vicuna-7B | 576 | 16 (36 ) | 57.5 | 63.3 | 55.8 | 1704 | 82.4 | 69.5 | 51.9 | 29.8 | 94.0 | 1.25 | 40.59 |
| Vanilla | Vicuna-7B | 2880 | 2880 | 63.7 | 66.2 | 59.7 | 1701 | 86.6 | 67.8 | 64.1 | 31.2 | 100.0 | 40.57 | 261.08 |
| QueCC | Vicuna-7B | 2880 | 80 (36 ) | 62.0 | 65.1 | 57.5 | 1762 | 85.2 | 68.6 | 59.1 | 29.4 | 97.7 | 3.99 | 66.73 |
| Method | #Vision tokens | Pre-training | Instruction-tuning |
| Vanilla | 576 | 3.5h | 10h |
| QueCC | 16 (36 ) | 0.7h | 7h |
| Braco | 16 (36 ) | 0.4h | 6.5h |
| Vanilla | 2880 | 14h | 28h |
| QueCC | 80 (36 ) | 1.1h | 8.2h |
| Braco | 80 (36 ) | 1h | 7.5h |
| Vision encoder / resolution | LLaVA-Pretrain | DocVQA |
| CLIP ViT-L/14-336 | 16 (16–16) | 25 (16–25) |
| SigLIP-SO400M/14-384 | 64 (64–64) | 81 (64–81) |
| SigLIP2-SO400M/16-512 | 36 (36–36) | 49 (36–49) |
| Tokens | Braco | QueCC | TokenPacker |
| 25 | — | ||
| 16 | |||
| 9 | |||
| 4 | — |