Visual-token compression is effective for improving the efficiency of vision-language models, but under extreme compression budgets, token pruning can break visual grounding while learned resamplers increase parameter count, attention cost, and training complexity. We revisit compression through a token parameterization lens, separating (i) basis transformation and structured truncation (retained subspace/compressibility) from (ii) coordinate organization (optimization and cross-modal alignment). This view yields two coupled objectives, compressibility and learnability, which we formalize as unified functionals. Guided by these objectives, we design Braco, a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling. Experiments show that Braco forms the favorable empirical accuracy-efficiency frontier under 23×--64× compression and remains competitive at 144×, reaching 95.2% accuracy while reducing prefill FLOPs by 84.2%--86.7% relative to the uncompressed upper bound. Against prior methods, Braco matches or improves accuracy while achieving up to approximately 36% end-to-end speedup and using 16.6×/78.8× lower compressor latency/FLOPs.
Figures & tables
Figure 1 : Our four-step token coder. Starting from an N×N visual-token grid, Braco transforms tokens into an orthonormal basis, keeps a fixed C×C low-frequency block, adds basis-coordinate embeddings, optionally changes coordinates within the same retained subspace, and appends learned spatial residual tokens pooled in parallel from the original grid.
Method
GQA
MMB EN
MMB CN
MME All
POPE F1
SQA
VQA-T
MMVet
Acc. (%)
FLOPs (T)
Lat. (ms)
Upper Bound, 576 Tokens ( 1 × )
Vanilla
62.9
65.5
60.7
1785
85.7
69.7
58.0
32.8
100.0
8.67
67.25
25 Retained Tokens ( 23 × )
PruMerge (ICCV25)
49.8
55.8
47.0
1515
58.5
68.7
50.4
20.2
80.2
1.37
44.10
DivPrune (CVPR25)
54.0
59.8
50.8
1510
75.2
68.5
49.5
26.0
87.0
1.40
41.00
MQT-LLaVA (NIPS24)
57.1
61.4
53.1
1689
79.9
69.4
50.2
27.7
91.3
1.37
44.20
Table 1 : Main results under matched visual-token budgets. We report eight benchmark scores, Acc. ( ↑ ; the mean score normalized by the 576-token Vanilla model), and single-image full-pipeline prefill FLOPs/latency ( ↓ ). QueCC is omitted at 25 tokens because its native grid does not support a 5×5 output on the fixed 24×24 visual-token lattice. Braco is on the leading empirical Acc.–cost frontier at 25/16/9 tokens and remains within 0.2 Acc. of QueCC at 4 tokens with lower cost.
Figure 2 : Compression vs. optimization probes. (a–b) Orthonormal basis choice controls energy retention under deployable structured truncation and oracle magnitude truncation. KLT is an oracle upper bound fitted to the diagnostic token second moment. (c) With an identical retained subspace, coordinate organization can still change optimization behavior.
Figure 3 : Compact compressibility diagnostics. (a) DCT concentrates energy into a compact low-frequency region. (b–c) CelebA linear probes test task-readability under the same structured truncation rule.
Method
Boundary
Acc. (%)
Lat. (ms)
FLOPs (G)
Mem. (MB)
QueCC (ICLR25)
pre-proj.
93.9
17.864
109.504
14.875
MQT-LLaVA (NIPS24)
pre-proj.
90.2
5.154
2.674
11.336
PruMerge (ICCV25)
pre-proj.
76.2
3.300
0.724
13.749
Braco (ours)
pre-proj.
94.0
1.073
1.389
8.731
TokenPacker (IJCV25)
post-proj.
93.7
1.004
15.271
15.125
DivPrune (CVPR25)
post-proj.
83.2
1.061
29.603
23.008
Table 2 : Compressor cost at 16 tokens. Boundary denotes pre- or post-projector measurement.
Figure 4 : Ablations. (a) Backbone–residual allocation under the 9-token budget. (b) Basis-coordinate embedding variants under the 16-token budget. (c) Basis and coordinate substitutions under the 16-token budget.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Design axis
Failure mode diagnosed
Braco choice
End-to-end confirmation
Retained subspace
Tiny structured budgets may discard task-relevant visual directions.
Use a DCT low-frequency backbone.
DCT has higher energy/readability in Figure 3 ; replacing it with spatial/Haar tokens drops Acc. in Figure 4(c) ; functional rankings are validated in Section J.6 .
Coordinate organization
Equal-information coordinates can differ in optimization and alignment.
Use a budget-dependent coordinate rule.
Figure 2(c) isolates this effect within the same retained subspace; wrong coordinates or random rotations reduce Acc. in Figure 4(c) .
Token identity
Transform coefficients no longer carry ordinary patch-position semantics.
Add basis-coordinate positional embeddings.
Replacing the basis-coordinate embedding with learned or 2D sine-cosine alternatives reduces Acc. in Figure 4(b) .
Local evidence
A pure low-pass backbone can miss sparse localized evidence.
Allocate a small spatial residual budget.
The hybrid c2s5 interface outperforms backbone-only and residual-only variants in Figures 4(a) and 6 .
Compressor overhead
Learned resamplers can recover accuracy while adding module cost.
Keep the interface prompt-independent and lightweight.
Braco matches QueCC-level Acc. with lower compressor latency/FLOPs in Tables 2 and 7 .
Appendix
Table 3 : Design-validation map. Each diagnostic identifies a failure mode, motivates a Braco design choice, and connects it to the corresponding end-to-end experiment.
Budget K
Config.
Backbone C2
Residual S
Coordinate org.
PE
4
c1s3
1
3
vanilla
yes
9
c2s5
4
5
vanilla
yes
16
c3s7
9
7
vanilla
yes
25
c4s9
16
9
idct
yes
Appendix
Table 4 : Braco configurations used in the main table. c{C}s{S} denotes a C×C low-frequency backbone and S spatial residual tokens; PE denotes the basis-coordinate positional embedding.
Allocation
Step 1
Step 2
Step 3
Step 4
c3s7
56.623
38.928
0
1236.293
c4s9
56.623
38.928
0.262
1240.663
Appendix
Table 5 : Analytic FLOPs decomposition of Braco stages. Values are reported in MFLOPs for the two representative allocations used in the main paper.
Method
Acc. (%)
FLOPs (G)
Latency (ms)
Peak memory (MB)
c2s5
93.2
1.382
1.067
8.729
c2s0
85.7
0.152
0.288
4.504
c0s5
88.2
1.382
1.091
8.729
c3s0
89.8
0.152
0.289
4.504
c0s9
92.0
1.396
1.077
8.734
Appendix
Table 6 : Ablation cost: residual-only variants at the compressor boundary. Residual-only variants are weaker and not cheaper overall than the hybrid backbone–residual allocation.
Method
#Tokens
FLOPs (G)
Latency (ms)
Peak memory (MB)
Braco
4
1.375
1.060
8.727
Braco
9
1.382
1.067
8.729
Braco
25
1.396
1.197
8.734
Appendix
Table 7 : Multi-budget Braco compressor cost. Peak memory varies little across retained-token budgets.
Method
LLM
Input
Retain
GQA
MMB EN
MMB CN
MME
POPE
SQA
VQA Text
MMVet
Acc.
FLOPs
Lat.
tokens
tokens
(%)
(T)
(ms)
Vanilla
Vicuna-7B
576
576
62.9
65.5
60.7
1785
85.7
69.7
58.0
32.8
100.0
8.67
67.25
QueCC
Vicuna-7B
576
16 (36 × )
59.0
63.1
54.6
1668
83.5
70.6
52.8
28.8
93.9
1.36
63.22
Braco
Vicuna-7B
576
16 (36 × )
57.5
63.3
55.8
1704
82.4
69.5
51.9
29.8
94.0
1.25
40.59
Vanilla
Vicuna-7B
2880
2880
63.7
66.2
59.7
1701
86.6
67.8
64.1
31.2
100.0
40.57
261.08
QueCC
Vicuna-7B
2880
80 (36 × )
62.0
65.1
57.5
1762
85.2
68.6
59.1
29.4
97.7
3.99
66.73
Appendix
Table 8 : Generalization to larger visual-token inputs. FLOPs and latency include vision encoder, projector, and LLM prefill.
Method
#Vision tokens
Pre-training
Instruction-tuning
Vanilla
576
3.5h
10h
QueCC
16 (36 × )
0.7h
7h
Braco
16 (36 × )
0.4h
6.5h
Vanilla
2880
14h
28h
QueCC
80 (36 × )
1.1h
8.2h
Braco
80 (36 × )
1h
7.5h
Appendix
Table 9 : Training time on 8 NVIDIA A100 GPUs. All entries use the same two-stage training recipe.
Figure 5 : Latency scaling curves on A100 and A800. Comparison of Braco and Vanilla across batch sizes. Braco consistently achieves lower latency than Vanilla, and both methods exhibit increasing latency as batch size grows. Vanilla runs out of memory (OOM) at batch size 128 and above, so its curves stop at batch size 64. Curves show mean latency, with shaded bands indicating standard deviation over repeated measurements. This batch-scaling profile uses the same measurement protocol within each curve and is intended for scaling comparison.
Table 10 : Latency comparison across batch sizes on A100 and A800. Reported values are mean latency (ms) with standard deviation; OOM denotes Vanilla out-of-memory at batch size 128 and above.
Vision encoder / resolution
LLaVA-Pretrain Kb⋆
DocVQA Kb⋆
CLIP ViT-L/14-336
16 (16–16)
25 (16–25)
SigLIP-SO400M/14-384
64 (64–64)
81 (64–81)
SigLIP2-SO400M/16-512
36 (36–36)
49 (36–49)
Appendix
Table 11 : Coordinate-rule calibration across encoders and distributions. Entries report the median crossover budget Kb⋆ and its range across three resamples of 1,024 unlabeled images.
Tokens
Braco
QueCC
TokenPacker
25
95.18±0.10
—
94.24±0.13
16
94.12±0.10
93.88±0.12
93.70±0.15
9
93.29±0.12
93.02±0.14
91.48±0.18
4
91.28±0.15
91.35±0.16
—
Appendix
Table 12 : Accuracy across three matched training seeds. Values are sample mean ± sample standard deviation. QueCC does not support 25 tokens on the fixed 24×24 lattice; TokenPacker is not included in the four-token repeated-training comparison.