This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly unstable: it improves accuracy by up to 82.5% relative on some datasets and degrades it by up to 100% on others. We trace this instability to spurious inversion: background patches receive higher CLIP text-similarity than the true object when the spurious attribute is background-separable, inverting the assumption every text- and attention-guided pruning method relies on. We introduce the Spurious Inversion Metric (SIM), a label-free, pre-deployment diagnostic whose sign predicts this effect with statistical significance (binomial p=0.035) across all 8 datasets, and remains dependable across 6 CLIP architectures with a clean foreground/background split. Naive masking is itself a major source of risk: it causes the largest average-accuracy loss of any method we evaluate, and its own per-image segmentation step is a significant runtime bottleneck. To address this, we design a batched, synchronization-free GPU segmentation routine that cuts this overhead from 3.5× to 1.75× baseline. Gating deployment by SIM's sign recovers masking's benefits while avoiding its worst failures, matching or exceeding a strong pruning baseline on 7 of 8 datasets.
Figures & tables
Figure 1: Foreground vs. background text-similarity on Waterbirds (background-spurious) and CelebA (in-object-spurious). Background patches outrank the object on Waterbirds across every group; the reverse holds on CelebA. Error bars: one standard deviation across images within each group.
Figure 2: The SIM pipeline on two real, representative images (selected as the image within each dataset closest to that dataset’s own mean SIM, not a cherry-picked extreme): original image, the unsupervised foreground/background mask, the per-patch text-similarity heatmap, and the resulting foreground/background means, SIM value, and gate decision. On Waterbirds, background patches (sky, water, wall) outrank the bird itself, giving positive SIM and a MARS decision; on CelebA, the foreground face outranks its surroundings, giving negative SIM and a fallback-to-FastV decision.
Method
Rel. Δ WG (%)
Avg. Acc. (%)
Latency (ms)
Hurt
Helped
Baseline (uncompressed)
-
79.9
15.7
-
-
FastV [ 2 ]
+2.3%
79.7
12.4
2/8
3/8
FiCoCo [ 10 ]
+2.3%
78.9
15.7
4/8
3/8
EViT [ 3 ]
−4.3%
79.7
12.7
4/8
2/8
PatchRank [ 11 ]
−5.2%
80.8
15.1
4/8
2/8
PACT [ 9 ]
−7.1%
78.5
34.1
5/8
2/8
Table 1: Mean relative change in worst-group accuracy vs. the uncompressed baseline (%), average accuracy (%), and per-image latency (batch size 1; MARS latency uses the original CPU segmentation step, Appendix G ), averaged over 8 datasets, for each pruning method at a 50% token budget. Hurt/Helped count datasets with strictly negative/positive Δ WG; the remainder show negligible ( ≈ 0) change.
Figure 3: SIM against MARS’s measured change in worst-group accuracy. Seven of eight datasets fall on the side of zero that SIM predicts; the single exception, MetaShifts, is analyzed as a structural boundary case in Section 5.4 .
Dataset
Baseline (256 tok)
Blind MARS
Blind FastV
SIM-gated
Waterbirds [ 4 ]
50.8
67.3
59.3
67.3
UrbanCars [ 6 ]
22.8
40.3
26.2
40.3
CelebA [ 5 ]
79.6
64.7
75.6
75.6
MetaShifts [ 7 ]
86.7
52.2
87.8
52.2
ImageNet-9 [ 8 ]
87.5
26.3
85.3
85.3
OxfordPets [ 13 ]
52.6
0.0
52.6
52.6
Table 2: Worst-group accuracy (%) under three deployment policies, all at 128 tokens except the uncompressed baseline (256 tokens).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Method
WG acc. (%)
Avg. acc. (%)
Tokens
Waterbirds [ 4 ]
Baseline
50.8±2.4
77.8±1.0
256
Masked (no prune)
66.5±3.3
81.0±0.7
256
Semantic-prune-only
66.3±2.9
80.1±0.7
128
MARS (hybrid)
67.3±2.7
80.3±0.7
128
Text-guided [ 12 ]
18.5±0.8
56.5±0.3
128
FastV [ 2 ]
59.3±3.1
79.4±1.1
128
Appendix
Table 3: Worst-group and average accuracy (%, mean ± std over 3 seeds) for all evaluated methods on the three datasets with seeded evaluation.
Figure 4: Sign-test correctness across 6 CLIP variants and 8 datasets (48 combinations, 31 correct, binomial p=0.030 ). (a) Full per-architecture breakdown: MetaShifts (column 4) is incorrect for every architecture, a reproducible structural boundary rather than architecture-specific noise. (b) The same data collapsed to a per-dataset correctness rate against the 3/6 chance level: three of the datasets the diagnostic was designed around sit at or near 6/6, MetaShifts sits at 0/6 (below chance), and the fine-grained classification benchmarks fall in between, where masking-based methods collapse regardless of the gate’s decision.
Architecture
Dataset
SIM
MARS WG
FastV WG
Correct
ViT-B/32 (openai)
Waterbirds [ 4 ]
+0.0287
63.0
37.0
✓
UrbanCars [ 6 ]
+0.0011
35.0
27.0
✓
CelebA [ 5 ]
-0.0164
70.0
79.0
✓
MetaShifts [ 7 ]
+0.0089
56.0
76.0
×
ImageNet-9 [ 8 ]
-0.0039
50.0
73.0
✓
OxfordPets [ 13 ]
-0.0081
1.8
1.8
×
Appendix
Table 4: Full architecture generalization results. “Correct” indicates the sign test ( SIM>0⇒ MARS) selects the method with higher worst-group accuracy.
Figure 5: Worst-group accuracy under blind MARS, blind FastV, and the SIM-gated policy. The gate tracks whichever of MARS or FastV is preferable on 7 of 8 datasets; on MetaShifts (bars 4) it follows MARS, reproducing the one structural miss identified in Table 5 .
Dataset
SIM
Rel. Δ WG (%)
Predicted
Actual
Correct
Waterbirds [ 4 ]
+0.044
+25.6%
help
help
✓
UrbanCars [ 6 ]
+0.051
+82.5%
help
help
✓
CelebA [ 5 ]
−0.019
−19.6%
hurt
hurt
✓
MetaShifts [ 7 ]
+0.025
−39.7%
help
hurt
×
ImageNet-9 [ 8 ]
−0.021
−70.0%
hurt
hurt
✓
OxfordPets [ 13 ]
−0.024
−100.0%
hurt
hurt
✓
Appendix
Table 5: SIM and MARS’s measured relative change in worst-group accuracy vs. the uncompressed baseline, per dataset, on OpenCLIP ViT-L/14.
Dataset
SIM (stratified)
SIM (non-stratified)
Rel. Δ WG (%)
Correct
Waterbirds [ 4 ]
+0.0443
+0.0478
+25.6%
✓
UrbanCars [ 6 ]
+0.0513
+0.0549
+82.5%
✓
CelebA [ 5 ]
−0.0194
−0.0246
−19.6%
✓
MetaShifts [ 7 ]
+0.0249
+0.0233
−39.7%
×
ImageNet-9 [ 8 ]
−0.0213
−0.0213
−70.0%
✓
OxfordPets [ 13 ]
−0.0240
−0.0243
−100.0%
✓
Appendix
Table 6: Non-stratified vs. group-stratified SIM, 200 unstratified images per dataset. “Correct” indicates the non-stratified estimate’s sign matches the known direction of MARS’s effect on worst-group accuracy.
Component
Latency (ms)
Tokens
Baseline forward pass
15.7
256
First-pass forward for segmentation embeddings
≈ 15.7
256
Segmentation step (CPU, original)
28.5
-
Segmentation step (batched GPU, ours, b =128)
0.81
-
MARS second-pass forward (masked, pruned)
11.0
128
FastV forward pass (single pass, no masking)
12.4
128
Appendix
Table 7: Per-image latency (ms) on a single NVIDIA A100, OpenCLIP ViT-L/14. All rows are batch size 1 except the batched GPU segmentation row (batch 128; see Appendix H for the full batch-size sweep).
Batch size
Per-image-adaptive PCA
Shared PCA basis
1
9.06
7.65
8
1.85
0.95
32
0.98
0.25
64
0.87
0.14
128
0.81
unreliable †
Appendix
Table 8: Per-image segmentation cost (ms) by batch size, batched GPU reimplementation, OpenCLIP ViT-L/14 patch embeddings on real Waterbirds images (CUDA-event timing, 10 warmup + 30 measured runs). Reference: CPU (sklearn) path, 28.5ms/image at batch size 1 (Table 7 ); original unbatched single-image GPU attempt, 161ms/image.
Dataset
Per-image-adaptive PCA
Shared PCA basis
Mask IoU
Rel. Δ WG (%)
Mask IoU
Rel. Δ WG (%)
Waterbirds [ 4 ]
0.680
+6.5%
0.666
+21.8%
UrbanCars [ 6 ]
0.596
+0.0%
0.581
−9.4%
CelebA [ 5 ]
0.626
+2.3%
0.633
+4.5%
Appendix
Table 9: Correctness validation for the batched GPU segmentation, 300 held-out images per dataset, against the CPU (sklearn) reference. Rel. Δ WG is the batched variant’s worst-group accuracy relative to the CPU reference’s, on the same images.
Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low token budget. Our analysis shows that as the pruning budget decreases, accuracy degradation is often accompanied by larger feature distribution shifts. Critically, the degree of this distribution shift strongly correlates with performance degradation. To better characterize this phenomenon, we introduce a lightweight distribution consistency metric to estimate the distribution shift between retained and full tokens. Motivated by these observations, we propose a two-stage pruning framework consisting of Anchor-Context Graph Recovery (ACGR) and Text-Aware Token Cluster Selection (TATCS). Specifically, ACGR transfers contextual information before token removal, while TATCS dynamically re-selects representative tokens when severe distribution shift is detected. Extensive experiments demonstrate that our method achieves superior and more stable performance under ultra-low token budget. Notably, it retains 92.1% of the upper-bound average performance on LLaVA-1.5-7B with only 16 visual tokens.
Xifeng Xue, Xiaokang Wang, Zirui Li +2
College of Computer Science, Nankai University, Tianjin, China · Nanjing University of Posts and Telecommunications, Nanjing, China
Visual token pruning reduces the inference cost of vision-language models (VLMs), but most methods only ask which tokens to keep. This retained-token view can keep redundant high-scoring tokens while leaving discarded evidence without a close representative. We propose CoverPruner, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM? CoverPruner formulates pruning as Representational Coverage Maximization (RCM), covering the full projected visual-token set with query-weighted demand. It instantiates RCM with projector-space coverage and a lightweight first-layer attention probe. Across multiple VLM architectures and compression rates, CoverPruner achieves the best average accuracy among all compared methods, with the largest gains usually appearing under aggressive compression.
Qingchan Zhu, Weihang You, Hanqi Jiang +3
School of Computing, University of Georgia · College of Engineering, Northeastern University
Vision Transformers (ViTs) are strong backbones for semantic segmentation, but their computational cost limits deployment. Recent token compression methods for efficient transformer-based segmentation reduce this cost by decreasing the number of tokens. However, existing evaluations primarily focus on low-to-moderate compression, leaving their behavior under aggressive compression and corrupted inputs unclear. Meanwhile, structural pruning provides an orthogonal route to efficiency by removing redundant components in the ViT architecture, but is rarely compared to token compression under a unified protocol. To bridge this gap, we benchmark representative token compression and structural pruning methods for ViT-based semantic segmentation under matched FLOPs on ADE20K and Cityscapes, together with their common-corruption variants ADE20K-C and Cityscapes-C. Our results reveal a consistent trend on both clean and corrupted inputs: token compression is highly effective at mild reductions but degrades sharply when compression becomes severe, consistent with substantial information loss from overly aggressive token reduction. In contrast, structural pruning exhibits a smoother degradation curve and is more stable at high compression. Motivated by these findings, we study a prune-then-merge pipeline that applies moderate token compression on top of a moderately pruned backbone. At comparable FLOPs, this combined strategy consistently achieves a better accuracy-robustness trade-off at high compression, offering a practical recipe for deployment-oriented ViT segmentation. Code is available at https://github.com/phatnguyencs/vit-seg-compression.
Tien-Phat Nguyen, Ngai-Man Cheung
Temasek Laboratories, Singapore University of Technology and Design, Singapore