This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly unstable: it improves accuracy by up to 82.5% relative on some datasets and degrades it by up to 100% on others. We trace this instability to spurious inversion: background patches receive higher CLIP text-similarity than the true object when the spurious attribute is background-separable, inverting the assumption every text- and attention-guided pruning method relies on. We introduce the Spurious Inversion Metric (SIM), a label-free, pre-deployment diagnostic whose sign predicts this effect with statistical significance (binomial p=0.035) across all 8 datasets, and remains dependable across 6 CLIP architectures with a clean foreground/background split. Naive masking is itself a major source of risk: it causes the largest average-accuracy loss of any method we evaluate, and its own per-image segmentation step is a significant runtime bottleneck. To address this, we design a batched, synchronization-free GPU segmentation routine that cuts this overhead from 3.5× to 1.75× baseline. Gating deployment by SIM's sign recovers masking's benefits while avoiding its worst failures, matching or exceeding a strong pruning baseline on 7 of 8 datasets.
Figures & tables
Figure 1: Foreground vs. background text-similarity on Waterbirds (background-spurious) and CelebA (in-object-spurious). Background patches outrank the object on Waterbirds across every group; the reverse holds on CelebA. Error bars: one standard deviation across images within each group.
Figure 2: The SIM pipeline on two real, representative images (selected as the image within each dataset closest to that dataset’s own mean SIM, not a cherry-picked extreme): original image, the unsupervised foreground/background mask, the per-patch text-similarity heatmap, and the resulting foreground/background means, SIM value, and gate decision. On Waterbirds, background patches (sky, water, wall) outrank the bird itself, giving positive SIM and a MARS decision; on CelebA, the foreground face outranks its surroundings, giving negative SIM and a fallback-to-FastV decision.
Method
Rel. Δ WG (%)
Avg. Acc. (%)
Latency (ms)
Hurt
Helped
Baseline (uncompressed)
-
79.9
15.7
-
-
FastV [ 2 ]
+2.3%
79.7
12.4
2/8
3/8
FiCoCo [ 10 ]
+2.3%
78.9
15.7
4/8
3/8
EViT [ 3 ]
−4.3%
79.7
12.7
4/8
2/8
PatchRank [ 11 ]
−5.2%
80.8
15.1
4/8
2/8
PACT [ 9 ]
−7.1%
78.5
34.1
5/8
2/8
Table 1: Mean relative change in worst-group accuracy vs. the uncompressed baseline (%), average accuracy (%), and per-image latency (batch size 1; MARS latency uses the original CPU segmentation step, Appendix G ), averaged over 8 datasets, for each pruning method at a 50% token budget. Hurt/Helped count datasets with strictly negative/positive Δ WG; the remainder show negligible ( ≈ 0) change.
Figure 3: SIM against MARS’s measured change in worst-group accuracy. Seven of eight datasets fall on the side of zero that SIM predicts; the single exception, MetaShifts, is analyzed as a structural boundary case in Section 5.4 .
Dataset
Baseline (256 tok)
Blind MARS
Blind FastV
SIM-gated
Waterbirds [ 4 ]
50.8
67.3
59.3
67.3
UrbanCars [ 6 ]
22.8
40.3
26.2
40.3
CelebA [ 5 ]
79.6
64.7
75.6
75.6
MetaShifts [ 7 ]
86.7
52.2
87.8
52.2
ImageNet-9 [ 8 ]
87.5
26.3
85.3
85.3
OxfordPets [ 13 ]
52.6
0.0
52.6
52.6
Table 2: Worst-group accuracy (%) under three deployment policies, all at 128 tokens except the uncompressed baseline (256 tokens).
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Method
WG acc. (%)
Avg. acc. (%)
Tokens
Waterbirds [ 4 ]
Baseline
50.8±2.4
77.8±1.0
256
Masked (no prune)
66.5±3.3
81.0±0.7
256
Semantic-prune-only
66.3±2.9
80.1±0.7
128
MARS (hybrid)
67.3±2.7
80.3±0.7
128
Text-guided [ 12 ]
18.5±0.8
56.5±0.3
128
FastV [ 2 ]
59.3±3.1
79.4±1.1
128
Appendix
Table 3: Worst-group and average accuracy (%, mean ± std over 3 seeds) for all evaluated methods on the three datasets with seeded evaluation.
Figure 4: Sign-test correctness across 6 CLIP variants and 8 datasets (48 combinations, 31 correct, binomial p=0.030 ). (a) Full per-architecture breakdown: MetaShifts (column 4) is incorrect for every architecture, a reproducible structural boundary rather than architecture-specific noise. (b) The same data collapsed to a per-dataset correctness rate against the 3/6 chance level: three of the datasets the diagnostic was designed around sit at or near 6/6, MetaShifts sits at 0/6 (below chance), and the fine-grained classification benchmarks fall in between, where masking-based methods collapse regardless of the gate’s decision.
Architecture
Dataset
SIM
MARS WG
FastV WG
Correct
ViT-B/32 (openai)
Waterbirds [ 4 ]
+0.0287
63.0
37.0
✓
UrbanCars [ 6 ]
+0.0011
35.0
27.0
✓
CelebA [ 5 ]
-0.0164
70.0
79.0
✓
MetaShifts [ 7 ]
+0.0089
56.0
76.0
×
ImageNet-9 [ 8 ]
-0.0039
50.0
73.0
✓
OxfordPets [ 13 ]
-0.0081
1.8
1.8
×
Appendix
Table 4: Full architecture generalization results. “Correct” indicates the sign test ( SIM>0⇒ MARS) selects the method with higher worst-group accuracy.
Figure 5: Worst-group accuracy under blind MARS, blind FastV, and the SIM-gated policy. The gate tracks whichever of MARS or FastV is preferable on 7 of 8 datasets; on MetaShifts (bars 4) it follows MARS, reproducing the one structural miss identified in Table 5 .
Dataset
SIM
Rel. Δ WG (%)
Predicted
Actual
Correct
Waterbirds [ 4 ]
+0.044
+25.6%
help
help
✓
UrbanCars [ 6 ]
+0.051
+82.5%
help
help
✓
CelebA [ 5 ]
−0.019
−19.6%
hurt
hurt
✓
MetaShifts [ 7 ]
+0.025
−39.7%
help
hurt
×
ImageNet-9 [ 8 ]
−0.021
−70.0%
hurt
hurt
✓
OxfordPets [ 13 ]
−0.024
−100.0%
hurt
hurt
✓
Appendix
Table 5: SIM and MARS’s measured relative change in worst-group accuracy vs. the uncompressed baseline, per dataset, on OpenCLIP ViT-L/14.
Dataset
SIM (stratified)
SIM (non-stratified)
Rel. Δ WG (%)
Correct
Waterbirds [ 4 ]
+0.0443
+0.0478
+25.6%
✓
UrbanCars [ 6 ]
+0.0513
+0.0549
+82.5%
✓
CelebA [ 5 ]
−0.0194
−0.0246
−19.6%
✓
MetaShifts [ 7 ]
+0.0249
+0.0233
−39.7%
×
ImageNet-9 [ 8 ]
−0.0213
−0.0213
−70.0%
✓
OxfordPets [ 13 ]
−0.0240
−0.0243
−100.0%
✓
Appendix
Table 6: Non-stratified vs. group-stratified SIM, 200 unstratified images per dataset. “Correct” indicates the non-stratified estimate’s sign matches the known direction of MARS’s effect on worst-group accuracy.
Component
Latency (ms)
Tokens
Baseline forward pass
15.7
256
First-pass forward for segmentation embeddings
≈ 15.7
256
Segmentation step (CPU, original)
28.5
-
Segmentation step (batched GPU, ours, b =128)
0.81
-
MARS second-pass forward (masked, pruned)
11.0
128
FastV forward pass (single pass, no masking)
12.4
128
Appendix
Table 7: Per-image latency (ms) on a single NVIDIA A100, OpenCLIP ViT-L/14. All rows are batch size 1 except the batched GPU segmentation row (batch 128; see Appendix H for the full batch-size sweep).
Batch size
Per-image-adaptive PCA
Shared PCA basis
1
9.06
7.65
8
1.85
0.95
32
0.98
0.25
64
0.87
0.14
128
0.81
unreliable †
Appendix
Table 8: Per-image segmentation cost (ms) by batch size, batched GPU reimplementation, OpenCLIP ViT-L/14 patch embeddings on real Waterbirds images (CUDA-event timing, 10 warmup + 30 measured runs). Reference: CPU (sklearn) path, 28.5ms/image at batch size 1 (Table 7 ); original unbatched single-image GPU attempt, 161ms/image.
Dataset
Per-image-adaptive PCA
Shared PCA basis
Mask IoU
Rel. Δ WG (%)
Mask IoU
Rel. Δ WG (%)
Waterbirds [ 4 ]
0.680
+6.5%
0.666
+21.8%
UrbanCars [ 6 ]
0.596
+0.0%
0.581
−9.4%
CelebA [ 5 ]
0.626
+2.3%
0.633
+4.5%
Appendix
Table 9: Correctness validation for the batched GPU segmentation, 300 held-out images per dataset, against the CPU (sklearn) reference. Rel. Δ WG is the batched variant’s worst-group accuracy relative to the CPU reference’s, on the same images.