CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach that keeps every CLIP parameter frozen. A compact MaskAdapter generates per-head, per-layer attention biases from an input mask and injects them into the frozen self-attention layers, steering attention toward the target region. Crucially, because the backbone remains strictly untouched, LAS-CLIP seamlessly reverts to vanilla CLIP when no mask is provided, preserving its foundational zero-shot capabilities. With approximately 116K to 145K trainable parameters and 100K training samples on two T4 GPUs, LAS-CLIP achieves competitive or superior results compared to Alpha-CLIP on ImageNet-S zero-shot classification and RefCOCO referring expression comprehension, despite the latter fine-tuning its entire encoder on millions of samples. Qualitative analysis further confirms stronger representational fidelity under incorrect masks and in downstream generation.
Figures & tables
Figure 1 : Overview of LAS-CLIP and the MaskAdapter architecture. The input image and a region mask are processed by the visual encoder and the MaskAdapter to produce region-aware features. The MaskAdapter consists of a mask encoder, a token encoder with layer conditioning, and a gated interaction module that computes attention biases to steer the frozen CLIP blocks.
Method
B/16
L/14
Top-1
Top-5
Top-1
Top-5
Original CLIP [ clip ]
66.5
88.9
73.5
91.6
MaskAdaptedCLIP [ maskadaptedclip ]
57.9
79.1
63.5
86.3
Red Circle [ redcircle ]
65.4
88.7
73.4
92.1
FALIP [ falip ]
68.2
89.5
74.9
91.9
Alpha-CLIP [ alphaclip ]
68.9
90.5
77.4
94.5
Table 1 : Zero-shot classification on ImageNet-S validation set. Mean per-class accuracy in percentage is reported. The best results are in bold and the second best results are underlined .
Method
RefCOCO
RefCOCO+
RefCOCOg
val
testA
testB
val
testA
testB
val
test
CPT [ cpt ]
32.2
36.1
30.3
31.9
35.2
28.8
36.7
36.5
ReCLIP [ reclip ]
45.8
46.1
47.1
47.9
50.1
45.1
59.3
59.0
Red Circle [ redcircle ]
49.8
58.6
39.9
55.3
63.9
45.4
59.4
58.9
FALIP [ falip ]
43.8
45.6
42.4
44.1
46.5
40.0
48.8
48.7
Alpha-CLIP [ alphaclip ]
55.7
61.1
50.3
55.6
62.7
46.4
61.2
62.0
Table 2 : Zero-shot referring expression comprehension accuracy in percentage on RefCOCO, RefCOCO+, and RefCOCOg. Following our primary baseline Alpha-CLIP, LAS-CLIP utilizes an ensemble of ViT-B/16 and ViT-L/14 backbones to ensure a strictly fair comparison, FALIP also uses the same setting. Note that prior baselines also report ensemble results using their strongest respective configurations (e.g., ReCLIP ensembles RN50x16 and ViT-B/32, Red Circle ensembles RN50x16 and ViT-L/14@336px). The best results are in bold and the second best results are underlined .
Layers
ImageNet-S Top-1
RefCOCO
B/16
L/14
val
testA
testB
K=1
70.7
78.1
54.9
60.9
49.6
K=2
70.0
77.7
55.8
62.2
50.3
K=3
69.9
77.6
57.5
64.1
50.3
K=4
69.6
76.6
57.3
64.3
50.3
Table 3 : Effect of the number of adapter layers K on ImageNet-S top-1 accuracy (ViT-B/16 and ViT-L/14) and RefCOCO accuracy (ViT-B/16 + ViT-L/14 ensemble) in percentage. The best results are in bold and the second best results are underlined .
Lid
ImageNet-S Top-1
RefCOCO
B/16
L/14
val
testA
testB
w/
69.6
76.6
57.3
64.3
50.3
w/o
66.4
76.1
58.2
65.2
50.6
Table 4 : Effect of identity regularization on ImageNet-S top-1 accuracy (per-backbone) and RefCOCO accuracy (ViT-B/16 + ViT-L/14 ensemble) in percentage, using K=4 layers. Best results are in bold .
Variant
ImageNet-S
RefCOCO
Top-1
Top-5
val
testA
testB
LAS-CLIP (default)
69.9
90.9
54.7
60.7
48.7
CLS-only
69.9
91.0
53.5
60.2
48.1
Head-agnostic
70.4
91.0
52.2
58.3
47.5
No token cond.
69.9
90.8
52.4
58.7
48.1
Table 5 : Adapter architecture ablation on ViT-B/16. Best in bold .
Figure 2 : Attention heatmaps visualized via text-based decomposition [ clipdecomp ] . Each row shows a different text query. The left group shows results using correct masks, and the right group shows results using incorrect masks. LAS-CLIP maintains better text-aligned attention under incorrect mask inputs compared to Alpha-CLIP.
Figure 3 : Subject-driven image generation using BLIP-Diffusion [ blipdiffusion ] with different CLIP backbones (ViT-L/14). Each row shows an input subject with its mask and the outputs generated by the baseline CLIP, Alpha-CLIP [ alphaclip ] , and LAS-CLIP. LAS-CLIP better preserves the visual identity of the input subject.
Figure S1 : Additional attention heatmap visualizations. Each row shows a different text query. The left group shows results using correct masks, and the right group shows results using incorrect masks.
Figure S2 : Additional subject-driven generation results using BLIP-Diffusion [ blipdiffusion ] . Each row shows an input subject, its mask, and the generated outputs.
Figure S3 : Mask robustness comparison on ImageNet-S (ViT-L/14). LAS-CLIP degrades more gracefully than Alpha-CLIP under dilation and spatial shift perturbations.
Method / Mask Source
ViT-B/16
ViT-L/14
Top-1
Top-5
Top-1
Top-5
Alpha-CLIP (SAM)
68.9
90.5
77.4
94.5
LAS-CLIP (YOLOE)
69.9
90.9
77.6
94.0
LAS-CLIP (SAM)
70.8
91.3
78.1
94.1
Table S1 : Training data source ablation on ImageNet-S. Top-1/Top-5 accuracy (%). Best in bold .
Method / Source
RefCOCO
RefCOCO+
RefCOCOg
val
testA
testB
val
testA
testB
val
test
Alpha-CLIP (SAM)
55.7
61.1
50.3
55.6
62.7
46.4
61.2
62.0
LAS-CLIP (YOLOE)
57.5
64.1
50.3
57.4
66.0
46.6
61.7
61.7
LAS-CLIP (SAM)
56.7
64.0
50.1
56.9
65.8
46.7
61.2
61.2
Table S2 : Training data source ablation on RefCOCO / RefCOCO+ / RefCOCOg. Accuracy (%). Best in bold .
Method
Batch
FPS
± std
Lat. (ms)
± std
Mem (MB)
GFLOPs
CLIP
1
70.7
3.1
14.2
0.6
974
81.01
CLIP
4
85.2
0.3
46.9
0.2
1000
20.25
CLIP
16
80.4
1.1
199.1
2.8
1096
5.06
CLIP
64
82.2
1.5
778.5
14.1
1480
1.27
Alpha-CLIP
1
58.8
0.4
17.0
0.1
1925
81.06
Alpha-CLIP
4
65.6
0.4
61.0
0.4
1959
20.27
Table S3 : Inference efficiency on ViT-L/14 (single T4 GPU). FPS, latency, memory, and GFLOPs are reported across batch sizes.
Large-scale pre-trained vision-language models like CLIP demonstrate remarkable zero-shot performance across diverse tasks. However, fine-tuning these models to improve downstream performance often degrades robustness against distribution shifts. Recent approaches have attempted to mitigate this trade-off, but often rely on computationally expensive text-guidance. We propose a novel method for robust fine-tuning, SAE-FT, which operates only on the model's visual representations. SAE-FT regularizes changes to these representations by penalizing the addition and removal of semantically meaningful features identified by a Sparse Autoencoder trained on the pre-trained model. This constraint prevents catastrophic forgetting and makes the fine-tuning process interpretable, enabling direct analysis of semantic changes. SAE-FT is both mechanistically transparent and computationally efficient, matching or exceeding state-of-the-art performance on ImageNet and its associated distribution shift benchmarks. Code is publicly available at: https://github.com/Fabian-Mor/sae-ft.
Contrastively trained vision-language models such as CLIP provide strong zero-shot transfer by aligning images and text in a shared embedding space. However, adapting these models to downstream tasks without degrading their open-vocabulary generalization remains challenging. Existing parameter-efficient adaptation methods typically improve task specialization through learned prompts, adapters, or multimodal transformations, where adaptation capacity is primarily expressed through additional trainable parameters. Inspired by recent latent reasoning methods in language models, we investigate a complementary perspective: can adaptation emerge from iterative reasoning on latent representations rather than from increasing parameter count alone? We introduce PERL (Parameter-Efficient Reasoning in CLIP Latent Space), a lightweight adaptation framework that augments a frozen CLIP model with a compact shared reasoning module applied recurrently across refinement steps. At each step, PERL generates a latent reasoning token conditioned on the current representation and injects it into an intermediate encoder layer, progressively refining higher-level semantic representations while preserving CLIP's pretrained multimodal structure. Across 15 benchmarks spanning base-to-novel generalization, cross-dataset transfer, and out-of-distribution ImageNet variants, PERL achieves the best parameter-performance trade-off among the compared methods under a fast-adaptation few-shot setting, combining strong novel-class accuracy and competitive transfer performance with only about 6K trainable parameters, up to 817x fewer than the largest compared approach. Overall, our results suggest that iterative latent reasoning provides a complementary adaptation mechanism to parameter scaling in discriminative vision-language models.
Vision-Language Models (VLMs) such as CLIP demonstrate strong zero-shot generalization, but their performance significantly degrades in cross-domain scenarios with scarce target-domain training data (Cross-Domain Few-Shot Learning, CDFSL). In this paper, we focus on the target-domain few-shot finetuning in the CLIP-based CDFSL task. Prevailing finetuning paradigms uniformly align all image patch tokens with their corresponding textual embeddings. However, we find a counterintuitive phenomenon: actively pushing away certain low-similarity image tokens, termed "tail tokens", from their textual embeddings consistently improves target-domain performance. We delve into this phenomenon and provide a novel interpretation: under great domain shifts and scarce training data, the model can hardly extract semantic information from visual inputs; therefore, the common belief of alignment is valid only for tokens already containing sufficient semantic information; for tail tokens, forcing the alignment would lead to excessive overfitting to the scarce training, while breaking the alignment is more useful. Motivated by this, we propose Adaptive Tail-Head Alignment (ATHA), a novel fine-tuning strategy for CLIP that transforms the conventional uniform alignment paradigm to an adaptive alignment paradigm, with both alignment strengthening and weakening. Extensive experiments on four challenging CDFSL benchmarks validate our state-of-the-art performance. Our code is available at https://github.com/shuaiyi308/ATHA.
Shuai Yi, Yixiong Zou, Yuhua Li +1
Institute of Artificial Intelligence, Huazhong University of Science and Technology, Wuhan, China · School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China.