CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach that keeps every CLIP parameter frozen. A compact MaskAdapter generates per-head, per-layer attention biases from an input mask and injects them into the frozen self-attention layers, steering attention toward the target region. Crucially, because the backbone remains strictly untouched, LAS-CLIP seamlessly reverts to vanilla CLIP when no mask is provided, preserving its foundational zero-shot capabilities. With approximately 116K to 145K trainable parameters and 100K training samples on two T4 GPUs, LAS-CLIP achieves competitive or superior results compared to Alpha-CLIP on ImageNet-S zero-shot classification and RefCOCO referring expression comprehension, despite the latter fine-tuning its entire encoder on millions of samples. Qualitative analysis further confirms stronger representational fidelity under incorrect masks and in downstream generation.
Figures & tables
Figure 1 : Overview of LAS-CLIP and the MaskAdapter architecture. The input image and a region mask are processed by the visual encoder and the MaskAdapter to produce region-aware features. The MaskAdapter consists of a mask encoder, a token encoder with layer conditioning, and a gated interaction module that computes attention biases to steer the frozen CLIP blocks.
Method
B/16
L/14
Top-1
Top-5
Top-1
Top-5
Original CLIP [ clip ]
66.5
88.9
73.5
91.6
MaskAdaptedCLIP [ maskadaptedclip ]
57.9
79.1
63.5
86.3
Red Circle [ redcircle ]
65.4
88.7
73.4
92.1
FALIP [ falip ]
68.2
89.5
74.9
91.9
Alpha-CLIP [ alphaclip ]
68.9
90.5
77.4
94.5
Table 1 : Zero-shot classification on ImageNet-S validation set. Mean per-class accuracy in percentage is reported. The best results are in bold and the second best results are underlined .
Method
RefCOCO
RefCOCO+
RefCOCOg
val
testA
testB
val
testA
testB
val
test
CPT [ cpt ]
32.2
36.1
30.3
31.9
35.2
28.8
36.7
36.5
ReCLIP [ reclip ]
45.8
46.1
47.1
47.9
50.1
45.1
59.3
59.0
Red Circle [ redcircle ]
49.8
58.6
39.9
55.3
63.9
45.4
59.4
58.9
FALIP [ falip ]
43.8
45.6
42.4
44.1
46.5
40.0
48.8
48.7
Alpha-CLIP [ alphaclip ]
55.7
61.1
50.3
55.6
62.7
46.4
61.2
62.0
Table 2 : Zero-shot referring expression comprehension accuracy in percentage on RefCOCO, RefCOCO+, and RefCOCOg. Following our primary baseline Alpha-CLIP, LAS-CLIP utilizes an ensemble of ViT-B/16 and ViT-L/14 backbones to ensure a strictly fair comparison, FALIP also uses the same setting. Note that prior baselines also report ensemble results using their strongest respective configurations (e.g., ReCLIP ensembles RN50x16 and ViT-B/32, Red Circle ensembles RN50x16 and ViT-L/14@336px). The best results are in bold and the second best results are underlined .
Layers
ImageNet-S Top-1
RefCOCO
B/16
L/14
val
testA
testB
K=1
70.7
78.1
54.9
60.9
49.6
K=2
70.0
77.7
55.8
62.2
50.3
K=3
69.9
77.6
57.5
64.1
50.3
K=4
69.6
76.6
57.3
64.3
50.3
Table 3 : Effect of the number of adapter layers K on ImageNet-S top-1 accuracy (ViT-B/16 and ViT-L/14) and RefCOCO accuracy (ViT-B/16 + ViT-L/14 ensemble) in percentage. The best results are in bold and the second best results are underlined .
Lid
ImageNet-S Top-1
RefCOCO
B/16
L/14
val
testA
testB
w/
69.6
76.6
57.3
64.3
50.3
w/o
66.4
76.1
58.2
65.2
50.6
Table 4 : Effect of identity regularization on ImageNet-S top-1 accuracy (per-backbone) and RefCOCO accuracy (ViT-B/16 + ViT-L/14 ensemble) in percentage, using K=4 layers. Best results are in bold .
Variant
ImageNet-S
RefCOCO
Top-1
Top-5
val
testA
testB
LAS-CLIP (default)
69.9
90.9
54.7
60.7
48.7
CLS-only
69.9
91.0
53.5
60.2
48.1
Head-agnostic
70.4
91.0
52.2
58.3
47.5
No token cond.
69.9
90.8
52.4
58.7
48.1
Table 5 : Adapter architecture ablation on ViT-B/16. Best in bold .
Figure 2 : Attention heatmaps visualized via text-based decomposition [ clipdecomp ] . Each row shows a different text query. The left group shows results using correct masks, and the right group shows results using incorrect masks. LAS-CLIP maintains better text-aligned attention under incorrect mask inputs compared to Alpha-CLIP.
Figure 3 : Subject-driven image generation using BLIP-Diffusion [ blipdiffusion ] with different CLIP backbones (ViT-L/14). Each row shows an input subject with its mask and the outputs generated by the baseline CLIP, Alpha-CLIP [ alphaclip ] , and LAS-CLIP. LAS-CLIP better preserves the visual identity of the input subject.
Figure S1 : Additional attention heatmap visualizations. Each row shows a different text query. The left group shows results using correct masks, and the right group shows results using incorrect masks.
Figure S2 : Additional subject-driven generation results using BLIP-Diffusion [ blipdiffusion ] . Each row shows an input subject, its mask, and the generated outputs.
Figure S3 : Mask robustness comparison on ImageNet-S (ViT-L/14). LAS-CLIP degrades more gracefully than Alpha-CLIP under dilation and spatial shift perturbations.
Method / Mask Source
ViT-B/16
ViT-L/14
Top-1
Top-5
Top-1
Top-5
Alpha-CLIP (SAM)
68.9
90.5
77.4
94.5
LAS-CLIP (YOLOE)
69.9
90.9
77.6
94.0
LAS-CLIP (SAM)
70.8
91.3
78.1
94.1
Table S1 : Training data source ablation on ImageNet-S. Top-1/Top-5 accuracy (%). Best in bold .
Method / Source
RefCOCO
RefCOCO+
RefCOCOg
val
testA
testB
val
testA
testB
val
test
Alpha-CLIP (SAM)
55.7
61.1
50.3
55.6
62.7
46.4
61.2
62.0
LAS-CLIP (YOLOE)
57.5
64.1
50.3
57.4
66.0
46.6
61.7
61.7
LAS-CLIP (SAM)
56.7
64.0
50.1
56.9
65.8
46.7
61.2
61.2
Table S2 : Training data source ablation on RefCOCO / RefCOCO+ / RefCOCOg. Accuracy (%). Best in bold .
Method
Batch
FPS
± std
Lat. (ms)
± std
Mem (MB)
GFLOPs
CLIP
1
70.7
3.1
14.2
0.6
974
81.01
CLIP
4
85.2
0.3
46.9
0.2
1000
20.25
CLIP
16
80.4
1.1
199.1
2.8
1096
5.06
CLIP
64
82.2
1.5
778.5
14.1
1480
1.27
Alpha-CLIP
1
58.8
0.4
17.0
0.1
1925
81.06
Alpha-CLIP
4
65.6
0.4
61.0
0.4
1959
20.27
Table S3 : Inference efficiency on ViT-L/14 (single T4 GPU). FPS, latency, memory, and GFLOPs are reported across batch sizes.
Institute of Artificial Intelligence, Huazhong University of Science and Technology, Wuhan, China · School of Computer Science and Technology, Huazhong University of Science and Technology, Wuhan, China.