Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space. To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring any image-text pairs. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space. Furthermore, driven by the empirical observations that CLIP's shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP's intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP's zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs. Project page at https://qzfm.github.io/salm_project_page/.
Figures & tables
Figure 1 : Comparison of feature visualizations and performance. Left: PCA visualization of features. In contrast to the fragmented and noisy outputs from CLIP and others, our method produces spatially coherent features with distinct semantic layouts. Right: The radar chart validates the superior performance of our method across multiple quantitative metrics.
Figure 2 : Visualization of CLIP’s shallow attention versus deep feature representation. CLIP demonstrates the capability to localize object textures ( e.g. , building edges) and detailed components ( e.g. , doors and windows) in its early layers, even though these details may not be effectively extracted in representations.
Figure 3 : Overview of SALM. First, the Dual-Matrix Alignment strategy ensures relative consistency in spatial relationships and magnitude distributions (Sec. 3.3 ). Subsequently, a latent masked modeling is employed to reconstruct the DINO features, thereby aggregating fine-grained information into global semantics (Sec. 3.4 ). Finally, reference regularization is applied to the source encoder to ensure the stability of the adaptation process within the original multimodal latent space (Sec. 3.5 ).
Figure 4 : Overview of SALM-Self.
Method
ImageNet
Zero-Shot
CIFAR-10
CIFAR-100
Pets
Caltech-101
RESISC45
PCam
DTD
EuroSAT
ImageNet-O
FER2013
ImageNetV2
Average
CLIP
76.55
94.92
74.33
93.68
83.43
63.78
60.72
55.69
61.50
32.75
49.11
70.89
67.35
RADIOv2.5 †
75.02
90.97
64.94
86.70
83.78
37.35
50.02
43.62
26.69
61.05
34.68
67.82
58.87
un 2 CLIP
71.25
91.84
69.95
91.74
85.45
57.33
61.78
52.93
62.46
36.15
51.49
64.91
66.00
KUEA
76.97
95.86
76.95
93.79
83.43
64.43
59.60
56.33
61.89
35.65
48.08
71.40
67.95
SALM-Self (Ours)
77.13
95.58
77.37
93.87
83.79
63.95
64.74
55.96
61.72
36.80
49.60
71.27
68.60
Table 1 : Zero-shot object recognition performance on various benchmarks. We report Top-1 accuracy (%) on ImageNet-1K and 11 other datasets. The best results are highlighted in bold . † indicates the model reproduced by us on ImageNet-1K.
Method
ZS. Fine-grained Tasks
Linear Probing Segmentation
SVHN
CLEVR Distance
CLEVR Counts
ADE20K
Cityscapes
VOC2012
COCO-Stuff
Context
CLIP
55.97
15.81
20.01
36.96
48.06
69.81
31.34
42.80
RADIOv2.5 †
41.25
19.45
15.60
42.35
51.82
75.31
35.11
48.73
un 2 CLIP
54.77
15.93
22.73
37.46
48.11
71.64
32.74
44.44
KUEA
57.74
15.95
20.81
36.38
46.74
69.02
31.46
42.45
SALM-Self (Ours)
57.92
15.83
22.78
39.23
49.43
72.33
34.60
45.61
Table 2 : Quantitative results on zero-shot fine-grained understanding and dense prediction. We report Top-1 accuracy (%) for zero-shot fine-grained understanding tasks, alongside mean IoU (mIoU) for semantic segmentation. † indicates the model reproduced by us on ImageNet-1K.
Method
ImageNet
CIFAR-10
CIFAR-100
Pets
SVHN
Caltech-101
RESISC45
PCam
DTD
EuroSAT
GTSRB
CLEVR D.
FER2013
CLEVR C.
Average
CLIP
80.09
97.42
85.49
94.73
78.5
95.97
96.01
84.46
80.58
97.14
92.76
55.88
71.92
75.62
84.76
RADIOv2.5 †
80.43
97.30
86.36
94.71
80.56
96.31
96.00
84.48
81.34
97.05
91.14
56.08
72.11
76.86
85.05
un 2 CLIP
80.02
97.02
84.51
93.76
81.63
96.53
95.82
84.14
80.37
97.70
92.5
60.59
72.21
77.43
85.30
KUEA
80.72
98.13
87.83
94.93
81.48
96.49
96.25
84.99
81.86
97.40
93.63
57.07
72.12
77.83
85.77
SALM-Self (Ours)
81.57
98.16
88.23
95.11
81.40
96.42
96.06
84.59
81.55
97.02
93.39
57.01
72.39
77.85
85.77
SALM (Ours)
80.94
98.17
88.28
95.20
81.61
96.45
96.90
85.49
81.86
97.73
93.66
57.88
72.50
79.75
86.17
Table 3 : Linear probing classification performance on various benchmarks. We report Top-1 accuracy (%) on 14 datasets. The best results are highlighted in bold . † indicates the model reproduced by us on ImageNet-1K.
Method
AI2D
POPE
TallyQA
VSR
RefCOCO
RefCOCO+
RefCOCOg
VQA v2
Average
LLaVA-7B
54.08
86.89
62.14
51.47
54.86
49.55
51.04
76.54
60.82
+ FT
53.95
86.90
62.45
52.70
66.42
58.64
58.93
77.23
64.65
un 2 CLIP
52.92
86.93
62.75
51.47
65.42
60.66
59.82
77.40
64.68
KUEA
53.33
87.30
61.21
52.04
66.94
60.74
60.48
77.27
64.92
SALM-Self (Ours)
54.19
86.92
62.53
53.44
66.99
60.74
61.37
76.99
65.40
SALM (Ours)
54.31
87.06
63.12
54.58
69.63
63.68
63.48
77.68
66.69
Table 4 : Performance evaluation on MLLM benchmarks. We compare the performance of different visual encoders when integrated into the LLaVA. “+FT” indicates that the model was fine-tuned using the same SFT data.
Method
ZS. Cls.
LP. Seg.
RESISC
PCam
SVHN
CLEVR Distance
ADE20K
CLIP
63.78
60.72
55.97
15.81
36.96
+FT
63.76
60.80
55.99
15.85
36.84
w/ DMA
63.48
63.00
57.51
16.13
42.07
w/ CGA
63.83
64.46
56.77
15.99
39.80
Ours
64.63
67.54
57.99
16.21
42.83
Table 5 : Ablation study on the effectiveness of proposed components. We analyze the impact of the DMA and CGA modules. “+FT” represents finetuning CLIP solely on ImageNet-1K with the text “This is a photo of {class name}”.
Figure 5 : Ablation analysis of data scale and masking ratio evaluated on ZS. datasets. Left: Performance scales consistently with increasing data volume. Right: SALM exhibits robustness across masking ratios, with 75% striking an optimal balance for capturing fine-grained semantics while retaining necessary spatial anchors.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value / Description
Dataset
ImageNet-1K (Train split)
Default Input Resolution
336×336
Global Batch Size
16
Cross-Guided Adapter Layers
4
Optimizer
AdamW
Learning Rate
1×10−5
Appendix
Table 6 : Detailed Hyperparameter Configurations.
Method
Settings
ImageNet
Zero-Shot Performance
Input Res.
Feat. Map
CIFAR-10
CIFAR-100
Pets
Caltech-101
RESISC45
PCam
DTD
EuroSAT
ImageNet-O
FER2013
ImageNetV2
Average
CLIP (ViT-L-14)
224 2
162
75.53
95.59
75.77
93.18
83.27
63.36
51.98
55.42
62.53
32.25
49.86
69.84
66.64
CLIP (ViT-L-14-336)
336 2
242
76.55
94.92
74.33
93.68
83.43
63.78
60.72
55.69
61.50
32.75
49.11
70.89
67.35
SigLIP
224 2
162
82.02
96.74
84.12
95.33
86.05
69.68
50.58
71.01
62.98
29.55
49.83
76.06
70.18
CLIP (L-14) + MAE
224 2 / 224 2
162/142
76.13
95.63
77.29
93.51
84.29
63.48
50.4
54.95
61.85
36.6
50.68
70.23
67.17
CLIP (L-14) + DINOv2
224 2 / 224 2
162 / 162
76.00
96.42
77.19
93.62
84.38
64.40
53.13
56.06
62.59
36.75
50.15
69.86
67.69
Appendix
Table 7 : Generalization and flexibility analysis. We report the performance across various model combinations. The columns Input Res. and Feat. Map denotes the input resolution and the extracted feature map size, respectively.
Method
ImageNet
Zero-Shot
CIFAR-10
CIFAR-100
Pets
Caltech-101
RESISC45
PCam
DTD
EuroSAT
ImageNet-O
FER2013
ImageNetV2
Average
CLIP
76.55
94.92
74.33
93.68
83.43
63.78
60.72
55.69
61.50
32.75
49.11
70.89
67.35
Simple Distillation
77.06
95.94
77.40
93.13
83.09
64.39
58.26
55.85
60.92
37.90
49.27
71.39
67.96
Ours w/o 3-stage
77.02
96.30
77.24
93.97
84.07
64.10
62.89
55.95
62.16
36.65
49.81
71.50
68.60
Ours w/ 3-stage
77.21
96.38
77.87
94.09
84.04
64.63
67.54
56.27
62.19
37.05
50.11
71.53
69.25
Appendix
Table 8 : Ablation analysis of simple feature distillation strategy and three-stage curriculum strategy on zero-shot benchmarks. We report Top-1 accuracy (%) on ImageNet-1K and 11 other datasets. The best results are highlighted in bold .
Method
ZS. Fine-grained Tasks
LP. Seg.
SVHN
CLEVR Distance
CLEVR Counts
ADE20K
CLIP
55.97
15.81
20.01
36.96
Simple Distillation
51.87
15.92
22.63
43.09
Ours w/o 3-stage
57.95
15.93
22.78
42.29
Ours w/ 3-stage
57.99
16.21
22.81
42.83
Appendix
Table 9 : Ablation analysis of simple feature distillation strategy and three-stage curriculum strategy on zero-shot fine-grained understanding and dense prediction. We report Top-1 accuracy (%) for zero-shot fine-grained understanding tasks, alongside mean IoU (mIoU) for semantic segmentation.
Figure 6 : Ablation analysis of the training schedule. We compare the weighted total objective loss curves of models trained with and without the three-stage curriculum strategy. Both settings successfully converge, showing that the model can still train and converge stably when the curriculum strategy is removed.
Method
ZS. Fine-grained Tasks
LP. Seg.
SVHN
CLEVR Distance
CLEVR Counts
ADE20K
w/o DMA
56.77
15.99
20.99
39.80
w/ K
56.73
16.05
22.17
40.98
w/ D
56.87
16.01
22.39
39.97
w/ DMA
57.99
16.21
22.81
42.83
Appendix
Table 10 : Ablation analysis of the Dual-Matrix Alignment (DMA) strategy on various downstream tasks. We decouple the effects of the Spatial Relation Matrix ( K ) and the Energy Difference Matrix ( D ).
Reference Layer
ImageNet-1K
Avg.
2
76.90
68.57
4
77.06
68.35
6
77.13
68.60
8
77.07
68.43
10
76.95
68.58
Appendix
Table 11 : Ablation study on the reference layer selection for SALM-Self.
Method
VOC2012
Context
COCO-Stuff
Cityscapes
ADE20K
CLIP
14.79
4.06
2.04
1.28
1.43
Simple Distillation
2.07
0.29
0.05
0.16
0.02
DeCLIP
85.16
39.37
28.75
33.13
21.92
SALM (Ours)
17.03
7.75
4.22
3.08
3.34
Appendix
Table 12 : Open-vocabulary segmentation results. We report mIoU (%) on five segmentation benchmarks.
Method
Average
CLIP
0.0
20.0
40.0
20.0
6.7
20.0
33.3
6.7
33.3
20.0
KUEA
6.6
26.7
40.0
13.3
6.7
40.0
26.6
13.3
20.0
21.5
Ours
13.3
20.0
46.7
13.3
13.3
53.3
33.3
13.3
26.7
25.9
Appendix
Table 13 : Performance of CLIP on MMVP-VLM benchmark. Symbols for visual patterns are inherited: : Orientation and Direction, : Presence of Specific Features, : State and Condition, : Quantity and Count, : Positional and Relational Context, : Color and Appearance, : Structural and Physical Characteristics, : Texts, : Viewpoint and Perspective.
Method
Image-to-Text Retrieval
Text-to-Image Retrieval
Flickr30K
MSCOCO
Flickr30K
MSCOCO
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
CLIP
87.30
98.29
99.30
57.92
81.20
87.86
67.34
89.02
93.28
37.07
61.64
71.51
KUEA
87.30
98.10
99.30
58.26
81.24
88.27
68.70
89.80
93.83
38.05
62.64
72.60
Ours
86.70
98.40
99.10
59.12
82.08
88.36
68.88
89.80
94.20
38.54
63.25
73.10
Appendix
Table 14 : Zero-shot image-text retrieval results on Flickr30K and MSCOCO benchmarks. We report Recall@K (K=1, 5, 10) for both Image-to-Text and Text-to-Image retrieval tasks.
Figure 7 : The t-SNE visualization of global semantic features. Compared to other methods, our approach generates more compact and well-separated clusters, demonstrating superior capability in learning discriminative global representations.
Figure 8 : PCA-based visualizations. Unlike other enhancement methods, our approach excels at capturing fine-grained semantic nuances and preserving discriminative local details.
Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning a shared image-text embedding space from large-scale contrastive pre-training. However, its image-level objective aligns text with a CLS-derived global representation, leaving local vision-language correspondence only indirectly constrained. Existing methods either introduce additional supervision, external models, or task-specific adaptation, while training-free approaches mainly recover dense responses from existing patch features without examining where local semantics become most accessible within CLIP. We introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence by isolating the patch-specific terms written into the CLS attention output. TraceCLIP further converts contribution-derived semantic responses into a semantic-geodesic topology gate that calibrates final-layer patch affinity for dense feature reconstruction. Diagnostic experiments show that these contribution features exhibit strong local semantic discrimination and text-conditioned spatial alignment. On eight zero-shot semantic segmentation benchmarks, TraceCLIP achieves gains of 1.3 to 4.5 points in average mIoU over the strongest prior training-free methods across both backbones and background settings, without additional training, external vision foundation models, or region-level supervision. More broadly, these findings suggest that spatially localized semantics may remain accessible within the internal construction of globally aligned representations.
Contrastive Language-Image Pre-training (CLIP) has been shown to have limitations in its fine-grained dense feature representation, due to its pre-training focusing on matching the whole image to a text description. Considering the large data and computational burden in pre-training a vision-language model from scratch, a series of works aim to enhance the fine-grained ability of CLIP through a fine-tuning scheme. However, existing works suffer from a variety of limitations: additional region annotations are usually required, which limits the semantic diversity due to the predefined categories and leads to a large effort to process the training data; and they usually sacrifice CLIP's original ability for global visual representation. To bypass these limitations, we propose SFF-CLIP (Self-annotated Fine-grained Fine-tuning for CLIP), which only uses image-text pairs as input to boost the fine-grained representation ability in the CLIP fine-tuning, while maintaining the global visual-semantic consistency. Concretely, a run-time region-phrase alignment scheme is designed, which obtains concept phrases from the input sentence, and aligns them with corresponding extracted region-based features using text-specific heat maps. Extensive experiments demonstrate that SFF-CLIP leads to significant performance improvements on fine-grained dense feature representation, as well as maintaining the performance of the original CLIP on image-level tasks. Code will be released later.
Chenyang Zhao, Wei Lin, Antoni B. Chan +1
Department of Computer Science, City University of Hong Kong · Division of Social Science, Hong Kong University of Science & Technology
Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis suggests this architecture suffers from an alignment challenge: visual features remain distant from the text space in the initial layers of the LLM, forcing the model to waste critical depth~\cite{zhang-etal-2024-investigating,artzy-schwartz-2024-attend} on superficial modality alignment rather than deep understanding and complex reasoning. In this work, we propose Deep Pre-Alignment (DPA), a novel architecture that replaces the standard ViT encoder with a small VLM as perceiver, ensuring visual features are deeply aligned with the text space of the target large language model. Comprehensive experiments demonstrate the effectiveness of DPA. On the 4B parameter scale, DPA outperforms baselines by 1.9 points across 8 multimodal benchmarks, with gains widening to 3.0 points at the 32B scale. Moreover, by offloading alignment to the perceiver, DPA achieves a 32.9% reduction in language capability forgetting over 3 text benchmarks. We further demonstrate that these gains are consistent across different LLM families including Qwen3 and LLaMA 3.2, highlighting the generality of our approach. Beyond performance, DPA also offers a seamless upgrade path for current VLM development, requiring only a modular replacement for the visual encoder with marginal computation overhead.
Tianyu Yu, Kechen Fang, Zihao Wan +5
1Tsinghua University · 2Taobao & Tmall Group of Alibaba · 3Shanghai Qi Zhi Institute.