Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space. To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring any image-text pairs. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space. Furthermore, driven by the empirical observations that CLIP's shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP's intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP's zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs. Project page at https://qzfm.github.io/salm_project_page/.
Figures & tables
Figure 1 : Comparison of feature visualizations and performance. Left: PCA visualization of features. In contrast to the fragmented and noisy outputs from CLIP and others, our method produces spatially coherent features with distinct semantic layouts. Right: The radar chart validates the superior performance of our method across multiple quantitative metrics.
Figure 2 : Visualization of CLIP’s shallow attention versus deep feature representation. CLIP demonstrates the capability to localize object textures ( e.g. , building edges) and detailed components ( e.g. , doors and windows) in its early layers, even though these details may not be effectively extracted in representations.
Figure 3 : Overview of SALM. First, the Dual-Matrix Alignment strategy ensures relative consistency in spatial relationships and magnitude distributions (Sec. 3.3 ). Subsequently, a latent masked modeling is employed to reconstruct the DINO features, thereby aggregating fine-grained information into global semantics (Sec. 3.4 ). Finally, reference regularization is applied to the source encoder to ensure the stability of the adaptation process within the original multimodal latent space (Sec. 3.5 ).
Figure 4 : Overview of SALM-Self.
Method
ImageNet
Zero-Shot
CIFAR-10
CIFAR-100
Pets
Caltech-101
RESISC45
PCam
DTD
EuroSAT
ImageNet-O
FER2013
ImageNetV2
Average
CLIP
76.55
94.92
74.33
93.68
83.43
63.78
60.72
55.69
61.50
32.75
49.11
70.89
67.35
RADIOv2.5 †
75.02
90.97
64.94
86.70
83.78
37.35
50.02
43.62
26.69
61.05
34.68
67.82
58.87
un 2 CLIP
71.25
91.84
69.95
91.74
85.45
57.33
61.78
52.93
62.46
36.15
51.49
64.91
66.00
KUEA
76.97
95.86
76.95
93.79
83.43
64.43
59.60
56.33
61.89
35.65
48.08
71.40
67.95
SALM-Self (Ours)
77.13
95.58
77.37
93.87
83.79
63.95
64.74
55.96
61.72
36.80
49.60
71.27
68.60
Table 1 : Zero-shot object recognition performance on various benchmarks. We report Top-1 accuracy (%) on ImageNet-1K and 11 other datasets. The best results are highlighted in bold . † indicates the model reproduced by us on ImageNet-1K.
Method
ZS. Fine-grained Tasks
Linear Probing Segmentation
SVHN
CLEVR Distance
CLEVR Counts
ADE20K
Cityscapes
VOC2012
COCO-Stuff
Context
CLIP
55.97
15.81
20.01
36.96
48.06
69.81
31.34
42.80
RADIOv2.5 †
41.25
19.45
15.60
42.35
51.82
75.31
35.11
48.73
un 2 CLIP
54.77
15.93
22.73
37.46
48.11
71.64
32.74
44.44
KUEA
57.74
15.95
20.81
36.38
46.74
69.02
31.46
42.45
SALM-Self (Ours)
57.92
15.83
22.78
39.23
49.43
72.33
34.60
45.61
Table 2 : Quantitative results on zero-shot fine-grained understanding and dense prediction. We report Top-1 accuracy (%) for zero-shot fine-grained understanding tasks, alongside mean IoU (mIoU) for semantic segmentation. † indicates the model reproduced by us on ImageNet-1K.
Method
ImageNet
CIFAR-10
CIFAR-100
Pets
SVHN
Caltech-101
RESISC45
PCam
DTD
EuroSAT
GTSRB
CLEVR D.
FER2013
CLEVR C.
Average
CLIP
80.09
97.42
85.49
94.73
78.5
95.97
96.01
84.46
80.58
97.14
92.76
55.88
71.92
75.62
84.76
RADIOv2.5 †
80.43
97.30
86.36
94.71
80.56
96.31
96.00
84.48
81.34
97.05
91.14
56.08
72.11
76.86
85.05
un 2 CLIP
80.02
97.02
84.51
93.76
81.63
96.53
95.82
84.14
80.37
97.70
92.5
60.59
72.21
77.43
85.30
KUEA
80.72
98.13
87.83
94.93
81.48
96.49
96.25
84.99
81.86
97.40
93.63
57.07
72.12
77.83
85.77
SALM-Self (Ours)
81.57
98.16
88.23
95.11
81.40
96.42
96.06
84.59
81.55
97.02
93.39
57.01
72.39
77.85
85.77
SALM (Ours)
80.94
98.17
88.28
95.20
81.61
96.45
96.90
85.49
81.86
97.73
93.66
57.88
72.50
79.75
86.17
Table 3 : Linear probing classification performance on various benchmarks. We report Top-1 accuracy (%) on 14 datasets. The best results are highlighted in bold . † indicates the model reproduced by us on ImageNet-1K.
Method
AI2D
POPE
TallyQA
VSR
RefCOCO
RefCOCO+
RefCOCOg
VQA v2
Average
LLaVA-7B
54.08
86.89
62.14
51.47
54.86
49.55
51.04
76.54
60.82
+ FT
53.95
86.90
62.45
52.70
66.42
58.64
58.93
77.23
64.65
un 2 CLIP
52.92
86.93
62.75
51.47
65.42
60.66
59.82
77.40
64.68
KUEA
53.33
87.30
61.21
52.04
66.94
60.74
60.48
77.27
64.92
SALM-Self (Ours)
54.19
86.92
62.53
53.44
66.99
60.74
61.37
76.99
65.40
SALM (Ours)
54.31
87.06
63.12
54.58
69.63
63.68
63.48
77.68
66.69
Table 4 : Performance evaluation on MLLM benchmarks. We compare the performance of different visual encoders when integrated into the LLaVA. “+FT” indicates that the model was fine-tuned using the same SFT data.
Method
ZS. Cls.
LP. Seg.
RESISC
PCam
SVHN
CLEVR Distance
ADE20K
CLIP
63.78
60.72
55.97
15.81
36.96
+FT
63.76
60.80
55.99
15.85
36.84
w/ DMA
63.48
63.00
57.51
16.13
42.07
w/ CGA
63.83
64.46
56.77
15.99
39.80
Ours
64.63
67.54
57.99
16.21
42.83
Table 5 : Ablation study on the effectiveness of proposed components. We analyze the impact of the DMA and CGA modules. “+FT” represents finetuning CLIP solely on ImageNet-1K with the text “This is a photo of {class name}”.
Figure 5 : Ablation analysis of data scale and masking ratio evaluated on ZS. datasets. Left: Performance scales consistently with increasing data volume. Right: SALM exhibits robustness across masking ratios, with 75% striking an optimal balance for capturing fine-grained semantics while retaining necessary spatial anchors.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value / Description
Dataset
ImageNet-1K (Train split)
Default Input Resolution
336×336
Global Batch Size
16
Cross-Guided Adapter Layers
4
Optimizer
AdamW
Learning Rate
1×10−5
Appendix
Table 6 : Detailed Hyperparameter Configurations.
Method
Settings
ImageNet
Zero-Shot Performance
Input Res.
Feat. Map
CIFAR-10
CIFAR-100
Pets
Caltech-101
RESISC45
PCam
DTD
EuroSAT
ImageNet-O
FER2013
ImageNetV2
Average
CLIP (ViT-L-14)
224 2
162
75.53
95.59
75.77
93.18
83.27
63.36
51.98
55.42
62.53
32.25
49.86
69.84
66.64
CLIP (ViT-L-14-336)
336 2
242
76.55
94.92
74.33
93.68
83.43
63.78
60.72
55.69
61.50
32.75
49.11
70.89
67.35
SigLIP
224 2
162
82.02
96.74
84.12
95.33
86.05
69.68
50.58
71.01
62.98
29.55
49.83
76.06
70.18
CLIP (L-14) + MAE
224 2 / 224 2
162/142
76.13
95.63
77.29
93.51
84.29
63.48
50.4
54.95
61.85
36.6
50.68
70.23
67.17
CLIP (L-14) + DINOv2
224 2 / 224 2
162 / 162
76.00
96.42
77.19
93.62
84.38
64.40
53.13
56.06
62.59
36.75
50.15
69.86
67.69
Appendix
Table 7 : Generalization and flexibility analysis. We report the performance across various model combinations. The columns Input Res. and Feat. Map denotes the input resolution and the extracted feature map size, respectively.
Method
ImageNet
Zero-Shot
CIFAR-10
CIFAR-100
Pets
Caltech-101
RESISC45
PCam
DTD
EuroSAT
ImageNet-O
FER2013
ImageNetV2
Average
CLIP
76.55
94.92
74.33
93.68
83.43
63.78
60.72
55.69
61.50
32.75
49.11
70.89
67.35
Simple Distillation
77.06
95.94
77.40
93.13
83.09
64.39
58.26
55.85
60.92
37.90
49.27
71.39
67.96
Ours w/o 3-stage
77.02
96.30
77.24
93.97
84.07
64.10
62.89
55.95
62.16
36.65
49.81
71.50
68.60
Ours w/ 3-stage
77.21
96.38
77.87
94.09
84.04
64.63
67.54
56.27
62.19
37.05
50.11
71.53
69.25
Appendix
Table 8 : Ablation analysis of simple feature distillation strategy and three-stage curriculum strategy on zero-shot benchmarks. We report Top-1 accuracy (%) on ImageNet-1K and 11 other datasets. The best results are highlighted in bold .
Method
ZS. Fine-grained Tasks
LP. Seg.
SVHN
CLEVR Distance
CLEVR Counts
ADE20K
CLIP
55.97
15.81
20.01
36.96
Simple Distillation
51.87
15.92
22.63
43.09
Ours w/o 3-stage
57.95
15.93
22.78
42.29
Ours w/ 3-stage
57.99
16.21
22.81
42.83
Appendix
Table 9 : Ablation analysis of simple feature distillation strategy and three-stage curriculum strategy on zero-shot fine-grained understanding and dense prediction. We report Top-1 accuracy (%) for zero-shot fine-grained understanding tasks, alongside mean IoU (mIoU) for semantic segmentation.
Figure 6 : Ablation analysis of the training schedule. We compare the weighted total objective loss curves of models trained with and without the three-stage curriculum strategy. Both settings successfully converge, showing that the model can still train and converge stably when the curriculum strategy is removed.
Method
ZS. Fine-grained Tasks
LP. Seg.
SVHN
CLEVR Distance
CLEVR Counts
ADE20K
w/o DMA
56.77
15.99
20.99
39.80
w/ K
56.73
16.05
22.17
40.98
w/ D
56.87
16.01
22.39
39.97
w/ DMA
57.99
16.21
22.81
42.83
Appendix
Table 10 : Ablation analysis of the Dual-Matrix Alignment (DMA) strategy on various downstream tasks. We decouple the effects of the Spatial Relation Matrix ( K ) and the Energy Difference Matrix ( D ).
Reference Layer
ImageNet-1K
Avg.
2
76.90
68.57
4
77.06
68.35
6
77.13
68.60
8
77.07
68.43
10
76.95
68.58
Appendix
Table 11 : Ablation study on the reference layer selection for SALM-Self.
Method
VOC2012
Context
COCO-Stuff
Cityscapes
ADE20K
CLIP
14.79
4.06
2.04
1.28
1.43
Simple Distillation
2.07
0.29
0.05
0.16
0.02
DeCLIP
85.16
39.37
28.75
33.13
21.92
SALM (Ours)
17.03
7.75
4.22
3.08
3.34
Appendix
Table 12 : Open-vocabulary segmentation results. We report mIoU (%) on five segmentation benchmarks.
Method
Average
CLIP
0.0
20.0
40.0
20.0
6.7
20.0
33.3
6.7
33.3
20.0
KUEA
6.6
26.7
40.0
13.3
6.7
40.0
26.6
13.3
20.0
21.5
Ours
13.3
20.0
46.7
13.3
13.3
53.3
33.3
13.3
26.7
25.9
Appendix
Table 13 : Performance of CLIP on MMVP-VLM benchmark. Symbols for visual patterns are inherited: : Orientation and Direction, : Presence of Specific Features, : State and Condition, : Quantity and Count, : Positional and Relational Context, : Color and Appearance, : Structural and Physical Characteristics, : Texts, : Viewpoint and Perspective.
Method
Image-to-Text Retrieval
Text-to-Image Retrieval
Flickr30K
MSCOCO
Flickr30K
MSCOCO
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
CLIP
87.30
98.29
99.30
57.92
81.20
87.86
67.34
89.02
93.28
37.07
61.64
71.51
KUEA
87.30
98.10
99.30
58.26
81.24
88.27
68.70
89.80
93.83
38.05
62.64
72.60
Ours
86.70
98.40
99.10
59.12
82.08
88.36
68.88
89.80
94.20
38.54
63.25
73.10
Appendix
Table 14 : Zero-shot image-text retrieval results on Flickr30K and MSCOCO benchmarks. We report Recall@K (K=1, 5, 10) for both Image-to-Text and Text-to-Image retrieval tasks.
Figure 7 : The t-SNE visualization of global semantic features. Compared to other methods, our approach generates more compact and well-separated clusters, demonstrating superior capability in learning discriminative global representations.
Figure 8 : PCA-based visualizations. Unlike other enhancement methods, our approach excels at capturing fine-grained semantic nuances and preserving discriminative local details.