cs.CVOct 6, 2026

Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked Modeling

Authors: Juntong Li, Lingwei Dang, Haomin Wu, Ziyan Qiu, Qingxin Xiao, Qingyao Wu

Organizations: School of Software Engineering, South China University of Technology

Abstract

Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space. To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring any image-text pairs. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space. Furthermore, driven by the empirical observations that CLIP's shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP's intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP's zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs. Project page at https://qzfm.github.io/salm_project_page/.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Fine-grained CLIP fine-tuning with self-annotated region alignment

    Jul 15, 2026Chenyang Zhao, Wei Lin, Antoni B. Chan +1Contrastive Language-Image Pre-Training ModelFine-Grained Perception

  2. Deep Pre-Alignment for VLMs

    May 14, 2026Tianyu Yu, Kechen Fang, Zihao Wan +5Vision-Language AlignmentVision Transformer