cs.LGOct 17, 2025

Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity

Authors: Naoki Yoshida, Satoshi Hayakawa, Yuhta Takida, Toshimitsu Uesaka, Hiromi Wakaki, Yuki Mitsufuji

Organizations: The University of Tokyo · Sony Group Corporation · Sony AI

Abstract

In this study, we propose an enhancement to the similarity computation mechanism in multimodal contrastive pretraining frameworks such as CLIP. Prior theoretical research has demonstrated that the optimal similarity metrics between paired modalities should correspond to the pointwise mutual information (PMI) between the two modalities. However, the current implementations of CLIP and its variants fail to fully utilize the underlying linear structure of PMI. We therefore propose KME-CLIP, which leverages this structure through the inner product in a reproducing kernel Hilbert space (RKHS). We theoretically prove that, under our assumptions, the KME-CLIP similarity can bring the contrastive loss arbitrarily close to its optimal value, which is attained by PMI, as the size of the point set grows, and we empirically evaluate KME-CLIP against CLIP and its kernel-based variants across several retrieval and classification tasks.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels

    Apr 17, 2026Hangke Sui, Yuqing Wang, Minh N DoContrastive AlignmentLarge Multimodal Models

  2. KODA: Contrastive Representation Comparison and Alignment for Vision-Language Foundation Models

    Jun 2, 2026Youqi Wu, Mohammad Jalali, Farzan FarniaContrastive LearningVision-Language Foundation Models

  3. CLIP-RD: Relational Distillation for Efficient CLIP Knowledge Distillation

    Date pendingJeannie Chung, Hanna Jang, Ingyeong Yang +2Contrastive Language-Image Pre-Training ModelSimilarity and Metrics