Theoretical Refinement of CLIP by Utilizing Linear Structure of Optimal Similarity
Organizations: The University of Tokyo · Sony Group Corporation · Sony AI
Abstract
In this study, we propose an enhancement to the similarity computation mechanism in multimodal contrastive pretraining frameworks such as CLIP. Prior theoretical research has demonstrated that the optimal similarity metrics between paired modalities should correspond to the pointwise mutual information (PMI) between the two modalities. However, the current implementations of CLIP and its variants fail to fully utilize the underlying linear structure of PMI. We therefore propose KME-CLIP, which leverages this structure through the inner product in a reproducing kernel Hilbert space (RKHS). We theoretically prove that, under our assumptions, the KME-CLIP similarity can bring the contrastive loss arbitrarily close to its optimal value, which is attained by PMI, as the size of the point set grows, and we empirically evaluate KME-CLIP against CLIP and its kernel-based variants across several retrieval and classification tasks.
Figures & tables
| CC3M | MSCOCO | Flickr30K | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Text retrieval | Model | Top1 | Top5 | Top10 | Top1 | Top5 | Top10 | Top1 | Top5 | Top10 |
| CC3M | CLIP | 21.95 | 41.97 | 51.01 | 13.64 | 32.34 | 43.57 | 26.46 | 52.18 | 62.68 |
| WPSE | 21.39 | 41.36 | 50.51 | 14.78 | 33.75 | 44.86 | 27.22 | 52.80 | 63.92 | |
| KME-CLIP | 24.22 | 44.85 | 53.51 | 14.14 | 34.25 | 45.65 | 28.18 | 55.72 | 65.92 | |
| CC12M | CLIP | 22.82 | 42.89 | 52.22 | 24.56 | 48.25 | 60.15 | 44.76 | 73.30 | 81.94 |
| WPSE | 22.59 | 42.70 | 51.94 | 24.29 | 48.85 | 61.07 | 45.80 | 73.74 | 82.40 | |
| Model | Average | ImageNet | CIFAR-10 | CIFAR-100 | STL-10 | Food-101 | Caltech-101 | Cars | Aircraft | Flowers | EuroSAT | DTD | Pets | SUN397 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CC3M | CLIP | 26.33 | 20.95 | 56.77 | 24.38 | 82.80 | 14.06 | 50.07 | 1.46 | 1.20 | 11.64 | 13.84 | 13.88 | 14.10 | 37.20 |
| WPSE | 27.38 | 22.17 | 59.40 | 30.78 | 81.64 | 14.80 | 50.70 | 1.27 | 1.13 | 14.01 | 16.20 | 13.83 | 12.78 | 37.27 | |
| KME-CLIP | 29.31 | 22.66 | 64.31 | 31.32 | 85.23 | 16.17 | 53.67 | 1.29 | 1.04 | 16.54 | 17.68 | 12.39 | 16.26 | 42.49 | |
| CC12M | CLIP | 43.48 | 36.99 | 75.41 | 42.77 | 92.36 | 45.97 | 71.60 | 17.44 | 2.48 | 25.58 | 31.72 | 19.47 | 55.59 | 47.88 |
| WPSE | 43.37 | 36.92 | 74.54 | 42.22 | 92.24 | 43.30 | 71.53 | 18.62 | 3.80 | 22.32 | 32.98 | 19.47 | 54.52 | 51.36 | |
| KME-CLIP | 45.67 | 39.07 | 78.06 | 46.63 | 92.74 | 49.14 | 76.62 | 17.52 | 3.45 | 25.00 | 31.28 | 21.60 | 60.67 | 51.94 |
| Image point-set size | (CLIP) | 2 | 10 | 50 | 100 | 197 |
|---|---|---|---|---|---|---|
| Top-1 accuracy | (23.66) | 23.71 | 23.93 | 24.35 | 24.79 | 25.57 |
| Dataset | Metric | CLIP | KME-CLIP (Bary) | KME-CLIP (LogSim) |
|---|---|---|---|---|
| MSCOCO | Gap ( ) | 0.508 | 0.402 | – |
| Align ( ) | 0.380 | 0.418 | 0.494 | |
| ImageNet | Gap ( ) | 0.426 | 0.357 | – |
| Align ( ) | 0.298 | 0.318 | 0.408 | |
| CIFAR-10 | Gap ( ) | 0.826 | 0.754 | – |
| Align ( ) | 0.300 | 0.278 | 0.432 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Config | Value |
|---|---|
| Total epochs | 50 |
| Warmup epochs | 2 |
| Warmup start learning rate | |
| Warmup end learning rate | |
| End learning rate | |
| AdamW |
| Model | Average | ImageNet | CIFAR-10 | CIFAR-100 | STL-10 | Food-101 | Caltech-101 | Cars | Aircraft | Flowers | EuroSAT | DTD | Pets | SUN397 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CC3M | CLIP | 72.29 | 59.59 | 88.75 | 70.33 | 92.94 | 67.02 | 82.34 | 37.45 | 44.20 | 92.02 | 95.88 | 66.01 | 72.88 | 70.32 |
| WPSE | 74.29 | 61.34 | 89.51 | 70.89 | 93.63 | 69.80 | 85.22 | 44.35 | 48.51 | 92.57 | 95.10 | 68.03 | 75.22 | 71.64 | |
| KME-CLIP | 74.44 | 62.02 | 89.73 | 70.94 | 93.10 | 69.76 | 84.62 | 45.69 | 49.28 | 93.16 | 95.80 | 67.82 | 74.64 | 71.13 | |
| CC12M | CLIP | 79.45 | 67.79 | 91.23 | 73.96 | 95.56 | 78.79 | 88.50 | 66.55 | 48.77 | 93.46 | 95.74 | 74.41 | 81.58 | 76.50 |
| WPSE | 81.17 | 69.71 | 91.98 | 74.78 | 95.69 | 79.81 | 89.93 | 71.45 | 57.60 | 94.93 | 96.02 | 72.66 | 83.62 | 77.05 | |
| KME-CLIP | 78.31 | 67.76 | 91.92 | 75.24 | 95.41 | 78.28 | 88.13 | 60.22 | 46.54 | 93.52 | 96.16 | 73.72 | 74.89 | 76.20 |
| Image point-set size | CLIP | 2 | 10 | 50 | 100 | 197 |
|---|---|---|---|---|---|---|
| Accuracy (%) | 23.66 | 23.71 | 23.93 | 24.35 | 24.79 | 25.57 |
| Inference time with two GPUs (sec) | 126 | 138 | 140 | 159 | 186 | 257 |
| Image point-set size | CLIP | 2 | 10 | 50 | 100 | 197 |
|---|---|---|---|---|---|---|
| Training time (sec/epoch) | 1604 | 1413 | 1439 | 925 | 722 | 945 |
| Number of GPUs for training | 2 | 2 | 2 | 4 | 8 | 8 |
| CC3M | MSCOCO | Flickr30K | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Top1 | Top5 | Top10 | Top1 | Top5 | Top10 | Top1 | Top5 | Top10 |
| Text retrieval | |||||||||
| CLIP | |||||||||
| WPSE | |||||||||
| KME-CLIP | |||||||||
| KME-no-log | |||||||||
| Dataset | CLIP | WPSE | KME-CLIP | KME-no-log |
|---|---|---|---|---|
| Average | ||||
| ImageNet | ||||
| CIFAR-10 | ||||
| CIFAR-100 | ||||
| STL-10 | ||||
| Food-101 |
| Final embedding dimension | ||||
| Conditional KL | ||||
| CLIP | WPSE | KME-no-log | KME-CLIP | Bayes | ||
|---|---|---|---|---|---|---|