CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a lightweight route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that, for the InfoNCE objective, the reported forgetting is driven primarily not by insufficient negatives but by an inappropriate magnitude of the contrastive temperature τ: with τ set sufficiently small, contrastive post-training improves rather than degrades the pretrained CLIP, which we explain through the temperature dependence of the InfoNCE gradient. Building on this finding, we propose \textbf{ComCLIP}, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2. Over multiple seeds, ComCLIP matches the self-distillation baseline CLIP-Refine on zero-shot classification while significantly improving the transferability of visual features, measured by linear probing (48.99 vs.\ 42.28 on ViT-B/16), and on ViT-L/14 it also improves MMVP over CLIP-Refine (24.20 vs.\ 19.01); CLIP-Refine remains stronger on image-text retrieval. Used as a drop-in vision encoder for LLaVA-1.5-7B without re-aligning the projector or LLM, ComCLIP yields no net change across 8 VLM benchmarks, i.e., the refinement does not break downstream compatibility. Code and models are available at https://github.com/showstarpro/ComCLIP.git.
Figures & tables
Figure 1: Multi-dimensional comparison on ViT-L/14 (values from Table 1 ). Relative to the original OpenAI CLIP, ComCLIP improves all four evaluation dimensions (zero-shot classification, retrieval, linear probing, MMVP), whereas prior post-training methods exhibit characteristic trade-offs. Each axis is min–max normalized per metric; ComCLIP’s MMVP is the 3 -seed mean (Sec. 4.3 ).
Figure 2: Overview of ComCLIP. Three parallel branches operate jointly: (i) frozen CLIP text encoder for image-text alignment ( Lclip ), (ii) trainable CLIP image encoder anchored to its frozen pretrained counterpart ( Lmse ), and (iii) frozen DINOv2 encoder providing relational supervision via batch-wise similarity matrices ( Lrkd ). All three losses operate on the post-projection embedding.
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
CLS Avg
I → T
T → I
I → T
T → I
Avg
Avg
Backbone: ViT-B/16
OpenAI CLIP (no post-training)
61.82
48.16
31.47
74.70
57.10
44.70
12.59
+ KUEA (ImageNet-1K)
62.07
48.48
31.99
75.00
57.62
44.65
11.85
+ KUEA (CC3M)
61.99
48.00
31.47
75.00
57.00
44.72
14.07
+ CLIP-Refine (COCO Caption)
62.87 ±0.11
54.17 ±0.28
37.72 ±0.07
81.30 ±0.40
64.45 ±0.21
42.28 ±0.62
16.05 ±1.14
Table 1: Comparison of post-training methods on two backbones (ViT-B/16 and ViT-L/14). The “+” denotes the application of a post-training method and dataset to the pre-trained OpenAI CLIP. Each method is reported under its strongest validated setting (Sec. 4.2 ): ComCLIP and CLIP-Refine are trained for 1 epoch and KUEA for 2 epochs (its default schedule). Cells with a subscript are means ± standard deviation over seeds ( 3 seeds; see Table 2 ); all other cells are single runs. MMVP has N=135 image pairs and a seed standard deviation of up to ∼3 points, so single-run MMVP differences should not be over-interpreted. The best and second-best results within each backbone are in bold and underline ; highlighting does not imply statistical significance.
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
CLS Avg
I → T
T → I
I → T
T → I
Avg
Avg
ViT-B/16 ( 3 seeds)
CLIP-Refine (COCO Caption)
62.87 ±0.11
54.17 ±0.28
37.72 ±0.07
81.30 ±0.40
64.45 ±0.21
42.28 ±0.62
16.05 ±1.14
ComCLIP (CC3M)
63.00 ±0.13
52.49 ±0.28
36.17 ±0.02
77.80 ±0.00
63.11 ±0.09
48.99 ±0.35
14.81 ±2.97
Welch p
0.26
2×10−3
3×10−4
4×10−3
3×10−3
4×10−4
0.56
ViT-L/14 ( 3 seeds, both on CC3M)
Table 2: Multi-seed comparison between ComCLIP and CLIP-Refine (mean ± standard deviation). Top: ViT-B/16, 3 seeds per method, each method in its best setting (ComCLIP on CC3M, CLIP-Refine on COCO Caption). Bottom: ViT-L/14 MMVP, 3 seeds per method, both trained on CC3M. p : two-sided Welch t -test. MMVP has N=135 image pairs; on ViT-B/16 the ComCLIP seeds correspond to {16,20,24} correct pairs.
Method
Rel. Δ
Ceil. Δ
W/T/L
AI2D
POPE
RefCOCOg
V ∗
MME-c
MME-p
SQA-Img
OCRBench
TextVQA
OpenAI CLIP
–
–
–
52.75
85.20
20.67
41.36
320.00
1350.79
67.63
27.80
34.57
KUEA *
−1.23
−0.51
3/0/6
52.30
85.46
20.50
38.74
317.14
1350.82
66.68
27.40
34.69
CLIP-Refine
−0.95
−0.37
3/0/6
52.20
85.57
20.61
39.79
308.21
1345.97
67.72
27.60
34.84
ComCLIP
+0.02
+0.04
4/2/3
52.75
85.14
20.63
41.36
307.50
1373.47
67.82
28.20
34.89
Table 3: LLaVA-1.5-7B with the vision encoder swapped for ComCLIP ViT-L/14 (post-trained on CC3M), KUEA ViT-L/14 (ImageNet-1K), CLIP-Refine ViT-L/14 (COCO Caption), or the original OpenAI CLIP ViT-L/14; the LLM and projector are kept frozen. ‘*’ indicates official KUEA weights ( Gong et al., 2025 ) . Because the benchmarks differ in scale by more than 60× (e.g., MME-p vs. RefCOCOg), we do not report a raw arithmetic mean. Instead, Rel. Δ is the mean per-benchmark relative change w.r.t. OpenAI CLIP (%), Ceil. Δ is the mean change after dividing each benchmark by its maximum attainable score (points), and W/T/L counts benchmarks improved/tied/degraded w.r.t. OpenAI CLIP. The best result in each benchmark column is in bold .
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
OpenAI CLIP (no post-training)
61.82
48.16
31.47
74.70
57.10
44.70
12.59
+ KUEA
61.99
48.00
31.47
75.00
57.00
44.72
14.07
+ (KUEA + Lclip )
62.24
48.84
34.20
75.30
60.80
45.25
13.33
+ Lclip
62.90
51.92
38.28
76.80
64.56
44.04
14.07
+ Lmse
61.82
48.18
31.45
74.70
57.12
44.79
11.85
Table 4: Loss decoupling ablation on ViT-B/16, post-trained on CC3M with batch size 128×8 . All configurations of {Lclip,Lmse,Lrkd} are trained for 1 epoch. KUEA and KUEA + Lclip use KUEA’s default 2 -epoch schedule and hyperparameters, which were not re-tuned (loss weight, learning rate, anchor weight) after adding Lclip . Temperature is fixed at τ=0.01 throughout. Entries are single runs except where a seed standard deviation is given. Best results per column (except MMVP) in bold . All MMVP values ( N=135 ) lie within the 3 -seed range of ComCLIP ( [11.85,17.78] ), so MMVP is not highlighted and no conclusion is drawn from it.
τ
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
OpenAI CLIP (no post-training)
61.82
48.16
31.47
74.70
57.10
44.70
12.59
0.07
57.90
39.10
34.12
65.70
58.86
43.95
16.30
0.03
61.51
48.94
37.01
74.50
63.54
43.12
13.33
0.02
62.57
51.54
38.33
76.70
65.00
43.98
13.33
0.01
62.90
52.50
38.26
76.80
64.56
44.04
14.07
Table 5: Ablation study on the initialization of the temperature parameter ( τ ) in the CLIP loss. We conduct this experiment exclusively using the CLIP loss on the pre-trained OpenAI CLIP ViT-B/16 model to isolate the impact of τ . We conduct post-training on the CC3M dataset. The global batch size is fixed at 128×8 (1024) across all settings. The best results are highlighted in bold . MMVP ( N=135 , single runs) is reported for completeness only: its variation across rows is within seed-to-seed variance (Table 2 ), so it is not highlighted and not used to select τ .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
CLS Avg
I → T
T → I
I → T
T → I
Avg
Avg
SigLIP (no post-training)
71.24
68.16
52.20
88.90
74.80
70.87
39.26
KUEA (ImageNet-1K) *
71.67
-
-
-
-
-
-
ComCLIP (CC3M)
72.15
70.06
53.05
89.50
75.78
70.88
41.48
Appendix
Table 6: Performance evaluation of post-training methods based on the SigLIP backbone. We compare our proposed method ( ComCLIP ) against the pre-trained SigLIP ViT-SO400M/14 baseline and KUEA (post-trained on ImageNet-1K). All methods are evaluated on the CC3M dataset for post-training, except for the KUEA baseline which is evaluated on its original setting. ‘*’ denotes results reported by the original authors ( Gong et al., 2025 ) . The best results in each column are highlighted in bold .
Figure 3: Multi-dimensional performance comparison on SigLIP ViT-SO400M/14.
Batch Size
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
16×8
62.00
50.56
37.26
76.10
63.22
42.15
15.56
32×8
62.50
51.26
37.79
76.50
64.16
42.98
11.85
64×8
62.65
51.44
38.10
76.50
64.12
43.65
14.81
128×8
62.69
52.50
38.26
77.40
64.52
43.68
15.56
256×8
63.03
52.52
38.22
77.80
64.80
44.63
18.52
Appendix
Table 7: Ablation study on the effect of batch size during post-training with the CLIP loss. Experiments are conducted on the pre-trained OpenAI CLIP ViT-B/16 model, utilizing the CC3M dataset exclusively for post-training. Only the CLIP loss is applied to update the visual encoder (λclip=1.0,λmse=0.0,λrkd=0.0) . The batch sizes are denoted as (per-GPUbatchsize×numberofGPUs) . The best results in each column are highlighted in bold . MMVP ( N=135 , single runs) varies within seed-to-seed variance (Table 2 ) and is not highlighted.
Dataset
Learnable
Batch Size
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
COCO Caption
Yes
128×8
62.25
49.86
34.74
77.20
61.38
45.30
14.81
No
128×8
62.25
49.86
34.74
77.20
61.40
45.27
14.81
CC3M
Yes
128×8
62.95
52.22
38.22
76.90
64.48
43.42
14.07
No
128×8
62.90
52.50
38.26
76.80
64.56
44.04
14.07
Appendix
Table 8: Ablation study on the learnability of τ and the impact of different post-training datasets (COCO Caption vs. CC3M). All experiments are conducted on the pre-trained OpenAI CLIP ViT-B/16 model, utilizing exclusively the CLIP loss (λclip=1.0,λmse=0.0,λrkd=0.0) . During post-training, only the parameters of the visual encoder are updated while the text encoder remains strictly frozen. The best results in each metric are highlighted in bold (MMVP, single runs within seed-to-seed variance, is not highlighted).
Linear probe with vs. without Lrkd ( 3 seeds)
Backbone
Lclip+Lmse
+Lrkd (ComCLIP)
Welch p
ViT-B/16
48.54±0.10
48.94±0.41
0.23
ViT-L/14
56.36±0.23
57.12±0.14
0.013
Teacher for Lrkd (single runs)
Configuration
Zero-shot (12 avg)
Linear Probe (5 avg)
ViT-B/16, Lclip+Lmse (no teacher)
63.35
48.60
Appendix
Table 9: Contribution of Lrkd and choice of teacher. Top: adding Lrkd (DINOv2-Large teacher) to Lclip+Lmse , linear probe (5-dataset average), mean ± standard deviation over 3 seeds, CC3M, 1 epoch; p : two-sided Welch t -test. Bottom: teacher choice for Lrkd (single runs, CC3M, 1 epoch). The ViT-B/16 teacher runs use the same configuration as Table 4 .
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
OpenAI CLIP (no post-training)
61.82
48.16
31.47
74.70
57.10
44.70
12.59
ComCLIP-bp
63.19
51.44
36.54
77.00
63.50
48.23
13.33
ComCLIP
63.15
52.20
36.18
77.80
63.22
48.74
14.81 ±2.97
Appendix
Table 10: Investigation of the alignment position for DINOv2 rkd loss features on ViT-B/16. We conduct post-training on the CC3M dataset. ComCLIP-bp denotes that the visual features of the student model used for rkd loss calculation are extracted prior to the projection layer. ComCLIP represents our default configuration, where visual features are projected into the joint image-text embedding space via the original projection layer. Single runs except where a seed standard deviation is given; the MMVP difference is within seed-to-seed variance (Table 2 ) and is not highlighted.
Loss Weights
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
λclip
λmse
λrkd
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
1.0
1.0
0.5
62.77
52.42
36.47
78.00
63.06
49.75
14.81
1.0
0.5
1.0
62.91
52.40
36.56
77.50
63.32
47.29
13.33
1.0
0.5
0.5
63.03
52.62
36.76
77.80
63.38
48.85
17.04
1.0
1.0
1.0
63.15
52.20
36.18
77.80
63.22
48.74
14.81 ±2.97
Appendix
Table 11: Ablation study on the balancing coefficients of the three loss components used during post-training. Setting the weight of the primary CLIP loss as the baseline ( λclip=1.0 ), we investigate the impact of varying the proportions of the MSE loss ( λmse ) and RKD loss ( λrkd ) coefficients. All experiments are conducted on the CC3M dataset based on the pre-trained OpenAI CLIP ViT-B/16 architecture. The best results in each column are highlighted in bold ; MMVP (single runs except where a seed standard deviation is given) is within seed-to-seed variance and is not highlighted.
Method
epochs
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
CLS Avg
I → T
T → I
I → T
T → I
Avg
Avg
OpenAI CLIP (Baseline)
-
61.82
48.16
31.47
74.70
57.10
44.70
12.59
ComCLIP
1
63.15
52.20
36.18
77.80
63.22
48.74
14.81 ±2.97
ComCLIP
2
62.84
52.02
36.24
77.20
63.44
49.98
16.30
Appendix
Table 12: Performance evaluation of post-training with ComCLIP on the CC3M dataset. Experiments are based on the pre-trained OpenAI CLIP ViT-B/16 architecture. We compare the results obtained after 1 and 2 epochs of post-training against the OpenAI CLIP baseline. The best results in each column are highlighted in bold ; MMVP (single run for 2 epochs, 3 -seed mean for 1 epoch) is within seed-to-seed variance and is not highlighted.
τ
logit Scale ( 1/τ )
Text Anchor KL ( ↓ )
0.07
14.29
5.50
0.03
33.33
2.18
0.02
50.00
1.25
0.01
100.00
0.48
0.007
142.86
0.56
0.005
200.00
1.15
Appendix
Table 13: Text-anchored geometric diagnostic of models post-trained using only Lclip across different temperature ( τ ) values. A lower Text Anchor KL ( ↓ , closer to 0) indicates that the post-trained visual feature manifold has not suffered from severe structural drift relative to the frozen text features.
Figure 4: Lclip Training Curve. With the temperature τ ( logit_scale=1/τ ) set to various values, post-training was performed using only Lclip . Convergence of the loss function was optimal when logit_scale=100 .
Model
Dataset
epochs
Method
Avg
IN-1k
Cifar10
Cifar100
CalTech
FER
Pets
DTD
RESISC
EuroSAT
PCAM
IN-S
IN-O
ViT-B/16
——
-
OpenAI CLIP
61.82
68.35
90.00
65.61
82.17
46.39
88.99
44.95
58.19
55.91
50.73
48.24
42.30
ImageNet-1k
2
KUEA
62.07
68.52
90.69
67.08
82.14
46.45
89.21
45.00
58.14
55.26
51.01
48.23
43.15
CC3M
2
KUEA
61.99
68.40
90.27
66.52
82.19
46.53
89.02
45.05
58.33
56.00
51.12
48.25
42.15
COCO Caption
1
CLIP-Refine
62.99
67.74
90.50
68.28
83.40
49.80
89.75
44.68
59.35
56.20
55.18
48.14
42.90
CC3M
1
CLIP-Refine
61.59
63.30
89.56
67.25
84.29
49.47
81.47
44.63
59.97
59.46
54.23
44.28
41.20
CC3M
2
CLIP-Refine
61.10
62.71
88.46
66.71
83.96
47.80
80.76
44.36
59.21
60.17
54.60
43.67
40.75
Appendix
Table 14: Top-1 zero-shot classification accuracy on 12 public benchmarks. The best and second-best results in each model block are highlighted in bold and underline , respectively.
Model
Dataset
epochs
Method
MSCOCO
Flickr30K
Image-to-Text
Text-to-Image
Image-to-Text
Text-to-Image
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
ViT-B/16
——
-
OpenAI CLIP
48.16
72.58
81.88
31.47
55.62
66.73
74.70
92.90
96.40
57.10
81.36
88.20
ImageNet-1k
2
KUEA
48.48
72.98
82.24
31.99
56.35
67.21
75.00
93.10
96.60
57.62
81.82
88.66
CC3M
2
KUEA
48.00
72.64
81.76
31.47
55.79
66.71
75.00
92.80
96.60
57.00
81.46
88.38
COCO Caption
1
CLIP-Refine
53.96
79.02
86.60
37.67
63.83
74.26
80.90
95.40
97.50
64.64
87.34
92.46
Appendix
Table 15: Zero-shot image-text and text-to-image retrieval results on MSCOCO and Flickr30K benchmarks. The best and second-best results in each model block are highlighted in bold and underline , respectively.
Model
Dataset
epochs
Method
Avg
IN-1K
SVHN
GTSRB
CLEVR Dist.
CLEVR Counts
ViT-B/16
——
-
OpenAI CLIP
44.70
66.34
45.19
56.94
31.37
23.68
IN-1k
2
KUEA
44.65
66.82
47.85
58.19
28.34
22.07
CC3M
2
KUEA
44.72
66.46
46.22
57.48
30.83
22.62
COCO Caption
1
CLIP-Refine
42.01
67.51
44.53
56.60
18.58
22.83
CC3M
1
CLIP-Refine
44.89
66.22
45.61
56.94
31.38
24.28
CC3M
2
CLIP-Refine
43.52
65.74
45.20
56.76
25.89
23.99
Appendix
Table 16: Linear probe evaluation results on various datasets. The best and second-best results in each model block are highlighted in bold and underline , respectively.
Model
Dataset
epochs
Method
Avg
Orientation and
Presence of Specific
State and Condition
Quantity and
Positional and
Color and Appearance
Structural and
Texts
Viewpoint and Perspective
ViT-B/16
——
-
OpenAI CLIP
12.59
6.67
0.00
26.67
13.33
13.33
20.00
13.33
0.00
20.00
ImageNet-1k
2
KUEA
11.85
0.00
0.00
20.00
13.33
13.33
26.67
13.33
0.00
20.00
CC3M
2
KUEA
14.07
6.67
0.00
26.67
20.00
13.33
20.00
20.00
0.00
20.00
COCO Caption
1
CLIP-Refine
17.04
13.33
6.67
20.00
0.00
20.00
46.67
33.33
0.00
13.33
CC3M
1
CLIP-Refine
19.26
13.33
0.00
26.67
6.67
26.67
53.33
33.33
0.00
13.33
CC3M
2
CLIP-Refine
20.74
13.33
0.00
33.33
20.00
20.00
40.00
33.33
6.67
20.00
Appendix
Table 17: Detailed MMVP benchmark results (single runs; each category has only 15 pairs, so one pair corresponds to 6.67 points). ‡ : per-category breakdown of one seed; on ViT-B/16 this seed is the maximum of three (3-seed mean 14.81±2.97 ), and on ViT-L/14 it is one of three seeds (3-seed mean 24.20±0.86 ). The multi-seed means in Table 2 are the figures we report; per-category and single-run differences are not used for any conclusion. The best results in each model block are highlighted in bold .
Method
Enhancement signal
IN-1K zero-shot
Retrieval (as reported)
MMVP-VLM
DIVA ( Wang et al., 2025 )
diffusion feedback
75.5→75.5
COCO R@1 I → T 56.4→56.7 , T → I 36.5→36.6
19.3→25.9
GenHancer ( Ma et al., 2025 )
generative reconstruction
75.5→75.6†
COCO text retrieval R@5 79.2→79.4
31.9
un2CLIP ( Li et al., 2025 )
inverted unCLIP
75.5→62.4
COCO T → I R@5 61.0→65.5 ; Flickr T → I R@5 87.3→90.1
32.6
ComCLIP (ours, CC3M)
contrastive + anchor + RKD, 1 epoch
75.54→76.10
COCO R@1 I → T 50.04→54.56 , T → I 34.29→39.40 ; Flickr T → I R@5 83.48→88.22
17.78→24.20±0.86
Appendix
Table 18: Comparison with generative-model-based enhancement methods on OpenAI CLIP ViT-L/14 (224px), using the numbers published by each paper (base → post-trained, as reported in that paper). Evaluation protocols and base values differ across papers (e.g., retrieval R@1 base values), so the rows are not strictly comparable; each row should be read as the change relative to its own base. “–”: not reported in the form used here. † A third-party reproduction from GenHancer’s official checkpoints reports ImageNet-1K 40.2 ( Li et al., 2025 ) ; the two sources disagree. ComCLIP MMVP is the 3 -seed mean (Table 2 ).
Contrastive vision-language models continue to be the dominant approach for image-text retrieval. Contrastive Language-Image Pre-training (CLIP) trains two neural networks to align their image and text embeddings in a shared latent space. As a challenging case-study for neurosymbolic AI, recent results evaluating CLIP on negated or paraphrased text have shown mixed performance as these are difficult to define formally for text data. Negation produces the opposite meaning using various possible but small lexical changes. Paraphrasing may use very different textual expressions to denote essentially the same thing. As a result, learning of paraphrasing and negation together poses a significant challenge because of the above mismatch between changes in syntax and intended meaning expected to be captured by distances in embedding space. This paper proposes a new CLIP contrastive loss function capable of balancing the requirements of having both paraphrasing and negation. It applies training triplets consisting of original, paraphrased and negated text generated by multiple large language models to the evaluation of CLIP models. The approach, called SemCLIP, aims to learn semantically-relevant and simple embeddings, placing paraphrased captions nearer to the original image embeddings while at the same time pushing negated captions farther away. Empirically, SemCLIP is shown to be capable of preserving roughly the same performance as CLIP augmented with either negation or paraphrasing. Although direct comparisons are difficult to make because the problem of learning with both negation and paraphrasing is different, an expected benefit of SemCLIP should be robustness when applied zero-shot to downstream image classification tasks. Our experiments confirm such robustness as measured by difference in accuracy (mean-accuracy delta) between original and negated captions on five downstream datasets.
Kwun Ho Ngan, Saman Sadeghi Afgeh, Joe Townsend +1
Fujitsu Research of Europe · United Kingdom · City St George’s, University of London
Contrastive Language-Image Pre-training (CLIP) relies on Vision Transformers whose attention mechanism is susceptible to spurious correlations, and scales quadratically with resolution. To address these limitations, We present CLIMP, the first fully Mamba-based contrastive vision-language model that replaces both the vision and text encoders with Mamba. The new architecture encodes sequential structure in both vision and language, with VMamba capturing visual spatial inductive biases, reducing reliance on spurious correlations and producing an embedding space favorable for cross-modal retrieval and out-of-distribution robustness-surpassing OpenAI's CLIP-ViT-B by 7.5% on ImageNet-O. CLIMP naturally supports variable input resolutions without positional encoding interpolation or specialized training, achieving up to 6.6% higher retrieval accuracy at 16x training resolution while using 5x less memory and 1.8x fewer FLOPs. The autoregressive text encoder further overcomes CLIP's fixed context limitation, enabling dense captioning retrieval. Our findings suggest that Mamba exhibits advantageous properties for vision-language learning, making it a compelling alternative to Transformer-based CLIP.The code and models are publicly available at https://github.com/NimrodShabtay/CLIMP}
Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embeddings, rendering them insensitive to negation. We attribute this failure to a phenomenon we call Representational Collapse: by tracking compositional divergence and visual alignment across the CLIP text encoder, we show that middle layers build compositional syntax, but the final layers collapse this structure as visual alignment rises, producing a syntax-blind final representation. To recover the lost negation signal without altering pretrained weights, we propose PeakPatch, a lightweight post-hoc correction system that intercepts the encoder at its compositional peak. An Embedding Correction Network (ECN) uses cross-attention to extract a negation-specific signal from the peak layer, anchored to a stable baseline, and predicts a deviation vector that re-injects the lost syntax into the final-layer embedding space. A complementary Score Correction Network (SCN) predicts bounded scalar score offsets for discriminative tasks. Both modules are trained jointly end-to-end while all CLIP parameters remain frozen, adding only 5.2M parameters (3.5% of the backbone) and preserving the standard cosine similarity interface. On NegBench, PeakPatch achieves 74.3% on COCO MCQ (+35.1 over CLIP, +17.8 over the best encoder fine-tuning method) and 65.5% on VOC MCQ, while outperforming all fine-tuning baselines on fully out-of-distribution negation retrieval despite training only 3.5% of the parameters. The corrected embeddings also transfer to text-to-image generation (+18.4 negation score) and generalize across ViT-B/32, ViT-L/14, and SigLIP backbones. Project URL: https://stevencylu.github.io/PeakPatch/.