CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a lightweight route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that, for the InfoNCE objective, the reported forgetting is driven primarily not by insufficient negatives but by an inappropriate magnitude of the contrastive temperature τ: with τ set sufficiently small, contrastive post-training improves rather than degrades the pretrained CLIP, which we explain through the temperature dependence of the InfoNCE gradient. Building on this finding, we propose \textbf{ComCLIP}, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2. Over multiple seeds, ComCLIP matches the self-distillation baseline CLIP-Refine on zero-shot classification while significantly improving the transferability of visual features, measured by linear probing (48.99 vs.\ 42.28 on ViT-B/16), and on ViT-L/14 it also improves MMVP over CLIP-Refine (24.20 vs.\ 19.01); CLIP-Refine remains stronger on image-text retrieval. Used as a drop-in vision encoder for LLaVA-1.5-7B without re-aligning the projector or LLM, ComCLIP yields no net change across 8 VLM benchmarks, i.e., the refinement does not break downstream compatibility. Code and models are available at https://github.com/showstarpro/ComCLIP.git.
Figures & tables
Figure 1: Multi-dimensional comparison on ViT-L/14 (values from Table 1 ). Relative to the original OpenAI CLIP, ComCLIP improves all four evaluation dimensions (zero-shot classification, retrieval, linear probing, MMVP), whereas prior post-training methods exhibit characteristic trade-offs. Each axis is min–max normalized per metric; ComCLIP’s MMVP is the 3 -seed mean (Sec. 4.3 ).
Figure 2: Overview of ComCLIP. Three parallel branches operate jointly: (i) frozen CLIP text encoder for image-text alignment ( Lclip ), (ii) trainable CLIP image encoder anchored to its frozen pretrained counterpart ( Lmse ), and (iii) frozen DINOv2 encoder providing relational supervision via batch-wise similarity matrices ( Lrkd ). All three losses operate on the post-projection embedding.
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
CLS Avg
I → T
T → I
I → T
T → I
Avg
Avg
Backbone: ViT-B/16
OpenAI CLIP (no post-training)
61.82
48.16
31.47
74.70
57.10
44.70
12.59
+ KUEA (ImageNet-1K)
62.07
48.48
31.99
75.00
57.62
44.65
11.85
+ KUEA (CC3M)
61.99
48.00
31.47
75.00
57.00
44.72
14.07
+ CLIP-Refine (COCO Caption)
62.87 ±0.11
54.17 ±0.28
37.72 ±0.07
81.30 ±0.40
64.45 ±0.21
42.28 ±0.62
16.05 ±1.14
Table 1: Comparison of post-training methods on two backbones (ViT-B/16 and ViT-L/14). The “+” denotes the application of a post-training method and dataset to the pre-trained OpenAI CLIP. Each method is reported under its strongest validated setting (Sec. 4.2 ): ComCLIP and CLIP-Refine are trained for 1 epoch and KUEA for 2 epochs (its default schedule). Cells with a subscript are means ± standard deviation over seeds ( 3 seeds; see Table 2 ); all other cells are single runs. MMVP has N=135 image pairs and a seed standard deviation of up to ∼3 points, so single-run MMVP differences should not be over-interpreted. The best and second-best results within each backbone are in bold and underline ; highlighting does not imply statistical significance.
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
CLS Avg
I → T
T → I
I → T
T → I
Avg
Avg
ViT-B/16 ( 3 seeds)
CLIP-Refine (COCO Caption)
62.87 ±0.11
54.17 ±0.28
37.72 ±0.07
81.30 ±0.40
64.45 ±0.21
42.28 ±0.62
16.05 ±1.14
ComCLIP (CC3M)
63.00 ±0.13
52.49 ±0.28
36.17 ±0.02
77.80 ±0.00
63.11 ±0.09
48.99 ±0.35
14.81 ±2.97
Welch p
0.26
2×10−3
3×10−4
4×10−3
3×10−3
4×10−4
0.56
ViT-L/14 ( 3 seeds, both on CC3M)
Table 2: Multi-seed comparison between ComCLIP and CLIP-Refine (mean ± standard deviation). Top: ViT-B/16, 3 seeds per method, each method in its best setting (ComCLIP on CC3M, CLIP-Refine on COCO Caption). Bottom: ViT-L/14 MMVP, 3 seeds per method, both trained on CC3M. p : two-sided Welch t -test. MMVP has N=135 image pairs; on ViT-B/16 the ComCLIP seeds correspond to {16,20,24} correct pairs.
Method
Rel. Δ
Ceil. Δ
W/T/L
AI2D
POPE
RefCOCOg
V ∗
MME-c
MME-p
SQA-Img
OCRBench
TextVQA
OpenAI CLIP
–
–
–
52.75
85.20
20.67
41.36
320.00
1350.79
67.63
27.80
34.57
KUEA *
−1.23
−0.51
3/0/6
52.30
85.46
20.50
38.74
317.14
1350.82
66.68
27.40
34.69
CLIP-Refine
−0.95
−0.37
3/0/6
52.20
85.57
20.61
39.79
308.21
1345.97
67.72
27.60
34.84
ComCLIP
+0.02
+0.04
4/2/3
52.75
85.14
20.63
41.36
307.50
1373.47
67.82
28.20
34.89
Table 3: LLaVA-1.5-7B with the vision encoder swapped for ComCLIP ViT-L/14 (post-trained on CC3M), KUEA ViT-L/14 (ImageNet-1K), CLIP-Refine ViT-L/14 (COCO Caption), or the original OpenAI CLIP ViT-L/14; the LLM and projector are kept frozen. ‘*’ indicates official KUEA weights ( Gong et al., 2025 ) . Because the benchmarks differ in scale by more than 60× (e.g., MME-p vs. RefCOCOg), we do not report a raw arithmetic mean. Instead, Rel. Δ is the mean per-benchmark relative change w.r.t. OpenAI CLIP (%), Ceil. Δ is the mean change after dividing each benchmark by its maximum attainable score (points), and W/T/L counts benchmarks improved/tied/degraded w.r.t. OpenAI CLIP. The best result in each benchmark column is in bold .
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
OpenAI CLIP (no post-training)
61.82
48.16
31.47
74.70
57.10
44.70
12.59
+ KUEA
61.99
48.00
31.47
75.00
57.00
44.72
14.07
+ (KUEA + Lclip )
62.24
48.84
34.20
75.30
60.80
45.25
13.33
+ Lclip
62.90
51.92
38.28
76.80
64.56
44.04
14.07
+ Lmse
61.82
48.18
31.45
74.70
57.12
44.79
11.85
Table 4: Loss decoupling ablation on ViT-B/16, post-trained on CC3M with batch size 128×8 . All configurations of {Lclip,Lmse,Lrkd} are trained for 1 epoch. KUEA and KUEA + Lclip use KUEA’s default 2 -epoch schedule and hyperparameters, which were not re-tuned (loss weight, learning rate, anchor weight) after adding Lclip . Temperature is fixed at τ=0.01 throughout. Entries are single runs except where a seed standard deviation is given. Best results per column (except MMVP) in bold . All MMVP values ( N=135 ) lie within the 3 -seed range of ComCLIP ( [11.85,17.78] ), so MMVP is not highlighted and no conclusion is drawn from it.
τ
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
OpenAI CLIP (no post-training)
61.82
48.16
31.47
74.70
57.10
44.70
12.59
0.07
57.90
39.10
34.12
65.70
58.86
43.95
16.30
0.03
61.51
48.94
37.01
74.50
63.54
43.12
13.33
0.02
62.57
51.54
38.33
76.70
65.00
43.98
13.33
0.01
62.90
52.50
38.26
76.80
64.56
44.04
14.07
Table 5: Ablation study on the initialization of the temperature parameter ( τ ) in the CLIP loss. We conduct this experiment exclusively using the CLIP loss on the pre-trained OpenAI CLIP ViT-B/16 model to isolate the impact of τ . We conduct post-training on the CC3M dataset. The global batch size is fixed at 128×8 (1024) across all settings. The best results are highlighted in bold . MMVP ( N=135 , single runs) is reported for completeness only: its variation across rows is within seed-to-seed variance (Table 2 ), so it is not highlighted and not used to select τ .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
CLS Avg
I → T
T → I
I → T
T → I
Avg
Avg
SigLIP (no post-training)
71.24
68.16
52.20
88.90
74.80
70.87
39.26
KUEA (ImageNet-1K) *
71.67
-
-
-
-
-
-
ComCLIP (CC3M)
72.15
70.06
53.05
89.50
75.78
70.88
41.48
Appendix
Table 6: Performance evaluation of post-training methods based on the SigLIP backbone. We compare our proposed method ( ComCLIP ) against the pre-trained SigLIP ViT-SO400M/14 baseline and KUEA (post-trained on ImageNet-1K). All methods are evaluated on the CC3M dataset for post-training, except for the KUEA baseline which is evaluated on its original setting. ‘*’ denotes results reported by the original authors ( Gong et al., 2025 ) . The best results in each column are highlighted in bold .
Figure 3: Multi-dimensional performance comparison on SigLIP ViT-SO400M/14.
Batch Size
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
16×8
62.00
50.56
37.26
76.10
63.22
42.15
15.56
32×8
62.50
51.26
37.79
76.50
64.16
42.98
11.85
64×8
62.65
51.44
38.10
76.50
64.12
43.65
14.81
128×8
62.69
52.50
38.26
77.40
64.52
43.68
15.56
256×8
63.03
52.52
38.22
77.80
64.80
44.63
18.52
Appendix
Table 7: Ablation study on the effect of batch size during post-training with the CLIP loss. Experiments are conducted on the pre-trained OpenAI CLIP ViT-B/16 model, utilizing the CC3M dataset exclusively for post-training. Only the CLIP loss is applied to update the visual encoder (λclip=1.0,λmse=0.0,λrkd=0.0) . The batch sizes are denoted as (per-GPUbatchsize×numberofGPUs) . The best results in each column are highlighted in bold . MMVP ( N=135 , single runs) varies within seed-to-seed variance (Table 2 ) and is not highlighted.
Dataset
Learnable
Batch Size
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
COCO Caption
Yes
128×8
62.25
49.86
34.74
77.20
61.38
45.30
14.81
No
128×8
62.25
49.86
34.74
77.20
61.40
45.27
14.81
CC3M
Yes
128×8
62.95
52.22
38.22
76.90
64.48
43.42
14.07
No
128×8
62.90
52.50
38.26
76.80
64.56
44.04
14.07
Appendix
Table 8: Ablation study on the learnability of τ and the impact of different post-training datasets (COCO Caption vs. CC3M). All experiments are conducted on the pre-trained OpenAI CLIP ViT-B/16 model, utilizing exclusively the CLIP loss (λclip=1.0,λmse=0.0,λrkd=0.0) . During post-training, only the parameters of the visual encoder are updated while the text encoder remains strictly frozen. The best results in each metric are highlighted in bold (MMVP, single runs within seed-to-seed variance, is not highlighted).
Linear probe with vs. without Lrkd ( 3 seeds)
Backbone
Lclip+Lmse
+Lrkd (ComCLIP)
Welch p
ViT-B/16
48.54±0.10
48.94±0.41
0.23
ViT-L/14
56.36±0.23
57.12±0.14
0.013
Teacher for Lrkd (single runs)
Configuration
Zero-shot (12 avg)
Linear Probe (5 avg)
ViT-B/16, Lclip+Lmse (no teacher)
63.35
48.60
Appendix
Table 9: Contribution of Lrkd and choice of teacher. Top: adding Lrkd (DINOv2-Large teacher) to Lclip+Lmse , linear probe (5-dataset average), mean ± standard deviation over 3 seeds, CC3M, 1 epoch; p : two-sided Welch t -test. Bottom: teacher choice for Lrkd (single runs, CC3M, 1 epoch). The ViT-B/16 teacher runs use the same configuration as Table 4 .
Method
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
OpenAI CLIP (no post-training)
61.82
48.16
31.47
74.70
57.10
44.70
12.59
ComCLIP-bp
63.19
51.44
36.54
77.00
63.50
48.23
13.33
ComCLIP
63.15
52.20
36.18
77.80
63.22
48.74
14.81 ±2.97
Appendix
Table 10: Investigation of the alignment position for DINOv2 rkd loss features on ViT-B/16. We conduct post-training on the CC3M dataset. ComCLIP-bp denotes that the visual features of the student model used for rkd loss calculation are extracted prior to the projection layer. ComCLIP represents our default configuration, where visual features are projected into the joint image-text embedding space via the original projection layer. Single runs except where a seed standard deviation is given; the MMVP difference is within seed-to-seed variance (Table 2 ) and is not highlighted.
Loss Weights
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
λclip
λmse
λrkd
12 datasets Avg
I → T
T → I
I → T
T → I
5 datasets Avg
Avg
1.0
1.0
0.5
62.77
52.42
36.47
78.00
63.06
49.75
14.81
1.0
0.5
1.0
62.91
52.40
36.56
77.50
63.32
47.29
13.33
1.0
0.5
0.5
63.03
52.62
36.76
77.80
63.38
48.85
17.04
1.0
1.0
1.0
63.15
52.20
36.18
77.80
63.22
48.74
14.81 ±2.97
Appendix
Table 11: Ablation study on the balancing coefficients of the three loss components used during post-training. Setting the weight of the primary CLIP loss as the baseline ( λclip=1.0 ), we investigate the impact of varying the proportions of the MSE loss ( λmse ) and RKD loss ( λrkd ) coefficients. All experiments are conducted on the CC3M dataset based on the pre-trained OpenAI CLIP ViT-B/16 architecture. The best results in each column are highlighted in bold ; MMVP (single runs except where a seed standard deviation is given) is within seed-to-seed variance and is not highlighted.
Method
epochs
Zero-shot
MS COCO (R@1)
Flickr30K (R@1)
Linear Probe
MMVP
CLS Avg
I → T
T → I
I → T
T → I
Avg
Avg
OpenAI CLIP (Baseline)
-
61.82
48.16
31.47
74.70
57.10
44.70
12.59
ComCLIP
1
63.15
52.20
36.18
77.80
63.22
48.74
14.81 ±2.97
ComCLIP
2
62.84
52.02
36.24
77.20
63.44
49.98
16.30
Appendix
Table 12: Performance evaluation of post-training with ComCLIP on the CC3M dataset. Experiments are based on the pre-trained OpenAI CLIP ViT-B/16 architecture. We compare the results obtained after 1 and 2 epochs of post-training against the OpenAI CLIP baseline. The best results in each column are highlighted in bold ; MMVP (single run for 2 epochs, 3 -seed mean for 1 epoch) is within seed-to-seed variance and is not highlighted.
τ
logit Scale ( 1/τ )
Text Anchor KL ( ↓ )
0.07
14.29
5.50
0.03
33.33
2.18
0.02
50.00
1.25
0.01
100.00
0.48
0.007
142.86
0.56
0.005
200.00
1.15
Appendix
Table 13: Text-anchored geometric diagnostic of models post-trained using only Lclip across different temperature ( τ ) values. A lower Text Anchor KL ( ↓ , closer to 0) indicates that the post-trained visual feature manifold has not suffered from severe structural drift relative to the frozen text features.
Figure 4: Lclip Training Curve. With the temperature τ ( logit_scale=1/τ ) set to various values, post-training was performed using only Lclip . Convergence of the loss function was optimal when logit_scale=100 .
Model
Dataset
epochs
Method
Avg
IN-1k
Cifar10
Cifar100
CalTech
FER
Pets
DTD
RESISC
EuroSAT
PCAM
IN-S
IN-O
ViT-B/16
——
-
OpenAI CLIP
61.82
68.35
90.00
65.61
82.17
46.39
88.99
44.95
58.19
55.91
50.73
48.24
42.30
ImageNet-1k
2
KUEA
62.07
68.52
90.69
67.08
82.14
46.45
89.21
45.00
58.14
55.26
51.01
48.23
43.15
CC3M
2
KUEA
61.99
68.40
90.27
66.52
82.19
46.53
89.02
45.05
58.33
56.00
51.12
48.25
42.15
COCO Caption
1
CLIP-Refine
62.99
67.74
90.50
68.28
83.40
49.80
89.75
44.68
59.35
56.20
55.18
48.14
42.90
CC3M
1
CLIP-Refine
61.59
63.30
89.56
67.25
84.29
49.47
81.47
44.63
59.97
59.46
54.23
44.28
41.20
CC3M
2
CLIP-Refine
61.10
62.71
88.46
66.71
83.96
47.80
80.76
44.36
59.21
60.17
54.60
43.67
40.75
Appendix
Table 14: Top-1 zero-shot classification accuracy on 12 public benchmarks. The best and second-best results in each model block are highlighted in bold and underline , respectively.
Model
Dataset
epochs
Method
MSCOCO
Flickr30K
Image-to-Text
Text-to-Image
Image-to-Text
Text-to-Image
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
ViT-B/16
——
-
OpenAI CLIP
48.16
72.58
81.88
31.47
55.62
66.73
74.70
92.90
96.40
57.10
81.36
88.20
ImageNet-1k
2
KUEA
48.48
72.98
82.24
31.99
56.35
67.21
75.00
93.10
96.60
57.62
81.82
88.66
CC3M
2
KUEA
48.00
72.64
81.76
31.47
55.79
66.71
75.00
92.80
96.60
57.00
81.46
88.38
COCO Caption
1
CLIP-Refine
53.96
79.02
86.60
37.67
63.83
74.26
80.90
95.40
97.50
64.64
87.34
92.46
Appendix
Table 15: Zero-shot image-text and text-to-image retrieval results on MSCOCO and Flickr30K benchmarks. The best and second-best results in each model block are highlighted in bold and underline , respectively.
Model
Dataset
epochs
Method
Avg
IN-1K
SVHN
GTSRB
CLEVR Dist.
CLEVR Counts
ViT-B/16
——
-
OpenAI CLIP
44.70
66.34
45.19
56.94
31.37
23.68
IN-1k
2
KUEA
44.65
66.82
47.85
58.19
28.34
22.07
CC3M
2
KUEA
44.72
66.46
46.22
57.48
30.83
22.62
COCO Caption
1
CLIP-Refine
42.01
67.51
44.53
56.60
18.58
22.83
CC3M
1
CLIP-Refine
44.89
66.22
45.61
56.94
31.38
24.28
CC3M
2
CLIP-Refine
43.52
65.74
45.20
56.76
25.89
23.99
Appendix
Table 16: Linear probe evaluation results on various datasets. The best and second-best results in each model block are highlighted in bold and underline , respectively.
Model
Dataset
epochs
Method
Avg
Orientation and
Presence of Specific
State and Condition
Quantity and
Positional and
Color and Appearance
Structural and
Texts
Viewpoint and Perspective
ViT-B/16
——
-
OpenAI CLIP
12.59
6.67
0.00
26.67
13.33
13.33
20.00
13.33
0.00
20.00
ImageNet-1k
2
KUEA
11.85
0.00
0.00
20.00
13.33
13.33
26.67
13.33
0.00
20.00
CC3M
2
KUEA
14.07
6.67
0.00
26.67
20.00
13.33
20.00
20.00
0.00
20.00
COCO Caption
1
CLIP-Refine
17.04
13.33
6.67
20.00
0.00
20.00
46.67
33.33
0.00
13.33
CC3M
1
CLIP-Refine
19.26
13.33
0.00
26.67
6.67
26.67
53.33
33.33
0.00
13.33
CC3M
2
CLIP-Refine
20.74
13.33
0.00
33.33
20.00
20.00
40.00
33.33
6.67
20.00
Appendix
Table 17: Detailed MMVP benchmark results (single runs; each category has only 15 pairs, so one pair corresponds to 6.67 points). ‡ : per-category breakdown of one seed; on ViT-B/16 this seed is the maximum of three (3-seed mean 14.81±2.97 ), and on ViT-L/14 it is one of three seeds (3-seed mean 24.20±0.86 ). The multi-seed means in Table 2 are the figures we report; per-category and single-run differences are not used for any conclusion. The best results in each model block are highlighted in bold .
Method
Enhancement signal
IN-1K zero-shot
Retrieval (as reported)
MMVP-VLM
DIVA ( Wang et al., 2025 )
diffusion feedback
75.5→75.5
COCO R@1 I → T 56.4→56.7 , T → I 36.5→36.6
19.3→25.9
GenHancer ( Ma et al., 2025 )
generative reconstruction
75.5→75.6†
COCO text retrieval R@5 79.2→79.4
31.9
un2CLIP ( Li et al., 2025 )
inverted unCLIP
75.5→62.4
COCO T → I R@5 61.0→65.5 ; Flickr T → I R@5 87.3→90.1
32.6
ComCLIP (ours, CC3M)
contrastive + anchor + RKD, 1 epoch
75.54→76.10
COCO R@1 I → T 50.04→54.56 , T → I 34.29→39.40 ; Flickr T → I R@5 83.48→88.22
17.78→24.20±0.86
Appendix
Table 18: Comparison with generative-model-based enhancement methods on OpenAI CLIP ViT-L/14 (224px), using the numbers published by each paper (base → post-trained, as reported in that paper). Evaluation protocols and base values differ across papers (e.g., retrieval R@1 base values), so the rows are not strictly comparable; each row should be read as the change relative to its own base. “–”: not reported in the form used here. † A third-party reproduction from GenHancer’s official checkpoints reports ImageNet-1K 40.2 ( Li et al., 2025 ) ; the two sources disagree. ComCLIP MMVP is the 3 -seed mean (Table 2 ).