cs.CVOct 8, 2026

Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder

Authors: Zidan Wang, Yaqian Li, Xiaokai Zhang, Kaiwen Long, Kun He, Hanpeng Liu

Organizations: Huazhong University of Science and Technology · Li Auto Inc.

Abstract

CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a lightweight route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that, for the InfoNCE objective, the reported forgetting is driven primarily not by insufficient negatives but by an inappropriate magnitude of the contrastive temperature ττ: with ττ set sufficiently small, contrastive post-training improves rather than degrades the pretrained CLIP, which we explain through the temperature dependence of the InfoNCE gradient. Building on this finding, we propose \textbf{ComCLIP}, a lightweight single-epoch post-training recipe that freezes CLIP's text encoder---so the refined vision encoder is a drop-in replacement with unchanged architecture and inference cost---and trains the vision encoder with a properly-tempered contrastive loss, an MSE anchoring loss against the original CLIP, and a relational distillation loss from DINOv2. Over multiple seeds, ComCLIP matches the self-distillation baseline CLIP-Refine on zero-shot classification while significantly improving the transferability of visual features, measured by linear probing (48.9948.99 vs.\ 42.2842.28 on ViT-B/16), and on ViT-L/14 it also improves MMVP over CLIP-Refine (24.2024.20 vs.\ 19.0119.01); CLIP-Refine remains stronger on image-text retrieval. Used as a drop-in vision encoder for LLaVA-1.5-7B without re-aligning the projector or LLM, ComCLIP yields no net change across 88 VLM benchmarks, i.e., the refinement does not break downstream compatibility. Code and models are available at https://github.com/showstarpro/ComCLIP.git.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Contrastive vision-language learning with paraphrasing and negation

    Nov 20, 2025Kwun Ho Ngan, Saman Sadeghi Afgeh, Joe Townsend +1VLM RobustnessVLM Reasoning

  2. CLIMP: Contrastive Language-Image Mamba Pretraining

    Jan 11, 2026Nimrod Shabtay, Itamar Zimerman, Eli Schwartz +1Cross-Modal LearningVLM Robustness

  3. What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features

    Jul 25, 2026Chen-Yi Lu, Yueh-Shao Chen, Somali ChaterjiVision-Language ModelsVLM Adaptation