cs.CVJul 15, 2026

AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization

Authors: Yiyang YaoShanglin LiuJianming LvChengjun WangJinyi LiYuchan JieZhihua Jin

Organizations: School of Computer Science and Technology, South China University of Technology, Guangzhou, China

Abstract

Contrastive Language-Image Pretraining learns a shared representation space through large-scale contrastive learning. However, existing methods that enforce global consistency regularization overlook a key challenge: the inherent information asymmetry between images and text: captions typically describe only one specific aspect of an image, thus images with similar visual content can be paired with completely divergent textual content and semantic information. Consequently, global regularizers inadvertently impose constraints between visually similar images whose captions describe divergent aspects, introducing semantic distortion into the representation space. We propose AspectCLIP, a framework that reformulates consistency regularization to respect this one-to-many structure. AspectCLIP first partitions training samples into attribute clusters based on textual similarity to identify aspect-coherent groups, then applies full cyclic consistency within each cluster while restricting cross-cluster regularization to prototype-level comparisons. This aspect-guided regularization enforces strict geometric alignment only when images and texts describe a consistent facet, while allowing flexibility across divergent aspects. Extensive experiments on downstream tasks demonstrate that AspectCLIP consistently outperforms traditional methods and achieves a more structured representation space.

Explore similar work

CardsList
  1. CLIMP: Contrastive Language-Image Mamba Pretraining

    Jan 11, 2026Nimrod Shabtay, Itamar Zimerman, Eli Schwartz +1Vision TransformerCross-Modal Learning

  2. Contrastive vision-language learning with paraphrasing and negation

    Nov 20, 2025Kwun Ho Ngan, Saman Sadeghi Afgeh, Joe Townsend +1NegationParaphrasing