cs.CVOct 7, 2026

What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining?

Authors: Nikos Giakoumoglou, Paschalis Giakoumoglou, Andreas Floros, Kleanthis Marios Papadopoulos, Tania Stathaki

Organizations: Imperial College London, London, UK · Centre for Research and Technology Hellas, Thessaloniki, Greece

Abstract

Synthetic hard negatives generated in the representation space have proven effective for unimodal self-supervised learning, but transferring this idea to vision-language pretraining is not straightforward. We analyze six representation-space synthesis strategies and identify two failure modes in their transfer to vision-language pretraining: cross-modal constructions that produce overly easy negatives or pull them toward the query, and intra-modal constructions that incorporate the matched positive. We also observe logit-scale saturation when training with synthetic hard negatives and a learnable temperature, and find that fixing the temperature improves downstream performance. Using this geometric analysis we propose SNAP, which generates intra-modal hard negatives that never involve the positive from either modality, avoiding both failure modes entirely. SNAP is model-agnostic, requires no external generative models, and adds less than 10% training time overhead. Evaluated on top of CLIP and FLIP across multiple architectures and datasets, SNAP delivers consistent improvements on zero-shot retrieval, zero-shot classification, and linear probe evaluation.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SynCo: Synthetic Hard Negatives for Contrastive Visual Representation Learning

    Oct 3, 2024Nikos Giakoumoglou, Tania StathakiContrastive LearningImagenet

  2. What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features

    Jul 25, 2026Chen-Yi Lu, Yueh-Shao Chen, Somali ChaterjiNegationRepresentational Collapse