Contrastive learning relies on informative negatives to shape the representation space, yet obtaining hard negatives is costly, often requiring large batch sizes or extensive memory banks. We propose SynCo (Synthetic negatives in Contrastive learning), an approach that synthesizes hard negatives directly in the representation space from cached queue embeddings, with no additional forward passes or input-space processing. We find that six lightweight synthesis strategies, exhaustively covering the geometric, stochastic, and adversarial perturbation families, consistently improve learned representations at negligible computational cost. Although applicable to any InfoNCE-based contrastive objective, we demonstrate SynCo within the MoCo framework. On ImageNet ILSVRC-2012 linear evaluation at 200 epochs, SynCo yields improvements of +0.4% over MoCo-v2 and +1.0% over MoCHi. Unlike MoCHi, which degrades at extended pretraining schedules (underperforming MoCo-v2 by 2.4% at 800 epochs), SynCo does not: with a simple synthetic negative schedule, performance improves by +0.5% over MoCo-v2 at 800 epochs. SynCo also transfers well to a range of downstream tasks.
We introduce ViTAMINS, a method that integrates synthetic hard negatives into unsupervised vision transformer pretraining to improve representation quality. Our approach is thoroughly benchmarked on ImageNet and transfer learning, image retrieval, copy detection, and image, video segmentation tasks. Notably, our proposed negatives give rise to emergent properties, where learned representations contain explicit information about the semantic content of an image and serve as excellent classifiers (up to +11.3% over baselines). ViTAMINS achieves these benefits through simple modifications to existing contrastive frameworks and outperforms competing methods while being more resource efficient, e.g., our ViT-B surpasses V-JEPA with ViT-L. Our findings motivate reconsidering contrastive learning as a simpler yet powerful alternative to dominant generative and self-distillation approaches.
Nikos Giakoumoglou, Andreas Floros, Kleanthis-Marios Papadopoulos +1
Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embeddings, rendering them insensitive to negation. We attribute this failure to a phenomenon we call Representational Collapse: by tracking compositional divergence and visual alignment across the CLIP text encoder, we show that middle layers build compositional syntax, but the final layers collapse this structure as visual alignment rises, producing a syntax-blind final representation. To recover the lost negation signal without altering pretrained weights, we propose PeakPatch, a lightweight post-hoc correction system that intercepts the encoder at its compositional peak. An Embedding Correction Network (ECN) uses cross-attention to extract a negation-specific signal from the peak layer, anchored to a stable baseline, and predicts a deviation vector that re-injects the lost syntax into the final-layer embedding space. A complementary Score Correction Network (SCN) predicts bounded scalar score offsets for discriminative tasks. Both modules are trained jointly end-to-end while all CLIP parameters remain frozen, adding only 5.2M parameters (3.5% of the backbone) and preserving the standard cosine similarity interface. On NegBench, PeakPatch achieves 74.3% on COCO MCQ (+35.1 over CLIP, +17.8 over the best encoder fine-tuning method) and 65.5% on VOC MCQ, while outperforming all fine-tuning baselines on fully out-of-distribution negation retrieval despite training only 3.5% of the parameters. The corrected embeddings also transfer to text-to-image generation (+18.4 negation score) and generalize across ViT-B/32, ViT-L/14, and SigLIP backbones. Project URL: https://stevencylu.github.io/PeakPatch/.
Contrastive representation learning (CRL) underpins many modern foundation models. Despite recent theoretical progress, existing analyses suffer from several key limitations: (i) the statistical consistency of CRL remains poorly understood; (ii) available generalization bounds deteriorate as the number of negative samples increases, contradicting the empirical benefits of large negative sets; and (iii) the retrieval performance of CRL has received limited theoretical attention. In this paper, we develop a unified statistical learning theory for CRL. For downstream tasks, we evaluate retrieval quality using an AUC-type population criterion and show that the contrastive loss is \emph{statistically consistent} with optimal ranking. We further establish a \emph{calibration-style inequality} that quantitatively relates excess contrastive risk to excess retrieval suboptimality. For upstream training, we study both supervised and self-supervised contrastive objectives and derive generalization bounds of order O(1/m+1/n) and O(1/m+1/n), respectively, where m denotes the number of negative samples and n the number of anchor points. These bounds not only explain the empirical advantages of large negative sets but also reveal an explicit trade-off between m and n. Extensive experiments on large-scale vision--language models corroborate our theoretical predictions.