Synthetic hard negatives generated in the representation space have proven effective for unimodal self-supervised learning, but transferring this idea to vision-language pretraining is not straightforward. We analyze six representation-space synthesis strategies and identify two failure modes in their transfer to vision-language pretraining: cross-modal constructions that produce overly easy negatives or pull them toward the query, and intra-modal constructions that incorporate the matched positive. We also observe logit-scale saturation when training with synthetic hard negatives and a learnable temperature, and find that fixing the temperature improves downstream performance. Using this geometric analysis we propose SNAP, which generates intra-modal hard negatives that never involve the positive from either modality, avoiding both failure modes entirely. SNAP is model-agnostic, requires no external generative models, and adds less than 10% training time overhead. Evaluated on top of CLIP and FLIP across multiple architectures and datasets, SNAP delivers consistent improvements on zero-shot retrieval, zero-shot classification, and linear probe evaluation.
Figures & tables
Figure 1 : Architecture comparison for the image-to-text direction. (a) Standard CLIP encodes images X and captions C into embeddings V∈Rd×∣B∣ and T∈Rd×∣B∣ , computing the similarity matrix V⊤T∈∣B∣×∣B∣ using only batch negatives. (b) Our method generates synthetic text negatives T~∈Rd×∣Stxt∣ directly in the representation space and concatenates them with batch embeddings to form an augmented similarity matrix V⊤[T∣T~]∈∣B∣×(∣B∣+∣Stxt∣) . In the similarity matrices, green cells denote positive pairs, white cells denote batch negatives, and purple cells denote synthetic negatives. The text-to-image direction follows an analogous procedure with synthetic image negatives ( cf . Sec. 3.2 ).
Figure 2 : Geometric intuition of failure modes in vision-language pretraining. (a) Cross-modal failure: Blending image and text embeddings creates samples in the “modality gap” that are trivially rejected ( Sec. 4.1 ). (b) Positive Leakage: Including the positive pair in synthesis creates negatives that contain the actual signal, causing gradient conflict ( Sec. 4.1 ). (c) SNAP: By generating intra-modal negatives and filtering the positive anchor, we create high-quality hard negatives that respect the modality geometry ( Sec. 4.3 ).
Figure 3Figure 4
Figure 6 : PCA subspace of synthetic negatives (ViT-B/32, CC3M). (a) Cross-modal negatives land in the gap void between modality cones. (b) Intra-modal negatives involving the positive cluster near the positive anchor, causing leakage. (c) SNAP negatives ( s=3 , s=4 ) are confined within the text cone and away from the positive.
Variant
IN-val
IN-v2
Random-neg. dup.
10.0
8.0
Hard-neg. dup.
10.4
8.5
Hard-neg. wt.
10.5
8.7
s=3
10.8
9.2
s=4
10.6
8.8
s=3,4 (i2t only)
10.7
8.8
Table 4 : Synthesis and directionality ablations on CC3M.
Data
Method
Text Retrieval
Image Retrieval
Flickr30k
MSCOCO
Flickr30k
MSCOCO
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
ViT-B/32
CC3M
SigLIP [ 75 ]
7.2 .2
19.8 .3
26.5 .1
3.1 .3
10.0 .1
15.7 .2
5.3 .1
15.2 .2
22.3 .1
2.8 .1
9.0 .1
14.0 .1
NegCLIP [ 74 ] ‡
4.9
13.8
19.6
1.7
6.3
10.4
4.8
13.0
18.5
2.0
6.6
10.6
LaCLIP [ 10 ] ‡
3.7
10.9
16.0
1.6
5.1
8.9
3.5
10.8
16.0
1.7
6.0
9.5
Table 5 : Zero-shot image-text retrieval on Flickr30k and MSCOCO. We report Recall@K (R@K) metrics for text and image retrieval. Symbols : ‡ Taken from [ 43 ]
Data
Model
Food-101
CIFAR-10
CIFAR-100
SUN397
Cars
Aircraft
DTD
Pets
Caltech-101
Flowers
ImageNet
ViT-B/32
CC3M
SigLIP [ 75 ]
3.8 .3
23.8 .3
8.1 .2
12.4 .2
0.8 .2
0.9 .1
4.4 .2
4.1 .1
26.6 .2
4.2 .2
5.6 .1
NegCLIP [ 74 ] ‡
–
–
8.2
–
–
1.0
6.8
4.1
29.4
3.7
4.6
LaCLIP [ 10 ] ‡
–
–
8.0
–
–
0.8
5.5
3.1
24.6
3.4
3.7
TripletCLIP [ 43 ] ‡
–
–
11.8
–
–
0.9
8.9
5.7
38.5
8.2
7.3
CLIP [ 47 ]
5.9 .2
28.7 .2
10.0 .3
14.5 .1
0.8 .3
1.9 .3
8.2 .2
7.0 .3
41.5 .3
5.9 .1
8.0 .3
Table 6 : Zero-shot classification. Top-1 accuracy on ImageNet and 10 downstream datasets spanning fine-grained, scene, and object recognition. ‡ Taken from [ 43 ] .
Data
Model
Food-101
CIFAR-10
CIFAR-100
SUN397
Cars
Aircraft
DTD
Pets
Caltech-101
Flowers
ImageNet
ViT-B/32
CC3M
SigLIP [ 75 ]
35.4 .1
72.5 .1
48.4 .3
29.4 .1
9.5 .1
13.6 .2
31.9 .2
34.0 .2
61.6 .2
50.8 .2
30.2 .1
CLIP [ 47 ]
40.7 .3
77.2 .3
54.6 .1
32.7 .1
13.2 .2
20.1 .3
38.3 .3
40.7 .1
71.1 .1
61.5 .1
33.9 .1
+ SNAP
41.3 .2
78.5 .2
55.2 .3
33.1 .3
13.9 .1
20.7 .1
38.9 .3
41.4 .3
72.0 .1
62.7 .2
34.7 .1
FLIP [ 29 ]
40.8 .1
74.8 .2
52.8 .2
34.7 .3
13.9 .2
17.9 .2
39.8 .2
37.6 .2
69.7 .2
61.1 .2
32.9 .2
+ SNAP
41.4 .1
75.6 .2
53.6 .1
36.1 .3
14.5 .1
18.7 .1
41.3 .1
37.5 .2
70.8 .1
62.0 .2
33.5 .3
Table 7 : Linear probe evaluation. Top-1 accuracy of a linear classifier trained on frozen image encoder features across ImageNet and 10 transfer datasets.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Metric
Categories
Train Size
Test Size
Food-101 [ 2 ]
Accuracy
101
75,750
25,250
CIFAR-10 [ 25 ]
Accuracy
10
50,000
10,000
CIFAR-100 [ 25 ]
Accuracy
100
50,000
10,000
SUN397 [ 66 ]
Accuracy
397
19,850
19,850
Stanford Cars [ 24 ]
Accuracy
196
8,144
8,041
FGVC Aircraft [ 36 ]
Mean per class
100
6,667
3,333
Appendix
Table 8 : Details of downstream classification datasets.
Table 9 : Hyperparameters. (a) CLIP/FLIP training configuration, identical for ViT-B/16, ViT-B/32, and RN-50 on both CC3M and CC12M. (b) SNAP synthetic negative generation parameters.
Figure 7 : Comparison of vision-language pretraining methods using synthetic data. Top: Existing input-space methods rely on external foundation models (e.g., LLMs for text, Diffusion for images) to generate synthetic data, incurring significant computational overhead and risk of data leakage. Bottom: SNAP operates directly in the representation space , using in-batch embeddings to symmetrically generate hard negatives for both modalities.
Table 13
Accuracy
Paired gain Δ (pp)
Backbone
Baseline
Baseline
+SNAP
Seed 1
Seed 2
Seed 3
Mean ± std.
ViT-B/32
CLIP
8.0 .3
8.2 .2
+0.46
+0.02
+0.06
+0.18±0.24
FLIP
6.8 .3
7.2 .1
+0.10
+0.27
+0.74
+0.37±0.33
ViT-B/16
CLIP
10.4 .2
11.0 .2
+1.07
+0.30
+0.37
+0.58±0.43
FLIP
10.3 .3
11.0 .2
+0.88
+1.16
+0.12
+0.72±0.54
RN-50
CLIP
13.6 .1
14.7 .2
+0.99
+1.05
+1.32
+1.12±0.18
Appendix
Table 12: Paired seed-wise improvements on CC3M. ImageNet zero-shot accuracy summaries, with standard deviations shown as subscripts. Each paired gain Δ is the accuracy with SNAP minus the baseline accuracy for the same seed, in percentage points. The final column reports the mean and sample standard deviation of the three paired gains.
ARO
Winoground
Method
Relation
Attribute
Text
Image
Group
CLIP
51.2
52.1
24.5
9.8
6.3
+ SNAP
54.8
56.3
25.2
9.6
6.5
Appendix
Table 13 : Compositional evaluation. Scores (%) for ViT-B/16 pretrained on CC12M.
Contrastive learning relies on informative negatives to shape the representation space, yet obtaining hard negatives is costly, often requiring large batch sizes or extensive memory banks. We propose SynCo (Synthetic negatives in Contrastive learning), an approach that synthesizes hard negatives directly in the representation space from cached queue embeddings, with no additional forward passes or input-space processing. We find that six lightweight synthesis strategies, exhaustively covering the geometric, stochastic, and adversarial perturbation families, consistently improve learned representations at negligible computational cost. Although applicable to any InfoNCE-based contrastive objective, we demonstrate SynCo within the MoCo framework. On ImageNet ILSVRC-2012 linear evaluation at 200 epochs, SynCo yields improvements of +0.4% over MoCo-v2 and +1.0% over MoCHi. Unlike MoCHi, which degrades at extended pretraining schedules (underperforming MoCo-v2 by 2.4% at 800 epochs), SynCo does not: with a simple synthetic negative schedule, performance improves by +0.5% over MoCo-v2 at 800 epochs. SynCo also transfers well to a range of downstream tasks.
Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embeddings, rendering them insensitive to negation. We attribute this failure to a phenomenon we call Representational Collapse: by tracking compositional divergence and visual alignment across the CLIP text encoder, we show that middle layers build compositional syntax, but the final layers collapse this structure as visual alignment rises, producing a syntax-blind final representation. To recover the lost negation signal without altering pretrained weights, we propose PeakPatch, a lightweight post-hoc correction system that intercepts the encoder at its compositional peak. An Embedding Correction Network (ECN) uses cross-attention to extract a negation-specific signal from the peak layer, anchored to a stable baseline, and predicts a deviation vector that re-injects the lost syntax into the final-layer embedding space. A complementary Score Correction Network (SCN) predicts bounded scalar score offsets for discriminative tasks. Both modules are trained jointly end-to-end while all CLIP parameters remain frozen, adding only 5.2M parameters (3.5% of the backbone) and preserving the standard cosine similarity interface. On NegBench, PeakPatch achieves 74.3% on COCO MCQ (+35.1 over CLIP, +17.8 over the best encoder fine-tuning method) and 65.5% on VOC MCQ, while outperforming all fine-tuning baselines on fully out-of-distribution negation retrieval despite training only 3.5% of the parameters. The corrected embeddings also transfer to text-to-image generation (+18.4 negation score) and generalize across ViT-B/32, ViT-L/14, and SigLIP backbones. Project URL: https://stevencylu.github.io/PeakPatch/.
We introduce ViTAMINS, a method that integrates synthetic hard negatives into unsupervised vision transformer pretraining to improve representation quality. Our approach is thoroughly benchmarked on ImageNet and transfer learning, image retrieval, copy detection, and image, video segmentation tasks. Notably, our proposed negatives give rise to emergent properties, where learned representations contain explicit information about the semantic content of an image and serve as excellent classifiers (up to +11.3% over baselines). ViTAMINS achieves these benefits through simple modifications to existing contrastive frameworks and outperforms competing methods while being more resource efficient, e.g., our ViT-B surpasses V-JEPA with ViT-L. Our findings motivate reconsidering contrastive learning as a simpler yet powerful alternative to dominant generative and self-distillation approaches.
Nikos Giakoumoglou, Andreas Floros, Kleanthis-Marios Papadopoulos +1