Synthetic hard negatives generated in the representation space have proven effective for unimodal self-supervised learning, but transferring this idea to vision-language pretraining is not straightforward. We analyze six representation-space synthesis strategies and identify two failure modes in their transfer to vision-language pretraining: cross-modal constructions that produce overly easy negatives or pull them toward the query, and intra-modal constructions that incorporate the matched positive. We also observe logit-scale saturation when training with synthetic hard negatives and a learnable temperature, and find that fixing the temperature improves downstream performance. Using this geometric analysis we propose SNAP, which generates intra-modal hard negatives that never involve the positive from either modality, avoiding both failure modes entirely. SNAP is model-agnostic, requires no external generative models, and adds less than 10% training time overhead. Evaluated on top of CLIP and FLIP across multiple architectures and datasets, SNAP delivers consistent improvements on zero-shot retrieval, zero-shot classification, and linear probe evaluation.
Figures & tables
Figure 1 : Architecture comparison for the image-to-text direction. (a) Standard CLIP encodes images X and captions C into embeddings V∈Rd×∣B∣ and T∈Rd×∣B∣ , computing the similarity matrix V⊤T∈∣B∣×∣B∣ using only batch negatives. (b) Our method generates synthetic text negatives T~∈Rd×∣Stxt∣ directly in the representation space and concatenates them with batch embeddings to form an augmented similarity matrix V⊤[T∣T~]∈∣B∣×(∣B∣+∣Stxt∣) . In the similarity matrices, green cells denote positive pairs, white cells denote batch negatives, and purple cells denote synthetic negatives. The text-to-image direction follows an analogous procedure with synthetic image negatives ( cf . Sec. 3.2 ).
Figure 2 : Geometric intuition of failure modes in vision-language pretraining. (a) Cross-modal failure: Blending image and text embeddings creates samples in the “modality gap” that are trivially rejected ( Sec. 4.1 ). (b) Positive Leakage: Including the positive pair in synthesis creates negatives that contain the actual signal, causing gradient conflict ( Sec. 4.1 ). (c) SNAP: By generating intra-modal negatives and filtering the positive anchor, we create high-quality hard negatives that respect the modality geometry ( Sec. 4.3 ).
Figure 3Figure 4
Figure 6 : PCA subspace of synthetic negatives (ViT-B/32, CC3M). (a) Cross-modal negatives land in the gap void between modality cones. (b) Intra-modal negatives involving the positive cluster near the positive anchor, causing leakage. (c) SNAP negatives ( s=3 , s=4 ) are confined within the text cone and away from the positive.
Variant
IN-val
IN-v2
Random-neg. dup.
10.0
8.0
Hard-neg. dup.
10.4
8.5
Hard-neg. wt.
10.5
8.7
s=3
10.8
9.2
s=4
10.6
8.8
s=3,4 (i2t only)
10.7
8.8
Table 4 : Synthesis and directionality ablations on CC3M.
Data
Method
Text Retrieval
Image Retrieval
Flickr30k
MSCOCO
Flickr30k
MSCOCO
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
ViT-B/32
CC3M
SigLIP [ 75 ]
7.2 .2
19.8 .3
26.5 .1
3.1 .3
10.0 .1
15.7 .2
5.3 .1
15.2 .2
22.3 .1
2.8 .1
9.0 .1
14.0 .1
NegCLIP [ 74 ] ‡
4.9
13.8
19.6
1.7
6.3
10.4
4.8
13.0
18.5
2.0
6.6
10.6
LaCLIP [ 10 ] ‡
3.7
10.9
16.0
1.6
5.1
8.9
3.5
10.8
16.0
1.7
6.0
9.5
Table 5 : Zero-shot image-text retrieval on Flickr30k and MSCOCO. We report Recall@K (R@K) metrics for text and image retrieval. Symbols : ‡ Taken from [ 43 ]
Data
Model
Food-101
CIFAR-10
CIFAR-100
SUN397
Cars
Aircraft
DTD
Pets
Caltech-101
Flowers
ImageNet
ViT-B/32
CC3M
SigLIP [ 75 ]
3.8 .3
23.8 .3
8.1 .2
12.4 .2
0.8 .2
0.9 .1
4.4 .2
4.1 .1
26.6 .2
4.2 .2
5.6 .1
NegCLIP [ 74 ] ‡
–
–
8.2
–
–
1.0
6.8
4.1
29.4
3.7
4.6
LaCLIP [ 10 ] ‡
–
–
8.0
–
–
0.8
5.5
3.1
24.6
3.4
3.7
TripletCLIP [ 43 ] ‡
–
–
11.8
–
–
0.9
8.9
5.7
38.5
8.2
7.3
CLIP [ 47 ]
5.9 .2
28.7 .2
10.0 .3
14.5 .1
0.8 .3
1.9 .3
8.2 .2
7.0 .3
41.5 .3
5.9 .1
8.0 .3
Table 6 : Zero-shot classification. Top-1 accuracy on ImageNet and 10 downstream datasets spanning fine-grained, scene, and object recognition. ‡ Taken from [ 43 ] .
Data
Model
Food-101
CIFAR-10
CIFAR-100
SUN397
Cars
Aircraft
DTD
Pets
Caltech-101
Flowers
ImageNet
ViT-B/32
CC3M
SigLIP [ 75 ]
35.4 .1
72.5 .1
48.4 .3
29.4 .1
9.5 .1
13.6 .2
31.9 .2
34.0 .2
61.6 .2
50.8 .2
30.2 .1
CLIP [ 47 ]
40.7 .3
77.2 .3
54.6 .1
32.7 .1
13.2 .2
20.1 .3
38.3 .3
40.7 .1
71.1 .1
61.5 .1
33.9 .1
+ SNAP
41.3 .2
78.5 .2
55.2 .3
33.1 .3
13.9 .1
20.7 .1
38.9 .3
41.4 .3
72.0 .1
62.7 .2
34.7 .1
FLIP [ 29 ]
40.8 .1
74.8 .2
52.8 .2
34.7 .3
13.9 .2
17.9 .2
39.8 .2
37.6 .2
69.7 .2
61.1 .2
32.9 .2
+ SNAP
41.4 .1
75.6 .2
53.6 .1
36.1 .3
14.5 .1
18.7 .1
41.3 .1
37.5 .2
70.8 .1
62.0 .2
33.5 .3
Table 7 : Linear probe evaluation. Top-1 accuracy of a linear classifier trained on frozen image encoder features across ImageNet and 10 transfer datasets.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Metric
Categories
Train Size
Test Size
Food-101 [ 2 ]
Accuracy
101
75,750
25,250
CIFAR-10 [ 25 ]
Accuracy
10
50,000
10,000
CIFAR-100 [ 25 ]
Accuracy
100
50,000
10,000
SUN397 [ 66 ]
Accuracy
397
19,850
19,850
Stanford Cars [ 24 ]
Accuracy
196
8,144
8,041
FGVC Aircraft [ 36 ]
Mean per class
100
6,667
3,333
Appendix
Table 8 : Details of downstream classification datasets.
Table 9 : Hyperparameters. (a) CLIP/FLIP training configuration, identical for ViT-B/16, ViT-B/32, and RN-50 on both CC3M and CC12M. (b) SNAP synthetic negative generation parameters.
Figure 7 : Comparison of vision-language pretraining methods using synthetic data. Top: Existing input-space methods rely on external foundation models (e.g., LLMs for text, Diffusion for images) to generate synthetic data, incurring significant computational overhead and risk of data leakage. Bottom: SNAP operates directly in the representation space , using in-batch embeddings to symmetrically generate hard negatives for both modalities.
Table 13
Accuracy
Paired gain Δ (pp)
Backbone
Baseline
Baseline
+SNAP
Seed 1
Seed 2
Seed 3
Mean ± std.
ViT-B/32
CLIP
8.0 .3
8.2 .2
+0.46
+0.02
+0.06
+0.18±0.24
FLIP
6.8 .3
7.2 .1
+0.10
+0.27
+0.74
+0.37±0.33
ViT-B/16
CLIP
10.4 .2
11.0 .2
+1.07
+0.30
+0.37
+0.58±0.43
FLIP
10.3 .3
11.0 .2
+0.88
+1.16
+0.12
+0.72±0.54
RN-50
CLIP
13.6 .1
14.7 .2
+0.99
+1.05
+1.32
+1.12±0.18
Appendix
Table 12: Paired seed-wise improvements on CC3M. ImageNet zero-shot accuracy summaries, with standard deviations shown as subscripts. Each paired gain Δ is the accuracy with SNAP minus the baseline accuracy for the same seed, in percentage points. The final column reports the mean and sample standard deviation of the three paired gains.
ARO
Winoground
Method
Relation
Attribute
Text
Image
Group
CLIP
51.2
52.1
24.5
9.8
6.3
+ SNAP
54.8
56.3
25.2
9.6
6.5
Appendix
Table 13 : Compositional evaluation. Scores (%) for ViT-B/16 pretrained on CC12M.