Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality rather than by semantics. We propose ITO, a framework addressing this limitation through two complementary mechanisms with distinct roles. Multimodal multiple alignment enriches supervision by constructing diverse cross-modal correspondences from multi-view image augmentations, providing the primary source of discriminative gain. A lightweight training-time multimodal fusion module then acts as a training-time regularizer, encouraging the encoders to produce features that are compatible under fusion. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures while incurring training-time overhead only. Extensive experiments across pretraining scales from millions to billions of image--text pairs show that ITO outperforms strong contrastive baselines at moderate scale and yields consistent gains over CLIP under identical data, backbone, and compute at the 100M--1B scale, on classification, retrieval, and multimodal benchmarks. Controlled comparisons show that these gains are not explained by additional training compute or image augmentation alone. Our analysis further reveals that the contribution of fusion grows with data scale and that it stabilizes optimization, mitigating the late-stage overfitting commonly observed in aggressive contrastive learning. Code is available at https://github.com/showstarpro/ITO.
Figures & tables
Figure 1 : Overview of the proposed ITO training framework. Starting from standard image–text contrastive pretraining, ITO restructures supervision through multimodal multiple alignment and introduces a lightweight multimodal fusion module during training. Multiple augmented image–text pairs derived from the same sample are used to enrich instance-level alignment, while training-time fusion encourages encoders to produce features that are compatible under fusion, reducing modality-induced separation. Importantly, the fusion module is used only during training and is discarded at inference time, allowing ITO to retain a standard dual-encoder architecture for efficient deployment.
Pretraining Data
CLIP
SLIP
SigLIP
FLAIR
ITO
ITO_sub2
CC3M
19.3
21.4
19.3
23.2
24.9
24.9
CC12M
34.1
36.0
35.0
35.5
38.4
40.4
YFCC15M
30.5
–
–
–
34.5
35.4
Laion100M
53.5
51.3
–
–
56.1
–
DataComp-1B
54.2
–
–
–
56.6
57.0
Scaling: longer training and larger backbones on DataComp-1B
Table 1 : Average top-1 zero-shot classification accuracy across 26 public benchmarks following the EVA-CLIP Sun et al. (2023) protocol. All models use ViT-B/16 and are trained for 30 epochs unless otherwise noted. The upper block compares ITO against contrastive baselines under the same backbone and training schedule; the lower block reports scaling behavior with longer training (10 epochs) and a larger backbone (ViT-L/16) on DataComp-1B. Due to computational constraints, SigLIP and FLAIR are evaluated only on smaller datasets and SLIP up to Laion100M; CLIP is the only baseline run consistently across all pretraining scales, and at large scale the comparison is against CLIP under identical data, backbone, and compute. Per-benchmark results are provided in Appendix Table 12 . Best results are in bold .
YFCC15M
Laion100M
DataComp-1B (1ep)
DataComp-1B (10ep)
Method
IN-1k
Avg
IN-1k
Avg
IN-1k
Avg
IN-1k
Avg
CLIP
51.82
55.88
67.27
81.86
67.39
82.14
74.99
87.62
ITO
60.52
63.76
68.70
83.56
69.87
82.53
76.20
89.18
Table 2 : Linear probing top-1 accuracy of ITO and CLIP across pretraining scales. We report ImageNet-1K (IN-1k) and the average over five datasets (IN-1k, Flowers-102, Stanford Cars, Pets, Food-101). Per-dataset results, additional baselines, and ITO_sub2 variants are provided in Appendix Table 13 . Best results in each block are in bold .
CC12M
DataComp-1B (1 epoch)
MSCOCO
Flickr30k
MSCOCO
Flickr30k
Method
I → T
T → I
I → T
T → I
I → T
T → I
I → T
T → I
CLIP
34.17
23.04
62.23
45.74
47.08
29.51
72.68
54.73
ITO
42.30
29.62
72.49
54.50
49.26
31.30
75.94
56.79
ITO_sub2
43.94
30.51
72.29
56.65
49.50
31.89
73.37
56.79
Table 3 : Zero-shot image-text retrieval (Recall@1) on MSCOCO and Flickr30k. CC12M and DataComp-1B (ViT-B/16, 1 epoch) represent medium- and large-scale pretraining respectively. Comparisons against additional baselines (SLIP, SigLIP, FLAIR), all five pretraining scales, and full Recall@5/10 metrics are provided in Appendix Table 18 . Best results are in bold .
Dataset
Method
VQAv2
GQA
T-VQA
SciQA-I
VisWiz
MMB-en
MMB-cn
MMVet
POPE-r
POPE-p
POPE-a
MMMU
MMStar
100M
CLIP
66.53
56.36
47.55
65.44
36.80
53.69
46.13
18.30
84.09
82.90
76.40
33.60
28.87
ITO
68.47
57.23
48.71
66.29
44.25
54.12
46.65
19.80
84.74
83.63
77.30
33.90
29.13
1B
CLIP
70.42
57.93
50.24
65.00
45.28
48.45
55.67
18.40
83.63
82.71
78.63
34.00
29.33
ITO
73.19
59.99
50.89
66.24
42.92
50.52
58.85
21.90
85.46
84.23
80.46
33.90
31.13
Table 4 : The performance of ITO on a broad range of multimodal tasks. The best results are bold .
Dataset
Setting
ZS IN-1k
Linear IN-1k
MSCOCO I → T
MSCOCO T → I
15 M
OpenCLIP
36.4
51.8
26.4
15.1
+ Multi-view alignment ( λ=0 )
43.7 (+7.3)
60.2 (+8.4)
30.6 (+4.2)
19.6 (+4.5)
+ Fusion regularization ( λ=2 )
44.3 (+0.6)
60.5 (+0.3)
30.8 (+0.2)
19.5 ( − 0.1)
1B
OpenCLIP
63.1
67.4
47.1
29.5
+ Multi-view alignment ( λ=0 )
64.8 (+1.7)
69.6 (+2.2)
49.0 (+1.9)
30.7 (+1.2)
+ Fusion regularization ( λ=2 )
65.9 (+1.1)
69.9 (+0.3)
49.3 (+0.3)
31.3 (+0.6)
Table 5 : Decomposing the contributions of multi-view alignment and fusion regularization on YFCC15M and DataComp-1B (ViT-B/16, 1 epoch on 1 B). Δ values are gains over the previous row. Multi-view alignment provides the dominant gain at both scales; fusion regularization adds a smaller but consistent improvement that grows at billion scale.
Method
λ
Top-1
Baseline
-
36.4
ITO
0
43.7
ITO
2
44.3
ITO
4
44.0
ITO
6
43.2
ITO
8
43.6
Table 6 : Ablation study on YFCC15M. We analyze the effect of Multimodal Multiple Alignment and training-time multimodal fusion by varying the fusion weight λ (left) and the number of Sentence Blocks (right). Results are zero-shot Top-1 accuracy (%) on ImageNet-1K using ViT-B/16. The baseline is standard CLIP.
Model (epochs)
FLOPs
ZS
Lin.
T → I
CLIP (30)
1.0 ×
36.4
51.8
15.1
CLIP (60)
2.0 ×
37.3
52.8
15.4
ITO, λ=0 (30)
2.3 ×
43.7
60.2
19.6
ITO, λ=2 (30)
2.3 ×
44.3
60.5
19.5
ITO, λ=0 (60)
4.6 ×
44.6
61.7
18.7
ITO, λ=2 (60)
4.6 ×
45.3
62.1
19.6
Table 7 : Left: compute-matched comparison on YFCC15M (ViT-B/16); FLOPs are training FLOPs relative to CLIP at 30 epochs. Right: training time (h) and peak memory (GB/GPU) on 64 A800 GPUs, with ratios to OpenCLIP. Inference cost of ITO is identical to CLIP.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
CC3M Sharma et al. (2018)
CC12M Changpinyo et al. (2021)
YFCC15M Torralba and Efros (2011)
Laion100M Schuhmann et al. (2021)
DataComp-1B Gadre et al. (2023)
Origin
3.32M
12.42M
15.39M
361.02M
1.40B
Download
2.91M
10.97M
14.08M
107.75M
1.03B
Appendix
Table 8 : The original and actual downloaded number of image-text pairs in pretraining datasets.
Config
Value
Optimizer
AdamW
Learning rate
5e-4
Weight decay
0.1
Optimizer momentum
β1 =0.9, β2 =0.98
Batch size
16384(ViT-B), 8192(ViT-L)
Warm-up iterations
500
Appendix
Table 9 : Common ITO hyperparameters on Datacomp-1B.
Dataset
Fusion
ZS IN-1k
Linear IN-1k
MSCOCO I → T
MSCOCO T → I
CC3M
no fusion ( λ=0 )
23.0
—
—
—
cross-attention
23.0
47.5
19.7
13.6
self-attention (ours)
23.3
48.1
21.6
14.6
YFCC15M
no fusion ( λ=0 )
43.7
60.2
30.6
19.6
cross-attention
43.1
59.9
30.8
19.4
self-attention (ours)
44.3
60.5
30.8
19.6
Appendix
Table 10 : Self-attention vs cross-attention fusion (two-layer, λ=2 ). Cross-attention with an EOS-query design consistently underperforms self-attention and, on YFCC15M, falls below the no-fusion variant. Best results are in bold .
Table 12 : Top-1 zero-shot classification accuracy on 26 public benchmarks following the EVA-CLIP Sun et al. (2023) protocol. Benchmarks include ImageNet-1K Deng et al. (2009) , ImageNet-A Hendrycks et al. (2021b) , ImageNet-R Hendrycks et al. (2021a) , CIFAR-10/100 Krizhevsky et al. (2009) , Food-101 Bossard et al. (2014) , Pets Parkhi et al. (2012) , SUN397 Xiao et al. (2010) , FGVC Aircraft Maji et al. (2013) , EuroSAT Helber et al. (2019) , VOC2007 Everingham et al. (2015) , etc. The best results are highlighted in bold .
Figure 2 : Linear image classification of ITO and its variants pretrained on CC3M.
Figure 3 : Linear image classification of ITO and its variants pretrained on CC12M.
Dataset
Method
IN-1k
Flowers
Cars
Pets
Food
Avg
15M
CLIP
51.82
85.17
19.60
53.61
69.18
55.88
ITO
60.52
92.18
28.37
61.22
76.50
63.76
ITO_sub2
60.40
92.37
27.73
62.82
76.22
63.91
100M
CLIP
67.27
86.65
87.17
82.88
85.35
81.86
ITO
68.70
90.21
88.16
84.49
86.22
83.56
1B
CLIP
67.39
90.08
87.14
79.61
86.47
82.14
Appendix
Table 13 : Linear image classification of ITO and its variants pretrained on YFCC15M, Laion100M and DataComp-1B. The best results are highlighted in bold . † indicates models trained for 10 epochs using ViT-B/16. ∗ indicates models trained for 1 epochs using ViT-L/16.
ImageNet-1k
ImageNet-A
ImageNet-R
ImageNet-S
ImageNet-V2
ObjectNet
CIFAR-10
CIFAR-100
Flowers-102
Food-101
Pets
Stanford Cars
MNIST
Caltech
SUN397
FGVC Aircraft
Country-211
DTD
EuroSAT
FER2013
GTSRB
PCam
Rendered SST2
Resisc45
STL10
VOC2007
Avg
(a) CC3M Pre-training ViT-B
FLAIR
27.7
7.6
31.4
15.0
23.8
15.5
73.9
44.3
18.9
19.8
25.9
2.5
13.1
69.8
42.3
1.6
2.6
18.8
7.4
30.6
11.4
45.7
50.1
34.6
92.9
26.6
29.0
ITO_sub2
30.4
8.8
38.8
19.9
27.2
16.1
62.7
38.0
19.7
23.1
29.9
3.6
18.5
71.4
46.0
2.6
2.6
20.5
7.1
22.4
11.9
50.0
50.1
37.3
94.4
27.1
30.0
ITO_sub3
30.9
9.4
39.5
20.7
26.7
16.3
62.2
36.1
20.0
23.1
30.2
4.2
16.8
72.1
46.0
1.2
2.6
21.3
8.5
19.6
10.2
50.0
50.1
35.0
94.0
22.5
29.6
Appendix
Table 14 : Top-1 zero-shot classification accuracy on 26 public benchmarks following the EVA-CLIP Sun et al. (2023) protocol. Results are reported for models pretrained on CC3M-recap with both ViT-B/16 as vision encoders. The best results are highlighted in bold .
Dataset
Method
ImageNet-1k
Flowers-102
Stanford Cars
Pets
Food-101
CC3M-recap
FLAIR
45.65
76.57
25.05
67.05
59.38
ITO_sub2
52.43
73.64
23.96
65.09
64.15
ITO_sub3
53.05
74.30
25.32
65.55
64.70
Appendix
Table 15 : Linear image classification of ITO and its variants pretrained on CC3M-recap.
Dataset
Method
MSCOCO
Flickr30k
I → T
T → I
I → T
T → I
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
CC3M-recap
FLAIR
37.54
64.38
75.88
29.76
55.53
66.69
65.19
87.08
92.41
53.25
78.01
85.48
ITO_sub2
47.76
73.76
82.64
31.59
58.54
69.98
74.56
93.89
96.75
57.32
82.23
88.38
ITO_sub3
47.40
74.94
83.86
32.38
58.74
69.77
74.95
94.38
96.75
59.13
83.41
89.80
Appendix
Table 16 : Zero-shot image-text retrieval on validation splits for standard benchmarks (MSCOCO and Flickr30k).
Dataset
Method
DOCCI
I → T
T → I
R@1
R@5
R@10
R@1
R@5
R@10
CC3M-recap
FLAIR
20.16
43.56
54.66
10.65
24.23
31.60
ITO_sub2
27.82
54.14
65.90
10.70
23.69
30.84
ITO_sub3
27.54
53.68
65.30
10.68
24.04
31.27
Appendix
Table 17 : Zero-shot image-text retrieval on validation splits for fine-grained retrieval benchmarks (DOCCI).
Dataset
Method
MSCOCO
Flickr30k
I → T
T → I
I → T
T → I
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
R@1
R@5
R@10
3M
CLIP
12.36
30.98
41.76
8.14
22.68
31.98
23.57
50.69
62.03
16.98
37.65
48.68
SigLIP
13.70
32.66
43.16
9.28
23.48
32.60
28.90
53.25
65.09
18.32
39.43
49.45
FLAIR
17.86
38.90
51.18
12.59
30.38
40.94
35.40
62.13
74.46
25.74
49.86
60.81
ITO
21.56
45.36
57.22
14.56
33.71
44.90
42.6
71.2
80.87
29.45
53.89
64.36
Appendix
Table 18 : Zero-shot image-text retrieval on validation splits for standard benchmarks (MSCOCO and Flickr30k). The best results are highlighted in bold . † indicates models trained for 10 epochs using ViT-B/16. ∗ indicates models trained for 1 epochs using ViT-L/16.
Dataset
Method
DOCCI
I → T
T → I
R@1
R@5
R@10
R@1
R@5
R@10
3M
CLIP
4.10
12.56
18.72
2.13
6.71
10.07
SigLIP
4.98
14.38
20.88
2.28
6.91
10.32
FLAIR
7.12
19.12
28.00
3.23
9.70
14.13
ITO
8.24
22.00
31.64
3.48
9.97
14.39
Appendix
Table 19 : Zero-shot image-text retrieval on validation splits for fine-grained retrieval benchmarks (DOCCI). The best results are highlighted in bold . † indicates models trained for 10 epochs using ViT-B/16. ∗ indicates models trained for 1 epochs using ViT-L/16.
Dataset
Setting
Top-1
Top-3
Top-5
Top-10
ZS IN-1k
15M
ITO ( λ=0 )
32.1
31.9
33.7
35.3
43.7
ITO ( λ=2 )
32.7
32.4
34.3
35.7
44.3
Δ
+0.6
+0.5
+0.6
+0.4
+0.6
1B
ITO ( λ=0 )
46.1
45.1
47.5
49.1
64.8
ITO ( λ=2 )
46.9
45.7
48.0
49.8
65.9
Δ
+0.8
+0.6
+0.5
+0.7
+1.1
Appendix
Table 20 : Cross-modal consistency score (%) at top- K following CyCLIP Goel et al. (2022) . Δ is the change from adding fusion regularization ( λ=2 ) to the alignment-only variant ( λ=0 ). ZS IN-1k is reported alongside for reference.
Figure 4 : UMAP visualization. All models are trained on the CC3M dataset. For visualization, 8,192 image-text pairs are randomly sampled from CC12M. Blue points represent images, and red points represent texts. (a): CLIP exhibits a clear separation between modalities, with a distinct boundary between image and text embeddings. (b): FLAIR shows more compact text embeddings surrounded by image embeddings, likely due to its text-conditioned fusion mechanism. (c): ITO shows a star-shaped distribution in which image and text embeddings are more closely interleaved and the boundary between modalities is less pronounced in this projection.
Figure 5 : UMAP visualization. All models are trained on the CC3M dataset. For visualization, 8,192 image-text pairs are randomly sampled from CC12M. Blue points represent images, and red points represent texts.
Figure 6 : UMAP visualization. All models are trained on the CC3M-recap dataset. For visualization, 8,192 image-text pairs are randomly sampled from CC12M. Blue points represent images, and red points represent texts.
Figure 7 : UMAP visualization for DataComp-1B. All models are trained on the DataComp-1B dataset for 10 epochs. For visualization, 8,192 image-text pairs are randomly sampled from CC12M. Blue points represent images, red points represent texts, and green points denote unified multimodal tokens. .
Figure 8 : Training dynamics and overfitting behavior on YFCC15M. Zero-shot ImageNet-1K accuracy as a function of training epochs for CLIP, SLIP, and ITO variants using ViT-B/16. Both CLIP and SLIP exhibit late-stage performance degradation, indicating overfitting on moderately sized datasets. Introducing Multimodal Multiple Alignment alone ( λ=0 ) improves overall accuracy but does not fully prevent overfitting. In contrast, enabling training-time multimodal fusion ( λ=2 ) stabilizes training and maintains consistent zero-shot performance in later epochs, demonstrating the regularizing effect of fusion-based cross-modal interaction.
Figure 9 : Layer-wise attention visualization of CLIP and ITO on DataComp-1B. Each image corresponds to a different transformer layer, with the layer index indicated above each visualization. ITO shows progressively more focused and semantically consistent attention distributions than CLIP. Note that all visualizations are derived from held-out samples rather than those seen during pretraining.
Dataset
Method
Time (h)
Time ratio
Memory (GB)
Memory ratio
YFCC15M
OpenCLIP
3.6
1.0 ×
7.9
1.0 ×
ITO ( λ=0 )
8.2
2.3 ×
16.8
2.1 ×
ITO ( λ=2 )
8.2
2.3 ×
16.8
2.1 ×
DataComp-1B
OpenCLIP
8.1
1.0 ×
30.0
1.0 ×
ITO ( λ=0 )
∼18
2.2 ×
68.6
2.3 ×
ITO ( λ=2 )
∼18
2.2 ×
68.6
2.3 ×
Appendix
Table 21 : Training time (hours) and peak GPU memory (GB per GPU) of OpenCLIP and ITO. Ratios are relative to OpenCLIP at the same scale.
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model. This setup bottlenecks generative performance behind frozen embeddings. To bridge this gap, we revisit joint multimodal representation learning and generation to produce linearly interpolatable embeddings that are directly consumable by generative decoders. We present FLAT (Flexible-Length Aligned Transmodal representations), a representation pre-training framework that jointly optimizes a shared multimodal encoder alongside downstream text-to-image (T2I) and image-to-text (I2T) decoders. By combining contrastive alignment with bidirectional cross-modal generative objectives, FLAT ensures its representations function as both discriminative semantic descriptors and generative conditions. Architecturally, FLAT maps visual and textual inputs into a unified continuous 1D sequence space, applying nested dropout over prefix-K tokens to enable dynamic output lengths. A single pre-training stage allows FLAT to perform cross-modal retrieval and generation across variable prefix K, achieving a T2I GenEval score of 71.1. Task-specific fine-tuning aligns model performance with state-of-the-art baselines: 83.1 GenEval on T2I generation; 40.5 BLEU-4 and 138.6 CIDEr on MS-COCO image captioning; and Recall@5 scores of 86.8 (I2T) / 75.8 (T2I) on MS-COCO alongside 98.3 (I2T) / 93.6 (T2I) on Flickr30K. Finally, qualitative evaluations demonstrate that FLAT representations natively support linear interpolation, latent space arithmetic, and zero-shot composed retrieval.
The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world. Recent work exploits this convergence by aligning frozen pretrained vision and language models with lightweight alignment layers, but typically relies on contrastive losses and millions of paired samples. In this work, we ask whether meaningful alignment can be achieved with substantially less supervision. We introduce a semi-supervised setting in which pretrained unimodal encoders are aligned using a small number of image-text pairs together with large amounts of unpaired data. To address this challenge, we propose SOTAlign, a two-stage framework that first recovers a coarse shared geometry from limited paired data using a linear teacher, and then refines the alignment on unpaired samples via an optimal-transport-based divergence that transfers relational structure without overconstraining the target space. SOTAlign effectively leverages unpaired images and text, learning robust joint embeddings across datasets and encoder pairs, and significantly outperforming supervised and semi-supervised baselines. Code is available at https://github.com/ExplainableML/SOTAlign.
Simon Roschmann, Paul Krzakala, Sonia Mazelet +2
Helmholtz Munich · Technical University of Munich · Munich Center for Machine Learning +3
The platonic representation hypothesis suggests that sufficiently large models converge to a shared representation geometry, even across modalities. Motivated by this, we ask: Can the semantic knowledge of a language model efficiently improve a vision model? As an answer, we introduce TextTeacher, a simple auxiliary objective that injects text embeddings as additional information into image classification training. TextTeacher uses readily available image captions, a pre-trained and frozen text encoder, and a lightweight projection to produce semantic anchors that efficiently guide representations during training while leaving the inference-time model unchanged. On ImageNet with standard ViT backbones, TextTeacher improves accuracy by up to +2.7 percentage points (p.p.) and yields consistent transfer gains (on average +1.0 p.p.) under the same recipe and compute. It outperforms vision knowledge distillation, yielding more accuracy at a constant compute budget or similar accuracy, but 33% faster. Our analysis indicates that TextTeacher acts as a feature-space preconditioner, shaping deeper layers in the first stages of training, and aiding generalization by supplying complementary semantic cues. TextTeacher adds negligible overhead, requires no costly multimodal training of the target model and preserves the simplicity and latency of pure vision models. Project page with code and captions: https://nauen-it.de/publications/text-teacher
Tobias Christian Nauen, Stanislav Frolov, Brian Bernhard Moser +3
RPTU University Kaiserslautern-Landau · German Research Center for Artificial Intelligence (DFKI)