Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs. Across datasets, modalities, and unimodal models, we show that we can consistently align independently trained representations without pairs, and standard geometric alignment metrics accurately predict when this is possible. Nevertheless, we can naturally benefit from paired examples. In the very few-pair regime, our method substantially outperforms existing ones, while staying competitive with pair-based methods with more added examples. Finally, we demonstrate that the resulting alignments can enable text-to-image generation without paired examples. These results show that independently trained models often share enough geometry to establish cross-modal correspondence with little or no paired data.
Figures & tables
Figure 1: Shared geometry enables unpaired alignment. Across modalities, embedding spaces of independent models share a similar geometry. Prior works require a multimodal “Rosetta Stone” of paired cross-modal examples to align both spaces. In contrast, our Wasserstein Procrustes approach directly exploits this shared geometry for aligning without observing a single paired example. We here show the nearest neighbor of each image in the text embedding space of disjoint datasets.
Figure 2: Wasserstein Procrustes alignment. We first get a coarse correspondence by repeatedly clustering and matching the geometric structure of both embedding spaces (right). From this initialization, we alternate correspondence estimation and orthogonal Procrustes updates to refine the alignment (left). Paired examples, when available, enter both stages naturally as a linear term .
Figure 3: Shared-geometry for vision-language alignment without paired examples. ( 3(a) ) FOSCTTM of unpaired aligners on vision and language models on MS COCO, outperforming prior works on most model pairs. We can handle images and captions come from different datasets (ours†). ( 3(b) ) Our alignment transfers semantics enabling zero-shot classification. ( 3(c) ) Points are vision-language-dataset combinations. The more similar the embedding spaces, the better our unpaired alignment.
Figure 4: Shared geometry predicts unpaired alignment across scientific domains. ( 4 ) Our generic aligner is competitive with specialized methods on language (NQ, ten encoder pairs), biology (PBMC, RNA → ATAC) and neuroscience (NSD fMRI, fifty-six subject pairs). The non-aggregated metrics of each benchmark are reported in Sec. C.3 . ( 4 ) Each point corresponds to a modality pair spanning seven scientific domains. As in vision and language, modality pairs with more similar embedding geometry are consistently easier to align without paired data.
Figure 5: Shared geometry drastically reduces the amount of paired supervision beneficial for improved alignment. Alignment quality (FOSCTTM) and downstream zero-shot classification accuracy are shown as a function of the number of known image-text pairs. Starting from a strong unpaired solution (leftmost point), our geometry-based method consistently outperforms existing few-pair methods, with the largest gains below 100 pairs.
Figure 6: Unpaired alignment recovers semantic structure. We visualize with our aligner the shared space of DINOv2 ViT-B/14 and all-mpnet-base-v2 trained on MS COCO-train images and SPC captions. Retrieval images and captions are from MS COCO-val. Given a text, our aligner can retrieve matching images ( 6(a) ) and vice versa ( 6(b) ). We plot the first two principal components of the aligned space ( 6(c) ) along with the MS COCO supercategories: classes coarsely align. Text-to-image generations ( 6(d) ), using a diffusion model from embeddings to images, show that our unpaired aligner already generates images of the broad class and scene while additional pairs improve finer details.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Initialization
NQ
SNARE-seq
MS COCO
PCA heuristic
0.503±
0.055
0.321±
0.234
0.265±
0.116
sorted heuristic
0.499±
0.034
0.351±
0.134
0.472±
0.010
Gromov-Wasserstein (random)
0.455±
0.039
0.209±
0.129
0.432±
0.153
Gromov-Wasserstein (uniform)
0.527±
0.042
0.222±
0.128
0.316±
0.054
mini-vec2vec ( k -means + 2-opt, C=20 )
0.082±
0.071
0.194 ±
0.041
0.092±
0.016
mini-vec2vec ( k -means + 2-opt, C=30 )
0.070±
0.049
0.221±
0.007
0.092±
0.019
Appendix
Table 1: Initialization. FOSCTTM of every initialization with the same read-out and no refinement. The clustering and matching initialization outperforms all other initializations. For NQ and MS COCO, MPOpt outperforms the 2-opt solver from mini-vec2vec while being competitive on SNARE-seq.
Figure 7: QAP solvers on the class-matching benchmark of Schnaus et al. (2025) . ( 7(a) ) Gromov-Wasserstein cost of the returned permutation (solid) and lower bound of the solver (dashed), normalized by the squared problem size. MPOpt with the GRASP primal heuristic, the solver of our initialization, matches them there and is best beyond. Its bound lies below the axis. ( 7(b) ) Normalized cost, lower is better, and wall-clock time on one CPU, the best cost and the best time of every size in bold. The Hahn-Grant solver is stopped after 5400 seconds. The last column runs MPOpt with GRASP for as long as the Hahn-Grant solver. All solvers except MPOpt with GRASP are the published runs of Schnaus et al. (2025) .
Initialization
blobs k =10
blobs k =30
curve
SNARE-seq
splatter
uniform
random
0.0781 / 0.0590
0.0493 / 0.0412
0.0878 / 0.0319
0.0982 / 0.0795
0.0917 / 0.0765
0.0333 / 0.0278
trivial
0.0652 / 0.0336
0.0412 / 0.0238
0.0732 / 0.0323
0.0818 / 0.0814
0.0765 / 0.0765
0.0278 / 0.0264
k -means
0.0525 / 0.0373
0.0343 / 0.0241
0.0617 / 0.0343
0.0757 / 0.0593
0.0674 / 0.0568
0.0264 / 0.0227
lower bound
0.0430 / 0.0378
0.0245 / 0.0241
0.0277 / 0.0279
0.0773 / 0.0612
0.0753 / 0.0636
0.0198 / 0.0198
k -means + QAP
0.0211 / 0.0206
0.0183 / 0.0168
0.0129 / 0.0129
0.0647 / 0.0589
0.0446 / 0.0446
0.0251 / 0.0226
Appendix
Table 2: Cluster matching as a low-rank Gromov-Wasserstein solver. Gromov-Wasserstein cost of the low-rank plan before and after mirror descent ( before / after ), averaged over seeds and ranks. Lower is better. Matching cluster centers leads to a better initialization and a better optimum after mirror descent for structured data.
Read-out
NQ
SNARE-seq
MS COCO
direct
0.024±
0.007
0.246±
0.050
0.091±
0.020
top- k relative representations
0.000 ±
0.000
0.254±
0.101
0.085±
0.023
batched Hungarian matching
0.000 ±
0.000
0.216 ±
0.034
0.076 ±
0.023
Appendix
Table 3: Read-out. FOSCTTM after reading a map from the same averaged correspondence, without refinement. Our batched Hungarian matching slightly outperforms the other two methods while requiring fewer hyperparameters than top-k relative representations.
Refinement
NQ
SNARE-seq
MS COCO
no refinement
0.000 ±
0.000
0.216±
0.034
0.076±
0.023
mini-vec2vec
0.000 ±
0.000
0.209 ±
0.033
0.046±
0.039
ours
0.000 ±
0.000
0.209 ±
0.034
0.020 ±
0.015
Appendix
Table 4: Refinement. FOSCTTM after refining the same read-out. “mini-vec2vec” is its refine 1 and refine 2 ( Dar, 2025 ) , and “ours” is the batch alternation of Algorithm 1 . Our refinement leads to similar or better numbers on all three benchmarks. It also doesn’t introduce any additional hyperparameters.
Figure 8: All four hyperparameters improve with scale. We ablate one variable at a time while fixing all other variables. The sweeps over the batch size and the restarts use the read-out without refinement ( R=0 ), so they isolate the initialization. Increasing each hyperparameters improves performance, except for the batch size on SNARE-seq, which saturates because it only contains around 520 samples. In addition to an improved performance, more restarts of the initialization also reduce the variance.
Pair
Encoder 1
Encoder 2
Val.
MLIP ↔ MLIP (MP-20) ( Jain et al., 2013 )
mace_mp_small-pca256-structure ( Batatia et al., 2025 )
mace_mp_medium-pca256-structure ( Batatia et al., 2025 )
Table 5: The modality pairs of Sec. 4.3 . We test our aligner on a variety of independently trained models. The last column gives the number of paired samples the alignment is evaluated on.
Figure 9: Unpaired alignment on MS COCO. Top: FOSCTTM across all vision-language method combinations (white: chance level; blue: better alignment). Our method substantially outperforms vec2vec and mini-vec2vec across most vision-language combinations. Bottom: Zero-shot accuracy of our aligner (%, white: chance level; green: better alignment). Despite being trained without paired data, our method reaches zero-shot accuracies well above chance for most model pairs.
Figure 10: Unpaired alignment on SPC. Top: FOSCTTM across all vision-language method combinations (white: chance level; blue: better alignment). Our method substantially outperforms vec2vec and mini-vec2vec across most vision-language combinations. As observed across all detailed captioning datasets, none of the methods is able to reliably align Qwen3 with generative token pooling ( Wang et al., 2026 ) on SPC. Bottom: Zero-shot accuracy of our aligner (%, white: chance level; green: better alignment). Despite being trained without paired data, our method reaches zero-shot accuracies well above chance for most model pairs.
Figure 11: Unpaired alignment on DCI. Top: FOSCTTM across all vision-language method combinations (white: chance level; blue: better alignment). Our method substantially outperforms vec2vec and mini-vec2vec across most vision-language combinations. As observed across all detailed captioning datasets, none of the methods is able to reliably align Qwen3 with generative token pooling ( Wang et al., 2026 ) on DCI. Bottom: Zero-shot accuracy of our aligner (%, white: chance level; green: better alignment). Despite being trained without paired data, our method reaches zero-shot accuracies well above chance for most model pairs, although lower than on MS COCO.
Figure 12: Unpaired alignment on DOCCI. Top: FOSCTTM across all vision-language method combinations (white: chance level; blue: better alignment). Our method substantially outperforms vec2vec and mini-vec2vec across most vision-language combinations. As observed across all detailed captioning datasets, none of the methods is able to reliably align Qwen3 with generative token pooling ( Wang et al., 2026 ) on DOCCI. Bottom: Zero-shot accuracy of our aligner (%, white: chance level; green: better alignment). Despite being trained without paired data, our method reaches zero-shot accuracies well above chance for most model pairs, although lower than on MS COCO.
Figure 13: Zero-shot accuracy in the cross-dataset setting Zero-shot accuracy of our aligner (%, white: chance level; green: better alignment). We fit our aligner with images from MS COCO and captions from SPC. The FOSCTTM of this setting is the ours † column of Fig. 3(a) in the main text. Our aligner reaches zero-shot accuracies comparable to the single-dataset setting for MPNet and Qwen3-Embedding-8B, while generative token pooling drops to chance.
Figure 14: Shared geometry predicts alignment. FOSCTTM of our unpaired alignment against four measures of the geometric similarity of the paired embeddings, one point per vision-language-dataset combination, colored by dataset, with a least-squares fit and its cluster-bootstrap confidence band. Dashed: the 0.5 of a random correspondence. We observe that all four similarity measures predict the alignment even though CKA has the highest absolute correlation.
Text encoder
Method
Pairs
CLIPScore
VQAScore
TIFA
CycleReward
–
GT image
–
0.614±±0.005
0.875±±0.006
0.840±±0.009
1.632±±0.038
Qwen2.5-1.5B
Scale-RAE
–
0.660±±0.005
0.825±±0.008
0.821±±0.009
1.708±±0.034
MPNet
ours
0
0.466±±0.005
0.458±±0.012
0.552±±0.014
−0.431±±0.051
ours
1
0.477±±0.005
0.486±±0.012
0.583±±0.013
−0.429±±0.050
ours
10
0.491±±0.005
0.491±±0.013
0.583±±0.013
−0.356±±0.048
ours
100
0.510±±0.005
0.511±±0.012
0.615±±0.013
−0.120±±0.051
Appendix
Table 6: Faithfulness of the generated images on CyclePrefDB. Our aligner outperforms the linear map on all four measures, for every number of pairs and with both text encoders. The real images and Scale-RAE ( Tong et al., 2026 ) are included as reference points, but they are trained on tens of millions of image-text pairs and are not comparable to our aligner, which uses at most 100 pairs.
Text encoder
Method
Pairs
CLIPScore
VQAScore
TIFA
CycleReward
–
GT image
–
0.6389±±0.0005
0.8638±±0.0009
0.8695±±0.0008
0.2743±±0.0051
Qwen2.5-1.5B
Scale-RAE
–
0.6456±±0.0004
0.8264±±0.0010
0.8528±±0.0009
0.5283±±0.0050
MPNet
ours
0
0.5170±±0.0006
0.4284±±0.0014
0.5886±±0.0014
−1.3547±±0.0032
ours
1
0.5279±±0.0006
0.4644±±0.0014
0.6216±±0.0014
−1.2869±±0.0034
ours
10
0.5588±±0.0005
0.5197±±0.0014
0.6650±±0.0013
−1.1598±±0.0036
ours
100
0.5687±±0.0005
0.5412±±0.0013
0.6863±±0.0012
−1.1025±±0.0036
Appendix
Table 7: Faithfulness of the generated images on MS COCO. Our aligner outperforms the linear map on all four measures, for every number of pairs and with both text encoders. As in Tab. 6 , the real images and Scale-RAE are reference points and are not comparable to our aligner, which uses at most 100 pairs.
Figure 15: Unpaired text-to-image generation with our method and MPNet. We show the generated images for 8 sample captions with our aligner using DINOv2 ViT-B/14 and MPNet. Without pairs, our aligner can already produce images of the general semantic class and scene. Additional pairs improve fine-grained details.
Figure 16: Unpaired text-to-image generation with a linear map and MPNet. We show the generated images for 8 sample captions with a linear map fitted on the pairs using DINOv2 ViT-B/14 and MPNet. Without pairs, the linear map is random. As more pairs are used, the faithfulness of the generation improves.
Figure 17: Unpaired text-to-image generation with our method and Contriever. We show the generated images for 8 sample captions with our aligner using DINOv2 ViT-B/14 and Contriever. Without pairs, our aligner can already produce images of the general semantic class and scene. Additional pairs improve fine-grained details.
Figure 18: Unpaired text-to-image generation with a linear map and Contriever. We show the generated images for 8 sample captions with a linear map fitted on the pairs using DINOv2 ViT-B/14 and Contriever. Without pairs, the linear map is random. As more pairs are used, the faithfulness of the generation improves.
vec2vec
mini-vec2vec
ours
Pair
FOSCTTM ↓
mean rank ↓
FOSCTTM ↓
mean rank ↓
FOSCTTM ↓
mean rank ↓
granite → e5
0.5109
4185.7
0.0000
1.0
0.0000
1.1
granite → gte
0.4209
3448.6
0.0000
1.1
0.0000
1.2
granite → gtr
0.4070
3335.1
0.0000
1.2
0.0000
1.3
granite → stella
0.4875
3994.4
0.0000
1.1
0.0000
1.2
gte → e5
0.2861
2344.5
0.0000
1.1
0.0000
1.1
Appendix
Table 8: Per-pair results on NQ. We report the FOSCTTM and mean rank of every encoder pair evaluated on 8,192 validation samples. On average, mini-vec2vec and our aligner achieve nearly perfect alignment on all model pairs. Vec2vec is unstable and does not consistently demonstrate significant performance.
Method
FOSCTTM ↓
LTA ↑
SCOT+
0.121±
0.027
0.919±
0.015
ours
0.089 ±
0.003
0.942 ±
0.019
Appendix
Table 9: Result on PBMC. We evaluate RNA → ATAC with FOSCTTM and label transfer accuracy (LTA). Our aligner outperforms the domain-specific aligner SCOT+ ( Baker et al., 2026 ) in FOSCTTM and label transfer accuracy.
platonic brain
ours
Pair
FOSCTTM ↓
mean rank ↓
FOSCTTM ↓
mean rank ↓
sub01 → sub02
0.040
36.9
0.005
5.2
sub01 → sub03
0.038
35.8
0.042
38.7
sub01 → sub04
0.073
67.5
0.040
37.1
sub01 → sub05
0.047
43.3
0.046
42.9
sub01 → sub06
0.030
28.0
0.026
25.0
Appendix
Table 10: Per-pair results on NSD fMRI. We report the FOSCTTM score and the average rank of each pair of subjects, averaged over ten random seeds. On average, both methods perform similarly, even though our method is more stable and produces more consistent results.
FOSCTTM ↓ per method
shared geometry
Setting
Encoder (both sides)
identity
centred+norm
ours
CKA ↑
mutual k -NN ↑
MR ↔ CT (brain)
DINOv1 ViT-B/16
0.178
0.108
0.253
0.773
0.124
CBCT ↔ CT (brain)
DINOv2 ViT-B/14
0.156
0.048
0.054
0.710
0.226
RGB ↔ thermal
DINOv2 ViT-L/14
0.053
0.021
0.020
0.894
0.167
SAR ↔ optical
DINOv2 ViT-L/14
0.260
0.177
0.217
0.624
0.092
Appendix
Table 11: Alignment is not required for some domains. We consider multiple modality pairs from the CycleGAN ( Zhu et al., 2017 ) literature, where one encoder is applied to both modalities. We use DINOv1 ( Caron et al., 2021 ) , DINOv2 ( Oquab et al., 2024 ) , and WavLM ( Chen et al., 2022 ) as the encoders. The identity column applies no map, and the centered and normalized column subtracts the training mean from each side and normalizes the rows. Since they come from the same encoder, we observe that the embedding spaces are directly aligned, and centering and normalization help. This emphasizes that domains that can use the same encoders, such as RGB and thermal, are directly compatible in their embedding spaces when general self-supervised models are used. Although our aligner doesn’t assume a shared embedding space, it still recovers a coarse alignment for these modalities.
Domain
Pair
FOSCTTM ↓
CKA ↑
mutual k -NN ↑
TSI ↑
astronomy
galaxy ↔ spectrum (DESI)
0.57 ± 0.00
0.18 ± 0.00
0.01 ± 0.00
0.55 ± 0.00
biology
Cell Painting ↔ L1000 (Rosetta)
0.48±±0.10
0.25 ± 0.00
0.18 ± 0.00
0.58 ± 0.00
biology
human ↔ mouse scRNA
0.34±±0.04
0.32 ± 0.00
0.10 ± 0.00
0.60 ± 0.00
biology
scRNA ↔ ADT (CITE-seq)
0.22±±0.13
0.96 ± 0.00
0.14 ± 0.00
0.67 ± 0.00
biology
scRNA ↔ scATAC (PBMC)
0.09 ± 0.00
0.90 ± 0.00
0.06 ± 0.00
0.76 ± 0.00
biology
scRNA ↔ scATAC (SNARE-seq)
0.21±±0.03
0.86 ± 0.00
0.04 ± 0.00
0.74 ± 0.00
Appendix
Table 12: Complete results for all fourteen modality pairs of Fig. 4 . Modalities with a high similarity score can be aligned more easily without pairs.
Figure 19: What a good, a mediocre and a failed alignment look like. Three pairs of Tab. 12 , each fitted without any pairs. We use the median seed as a representative visualization. Every row shows both modalities after the map, projected by one PCA fitted on their union and colored by a label the aligner never saw. In addition, we show the label agreement of the nearest cross-modal neighbor (right), the FOSCTTM and three similarity scores (below). ( 19(a) ) The four cell lines of SNARE-seq separate, but BJ and K562 are matched to each other, so the matrix is diagonal up to that swap. ( 19(b) ) The lineages are recovered only in part. Neural cells are matched best, lymphoid and myeloid cells are mostly matched to each other, and muscle cells are drawn to the epithelial lineage. ( 19(c) ) Nothing is recovered.
Figure 20: Generative token pooling tends to help for short captions. We compare CKA alignment for different datasets. Generative token pooling tends to be more helpful for short captions than for long, detailed ones. WIT is an outlier where generative token pooling is exceptionally effective.
Figure 21: The same images, captions of decreasing detail. We compare different embedding strategies on DCI and DOCCI and find that generative token pooling can be helpful for short captions. Embedding models and normal mean pooling of the caption are generally superior for long captions.
Figure 22: Alignment between cluster centers, inside clusters and on random points. Both spaces are clustered jointly into C clusters. Every point of a curve scores n=C points, either the C cluster centers, C members of one cluster or C random samples, so the three curves differ only in how coarse the points are. The coarse cluster centers are consistently better aligned than random points or points within clusters.
Figure 23: Alignment in the top principal components. Each space is projected onto its top p principal components or onto a random p -dimensional subspace. The dotted line is the alignment of the two full spaces. Ten principal components capture most of the alignment and sometimes surpass the alignment of the original space. For mutual k -NN, around 100 principal components are required to approximate the full alignment.
Figure 24: Alignment after clipping one similarity kernel. The centered cosine-similarity kernel of one space is clipped from below or from above at the threshold τ , while the kernel of the other space is left unchanged. Points with a cosine similarity of 0.6 or higher and points farther away than orthogonal contribute only marginally to the alignment score.
Multimodal pre-training demonstrates strong generalization performance, but this paradigm is often impractical in domains where paired data are scarce. A promising alternative is post-hoc multimodal alignment, which aligns separately pre-trained unimodal encoders using a limited number of paired examples. However, existing methods focus primarily on aligning global representations, missing patch-token relations. This may hinder transfer to tasks that require fine-grained cross-modal matching beyond coarse sample-level semantics. To address this issue, we propose a post-hoc alignment method that learns token-level cross-modal structure using relative representations. Specifically, we represent images and texts through their token-level similarities to a set of learnable anchors in each modality space, which are trained to induce consistent cross-modal similarity patterns for matched pairs. Despite learning only the anchors without heavy projection layers, our approach consistently outperforms existing methods in zero-shot classification, cross-modal retrieval, and zero-shot segmentation by a substantial margin. This highlights the importance of modeling fine-grained cross-modal structure for effective post-hoc multimodal alignment with limited paired data.
The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world. Recent work exploits this convergence by aligning frozen pretrained vision and language models with lightweight alignment layers, but typically relies on contrastive losses and millions of paired samples. In this work, we ask whether meaningful alignment can be achieved with substantially less supervision. We introduce a semi-supervised setting in which pretrained unimodal encoders are aligned using a small number of image-text pairs together with large amounts of unpaired data. To address this challenge, we propose SOTAlign, a two-stage framework that first recovers a coarse shared geometry from limited paired data using a linear teacher, and then refines the alignment on unpaired samples via an optimal-transport-based divergence that transfers relational structure without overconstraining the target space. SOTAlign effectively leverages unpaired images and text, learning robust joint embeddings across datasets and encoder pairs, and significantly outperforming supervised and semi-supervised baselines. Code is available at https://github.com/ExplainableML/SOTAlign.
Simon Roschmann, Paul Krzakala, Sonia Mazelet +2
Helmholtz Munich · Technical University of Munich · Munich Center for Machine Learning +3
Despite the impressive results achieved by multimodal large language models (MLLMs), their training typically relies on jointly curated multimodal data, requiring substantial human effort to construct multi-way aligned datasets and thereby limiting scalability across domains. In this work, we explore training MLLMs by only leveraging multiple paired modalities as a surrogate for the full joint multimodal distribution. Specifically, we first provide a theoretical analysis of the conditions under which the representations are identifiable with only observing pairwise modalities. Building on this analysis, we propose a representation learning framework for aligning latent representations across modalities using only pairwise data. The framework consists of two stages: latent representation alignment and cross-modal recomposition. Specifically, in the first stage, we learn the shared latent space across modalities by both self-modal reconstruction and pair-wise contrastive learning. We also incorporate an inductive bias in the contrastive learning process by partially aligning and minimal latent specification. In stage two, we integrate the encoder of newly introduced modalities with the decoders of the pre-trained modalities to facilitate cross-modal transfer and generation. We evaluate our method by newly adding 3D point clouds and tactile modalities into pre-trained MLLMs with three modality pairs and show that, by learning an aligned latent representation space, our model achieves strong cross-modal performance.
Yan Li, Yunlong Deng, Yuewen Sun +3
Mohamed bin Zayed University of Artificial Intelligence · Carnegie Mellon University