In self-supervised learning, a third modality should improve, or at least preserve, performance. Across nine image-text-tabular datasets, we show that it instead harms performance: the trimodal model underperforms its own best bimodal subset in 55.6% of paired runs under VICReg. The same failure occurs in 51.1% of paired runs under SimSiam. We call this failure multimodal saturation. We propose that the failure lies in the alignment geometry. Riemannian VICReg (R-VICReg) generalizes classical VICReg: it aligns views by squared geodesic distance on learnable negative-curvature product factors and recovers VICReg exactly as curvature vanishes. Over the same 45 paired runs, R-VICReg raises the probability that the third modality helps from 44.4% to 64.4%, with gains concentrated where VICReg saturates.
Figures & tables
Figure 1: The alignment geometry decides whether aligned states stay readable. Masked-view alignment contracts each semantic state (one marker shape per state) inside the same radius budget, and a downstream head resolves states only at its margin scale ε=L1 . (a) Under Euclidean alignment, several states land in one ε -cell and merge: induced redundancy. (b) Under geodesic alignment at κ<0 , the same budget offers exponentially more ε -cells, so each state keeps its own. Lemmas 8 – 10 prove the two growth laws.
Figure 2: Three views render the learned metric of one trimodal CSGO run (seed 456). (a) In the flat VICReg projector, equal-distance rings stay evenly spaced. (b) In the most-utilized R-VICReg factor ( κ32=−0.988 ), the rings at the same geodesic distances crowd toward the boundary. (c) The same points sit on the hyperboloid above their Poincaré-ball shadow (construction in Appendix D ).
Figure 3: Panel (a) reports the per-dataset positive-gain rate p+(0)=Pr(Δtri>0) over five seeds, so rates are multiples of 0.2 . VICReg scores 0 of 5 on CSGO and Pokemon Secondary (marked 0 ). Panel (b) reports the paired difference in trimodal test accuracy between R-VICReg and VICReg: bars give the five-seed mean and the orange line the VICReg baseline.
Figure 4: Reliability p+(τ)=Pr(Δtri>τ) over N=45 paired runs per method, swept over τ∈{0,0.005,0.01} : R-VICReg stays above both baselines at every threshold, so the ranking is no artifact of τ=0 .
Table 1: κ -stereographic operations used by R-VICReg and their flat limits. Every limit is uniform on compacta (Lemma 12 ); the distance limit squares to the invariance-term limit dκ(x,y)2→4∥x−y∥22 of Proposition 4 .
Figure 5: A rescue is a paired run that VICReg saturates ( Δtri≤τ ) and R-VICReg does not; a harm is the reverse. (a) The quadrants at τ=0 hold 13 rescues, 4 harms, and 28 concordant runs. (b) The paired improvement δ is positive where VICReg saturates and negative where it does not; bars show 95% Student- t intervals.
Figure 6: The learned R-VICReg geometry stays hyperbolic on every dataset: curvature remains negative and boundary utilization stays high. The left axis reports umax , with bars spanning the five-seed minimum–maximum range; the right axis reports the mean learned curvature κ .
Hyperparameter
Value
Pretrain / finetune epochs
50 / 50
Batch size
64
Optimizer / learning rate / weight decay
AdamW / 10−4 / 0.02
Warmup epochs / gradient clip
5 / 1.0
Text / image / tabular reconstruction weights
(1.0,1.0,0.01)
Alignment weight λMMM
1.0
Appendix
Table 2: Hyperparameters shared by all experiments.
Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively underexplored. Pairwise CLIP-style losses decompose multi-modal alignment into independent two-way comparisons and therefore do not explicitly model higher-order dependencies among multiple modalities. Recent beyond-pairwise objectives approach this problem from statistical or geometric perspectives, but arbitrary-modality alignment still lacks a principled criterion for defining what each modality should preserve and compress relative to the others. We revisit arbitrary-modality alignment through the Information Bottleneck principle. In multi-modal learning, sufficiency should preserve information predictable from the remaining modalities, while minimality should compress modality-specific information not supported by them. This naturally leads to a One-vs-All view, where each modality is characterized with respect to the remaining modalities. We propose OVA-IB, an Information Bottleneck framework for arbitrary-modality alignment. OVA-IB optimizes a tractable One-vs-All contrastive lower bound for sufficiency connected to a Dual Total Correlation-style objective, uses a parameter-free geometry-aware projection score, and derives a tractable upper-bound regularizer for minimality by bounding each representation's dependence on its own input with representation distributions induced by the remaining modalities. Experiments on classification, regression, modality-agnostic evaluation, and cross-modal retrieval benchmarks demonstrate strong and robust performance.
Tianchao Li, Shujian Yu, Xinrui Zu +4
Hong Kong University of Science and Technology, Hong Kong · Vrije Universiteit Amsterdam, Netherlands · UiT – The Arctic University of Norway, Norway +3
Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of n-modal data remains challenging. In this work, we present BaryBind, which aims to transport the specific modality towards the Wasserstein barycenter (WB) optimized across all modalities and introduces a volumetric alignment objective to establish a unified semantic space around the WB embedding. Specifically, we project specific modalities to the WB, which minimizes the average Wasserstein distances to multimodal distributions and serves as the anchor for subsequent alignment. We then construct a barycenter simplex, whose volume is taken as a similarity metric for global alignment centered at the WB. Experiments show that BaryBind achieves competitive performance in text-video-audio retrieval, classification, videoQA, and cross-modal generation tasks, along with robustness under modality absence and scalability to more than three modalities. Code is released at https://github.com/xl-tang3/BaryBind.
Contrastive learning is increasingly moving toward settings with three or more modalities instead of image-text pairs. Yet, extending models from pairwise to higher-order multimodal alignment can introduce optimization and representation challenges. We identify encoder Jacobian conditioning as a key factor in trimodal contrastive learning: poorly conditioned encoders exhibit collapsing or amplified singular-value spectra, leading to exploding Jacobian condition numbers and degraded multimodal alignment. We introduce geometry-preserving encoders (GPEs) by directly conditioning the Jacobian through regularization and demonstrating that simple modifications like LeakyReLU activations and residual paths recover these geometric benefits. Across a synthetic benchmark and four real-world datasets including missing modalities, improving Jacobian conditioning boosts retrieval and linear probe performance across multiple contrastive objectives, whereas expressive objectives yield little benefit in linear probes. More broadly, our results show that multimodal contrastive learning depends not only on objective expressivity, but also on the geometric and optimization properties of the underlying encoders.
Tillmann Rheude, Roland Eils, Benjamin Wild
Berlin Institute of Health, Charité - Universitätsmedizin Berlin · Department of Mathematics and Computer Science, Freie Universität Berlin · Intelligent Medicine Institute, Fudan University