cs.LGOct 5, 2026

On the Geometry of Multimodal Saturation: Riemannian VICReg

Authors: Nessim Ben Abbes, Duc Han Le, Sabri Mtibaa, Van-Tam Nguyen

Organizations: Protectline · Télécom Paris

Abstract

In self-supervised learning, a third modality should improve, or at least preserve, performance. Across nine image-text-tabular datasets, we show that it instead harms performance: the trimodal model underperforms its own best bimodal subset in 55.6% of paired runs under VICReg. The same failure occurs in 51.1% of paired runs under SimSiam. We call this failure multimodal saturation. We propose that the failure lies in the alignment geometry. Riemannian VICReg (R-VICReg) generalizes classical VICReg: it aligns views by squared geodesic distance on learnable negative-curvature product factors and recovers VICReg exactly as curvature vanishes. Over the same 45 paired runs, R-VICReg raises the probability that the third modality helps from 44.4% to 64.4%, with gains concentrated where VICReg saturates.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment

    May 28, 2026Tianchao Li, Shujian Yu, Xinrui Zu +4Multimodal AlignmentModality Alignment

  2. Binding Multiple Modalities via Multimodal Wasserstein Barycenter

    Sep 27, 2026Xiaole Tang, Jiayi Xu, Xiang Gu +2Multimodal AlignmentComplementary Modalities

  3. Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning

    Jul 20, 2026Tillmann Rheude, Roland Eils, Benjamin WildMultimodal Contrastive LearningMultimodal Alignment