Organizations: School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics · Artificial Intelligence and Digital Finance Key Laboratory of Sichuan Province · Chengdu Everimaging Science and Technology Co., Ltd · X-Humanoid · Hong Kong Institute of AI for Science, City University of Hong Kong · Zhejiang Sci-Tech University · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · University of Electronic Science and Technology of China
Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture. Experiments on the controlled synergy task and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions. Code is released at https://github.com/brightest66/HRIL.
Figures & tables
Figure 1: PID decomposes joint mutual information I(X1,X2;Y) into shared R , unique U1/U2 , and synergistic S .
Figure 2: The overall pipeline of HRIL. Given M modalities, we apply data augmentation to construct two views X′ and X′′ , and encode them into unimodal representations {Zm′}m=1M and {Zm′′}m=1M , as well as fused representations Zf′ and Zf′′ . The contrastive objective LInfoNCE performs cross-view alignment between each unimodal representation and the fused one, while Lsyn and Lalign are computed from unimodal representations in both views.
Model
Redundancy ↑
Uniqueness ↑
Synergy ↑
Cross ♣ [ 36 ]
100.0
11.6
50.0
Cross+Self ♣ [ 52 ]
99.7
86.9
50.0
FactorCL ♣ [ 26 ]
99.8
62.5
46.5
CoMM [ 11 ]
99.9±0.06
86.8±2.99
71.4±3.47
HRIL (ours)
99.7±0.14
92.6±3.04
82.3±1.61
Table 1: Linear probing accuracy (in %) for redundancy (shape), uniqueness (texture), and synergy (color and texture) on the Trifeature dataset. ♣ denotes results from [ 11 ] .
Model
UR-FUNNY ↑
MOSI ↑
MOSEI ↑
Average
Tri_CLIP
60.6±0.79
60.1±3.43
64.1±0.42
61.6
CMC [ 39 ]
60.1±0.78
62.4±1.67
62.7±0.72
61.73
ConFu [ 23 ]
61.3±0.96
62.5±2.66
64.8±0.61
62.87
CoMM [ 11 ]
64.8±1.13
66.0±2.27
69.9±0.34
66.9
HRIL (ours)
65.5±0.56
67.2±1.28
70.5±0.20
67.73
Table 2: Linear probing top-1 accuracy (in %) for classification tasks on trimodal Multibench.
Model
Classification
MIMIC ↑
MOSI ↑
UR-FUNNY ↑
MUSTARD ↑
Average ∗↑
Cross♣ [ 36 ]
66.7±0.1
47.8±1.8
50.1±1.9
53.5±2.9
54.52
Cross+Self♣ [ 52 ]
65.49±0.0
49.0±1.1
59.9±0.9
53.9±4.0
57.07
FactorCL♣ [ 26 ]
67.3±0.0
51.2±1.6
60.5±0.8
55.80±0.9
58.7
CoMM [ 11 ]
66.4±0.4
63.7±2.5
63.3±0.5
64.4±1.1
64.45
HRIL (ours)
68.0±0.7
65.8±1.2
63.9±0.9
66.6±2.7
66.08
Table 3: Linear probing top-1 accuracy (in %) for classification tasks on bimodal Multibench. ♣ denotes results are from [ 26 ] .
Loss components
Redundancy ↑
Uniqueness ↑
Synergy probe ↑
LInfoNCE
Lalign
Lsyn
×
✓
✓
28.5±2.74
12.5±3.15
50.0±0.00
✓
×
×
99.9±0.10
87.2±2.13
50.0±0.00
✓
×
✓
99.9±0.05
86.6±2.79
72.4±3.16
✓
✓
×
99.9±0.09
85.1±6.86
78.0±0.44
✓
✓
✓
99.7±0.14
92.6±3.04
82.3±1.61
Table 4: Linear probing accuracy (in %) of multimodal interactions on the Trifeature dataset under different combinations of loss terms in Eq. 8 .
Lsyn
InterSHAP Value (%)
Average
UR-FUNNY
MOSI
MOSEI
×
7.78±4.56
6.82±3.73
1.35±0.53
5.32
✓
9.25±5.56
9.90±4.81
1.47±0.40
6.87
Improvement
+1.47
+3.08
+0.12
+1.55
Table 5: Quantitative evaluation of InterSHAP on Multibench [ 27 ] datasets. × denotes the model without Lsyn , and ✓ denotes the full objective of HRIL.
Fusion
R
U
S
Average
Concat+Linear
99.9±0.1
65.4±10.73
50.0±0.00
71.77
HRIL
99.7±0.14
92.6±3.04
82.3±1.61
91.83
Table 6: Linear probing accuracy (in %) of multimodal interactions on the Trifeature dataset. Concat+Linear: concatenate the modality embeddings and pass through a linear layer.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
UR-FUNNY
MOSI
CoMM
7.77±3.05
3.82±1.90
HRIL without Lsyn
7.78±4.56
6.82±3.73
Multilinear-rank control
8.17±6.33
5.87±4.26
HRIL
9.25±5.56
9.90±4.81
Appendix
Table 7: Further comparison with InterSHAP on the UR-FUNNY and MOSI datasets.
Hyperparameter
Value
UR-FUNNY
MOSEI
Tucker rank rm
8
65.0±0.54
69.9±0.31
16
64.9±1.34
70.0±0.47
24
65.5±0.56
70.5±0.20
Projection dimension rp
8
64.8±0.59
70.2±0.26
16
64.9±0.66
70.2±0.20
24
65.5±0.56
70.5±0.20
Appendix
Table 8: Sensitivity analysis: accuracy (%)
Figure 3: Performance sensitivity to mini-batch size on MOSEI, UR-FUNNY, and MOSI.
Model
Modalities
weighted-F1 ↑
macro-F1 ↑
SimCLR ♣△ [ 7 ]
V
40.35±0.23
27.99±0.33
CLIP ♣ [ 36 ]
V
51.5
40.8
L
51.0
43.0
V+L
58.9
50.9
SLIP ♣△ [ 32 ]
V+L
56.54±0.19
47.35±0.27
CLIP ♣△ [ 36 ]
V+L
54.49±0.19
44.94±0.30
Appendix
Table 9: Linear probing F1 scores (%) on MM-IMDb. △ indicates further training on unlabeled data. ♣ denotes results from [ 11 ] .
Model
V&T CP ↑
Cross
86.3±0.25
Cross+Self
87.6±0.26
CoMM
87.0±1.77
HRIL (ours)
88.6±0.14
Appendix
Table 10: Linear probing accuracy (%) on Vision&Touch contact prediction (V&T CP).
Dataset
Model
Training Efficiency Metrics
Step Time (ms)
Optimizer Time (ms)
Backward / Step Ratio (%)
UR-FUNNY
CoMM
90.73
3.86
53.82
HRIL
155.65
3.47
46.23
MOSI
CoMM
90.00
3.90
54.07
HRIL
156.95
3.47
46.96
MOSEI
CoMM
90.57
4.06
53.99
Appendix
Table 11: Training efficiency comparison between CoMM [ 11 ] and HRIL on trimodal Multibench [ 27 ] datasets. Both methods use the same multimodal backbone. We report the average wall-clock step time, optimizer update time, and the percentage of each step spent in backward computation.
A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone. While most approaches operate at the architectural level through larger or more complex fusion models, we propose a complementary axis: shaping the training objective itself. Standard training often emphasizes unimodal or redundant information, falling short on examples that require cross-modal reasoning. We formalize multimodal synergy through information theory and introduce the Synergistic Information Bottleneck (SynIB), a scalable objective that targets synergy directly. To prioritize learning synergy, SynIB motivates the model to predict accurately from all modalities while penalizing confidence when information from any modality is withheld. Alongside the standard task loss, the model runs forward passes with one modality masked at a time and is penalized for remaining confident, which would indicate reliance on unimodal cues rather than cross-modal interactions. We validate SynIB in two regimes. On synthetic XOR tasks where the ground-truth synergy is known by construction, standard training fails to recover it while SynIB does. On five real-world benchmarks, including three MultiBench affective tasks, Hateful Memes with CLIP-ViT and DeBERTa backbones, and a controllable irony extension of CREMA-D we introduce, SynIB improves accuracy on synergy-dependent examples by up to 7.8% and overall accuracy by up to 3.8%.
Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen +4
Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way to isolate these target-relative contributions within learned continuous representations. We introduce a framework that applies contrastive or masked objectives at intermediate layers, coupled with source-wise invertible normalizing flows and a supervised, low-rank latent variable model. This architecture explicitly factorizes the joint distribution into shared task-relevant variation, modality-specific predictive variation, and task-irrelevant dependence. Drawing connections to prior multimodal learning assumptions, our approach evaluates how modalities independently and jointly contribute to the target. Ultimately, this framework unites intermediate representation learning with structured likelihood-based guidance, offering a practical latent-variable lens for characterizing continuous multimodal interactions. Empirically, we demonstrate the effectiveness of our approach across diverse multimodal benchmarks, showing robust improvements in predictive performance.
Wanting Huang, Sanvesh Srivastava, Weiran Wang
Department of Computer Science University of Iowa · Department of Statistics University of Iowa
Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions. A critical yet underexplored challenge is that these implicit interactions vary dynamically across samples. In this work, we present the first systematic, information-theoretic analysis highlighting why learning these dynamic, sample-specific interactions is critical for effective multimodal learning. Our analysis further reveals deficits in conventional paradigms at learning these distinct interaction types: modality ensemble approaches struggle to capture synergy, while joint learning paradigms often under-utilize redundant information. This highlights the need for an approach that can adaptively learn from different interaction types on a per-sample basis. To this end, we propose Decomposition-based Multimodal Interaction Learning (DMIL), a novel paradigm that explicitly models and learns from sample-specific interactions. First, we design a variational decomposition architecture to isolate the constituent interaction components. Second, we employ a new learning strategy that leverages these explicit interaction components in a fine-tuning process to achieve comprehensive interaction learning. Extensive experiments across diverse tasks and architectures demonstrate that DMIL consistently achieves superior performance by adapting to holistic sample-specific interactions. Our framework is flexible and broadly applicable, establishing an interaction-centric paradigm for multimodal learning. The code is available at https://github.com/GeWu-Lab/DMIL.
Zequn Yang, Yake Wei, Haotian Ni +2
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing · Beijing Key Laboratory of Research on Large Models and Intelligent Governance · Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE +2