Organizations: School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics · Artificial Intelligence and Digital Finance Key Laboratory of Sichuan Province · Chengdu Everimaging Science and Technology Co., Ltd · X-Humanoid · Hong Kong Institute of AI for Science, City University of Hong Kong · Zhejiang Sci-Tech University · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · University of Electronic Science and Technology of China
Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture. Experiments on the controlled synergy task and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions. Code is released at https://github.com/brightest66/HRIL.
Figures & tables
Figure 1: PID decomposes joint mutual information I(X1,X2;Y) into shared R , unique U1/U2 , and synergistic S .
Figure 2: The overall pipeline of HRIL. Given M modalities, we apply data augmentation to construct two views X′ and X′′ , and encode them into unimodal representations {Zm′}m=1M and {Zm′′}m=1M , as well as fused representations Zf′ and Zf′′ . The contrastive objective LInfoNCE performs cross-view alignment between each unimodal representation and the fused one, while Lsyn and Lalign are computed from unimodal representations in both views.
Model
Redundancy ↑
Uniqueness ↑
Synergy ↑
Cross ♣ [ 36 ]
100.0
11.6
50.0
Cross+Self ♣ [ 52 ]
99.7
86.9
50.0
FactorCL ♣ [ 26 ]
99.8
62.5
46.5
CoMM [ 11 ]
99.9±0.06
86.8±2.99
71.4±3.47
HRIL (ours)
99.7±0.14
92.6±3.04
82.3±1.61
Table 1: Linear probing accuracy (in %) for redundancy (shape), uniqueness (texture), and synergy (color and texture) on the Trifeature dataset. ♣ denotes results from [ 11 ] .
Model
UR-FUNNY ↑
MOSI ↑
MOSEI ↑
Average
Tri_CLIP
60.6±0.79
60.1±3.43
64.1±0.42
61.6
CMC [ 39 ]
60.1±0.78
62.4±1.67
62.7±0.72
61.73
ConFu [ 23 ]
61.3±0.96
62.5±2.66
64.8±0.61
62.87
CoMM [ 11 ]
64.8±1.13
66.0±2.27
69.9±0.34
66.9
HRIL (ours)
65.5±0.56
67.2±1.28
70.5±0.20
67.73
Table 2: Linear probing top-1 accuracy (in %) for classification tasks on trimodal Multibench.
Model
Classification
MIMIC ↑
MOSI ↑
UR-FUNNY ↑
MUSTARD ↑
Average ∗↑
Cross♣ [ 36 ]
66.7±0.1
47.8±1.8
50.1±1.9
53.5±2.9
54.52
Cross+Self♣ [ 52 ]
65.49±0.0
49.0±1.1
59.9±0.9
53.9±4.0
57.07
FactorCL♣ [ 26 ]
67.3±0.0
51.2±1.6
60.5±0.8
55.80±0.9
58.7
CoMM [ 11 ]
66.4±0.4
63.7±2.5
63.3±0.5
64.4±1.1
64.45
HRIL (ours)
68.0±0.7
65.8±1.2
63.9±0.9
66.6±2.7
66.08
Table 3: Linear probing top-1 accuracy (in %) for classification tasks on bimodal Multibench. ♣ denotes results are from [ 26 ] .
Loss components
Redundancy ↑
Uniqueness ↑
Synergy probe ↑
LInfoNCE
Lalign
Lsyn
×
✓
✓
28.5±2.74
12.5±3.15
50.0±0.00
✓
×
×
99.9±0.10
87.2±2.13
50.0±0.00
✓
×
✓
99.9±0.05
86.6±2.79
72.4±3.16
✓
✓
×
99.9±0.09
85.1±6.86
78.0±0.44
✓
✓
✓
99.7±0.14
92.6±3.04
82.3±1.61
Table 4: Linear probing accuracy (in %) of multimodal interactions on the Trifeature dataset under different combinations of loss terms in Eq. 8 .
Lsyn
InterSHAP Value (%)
Average
UR-FUNNY
MOSI
MOSEI
×
7.78±4.56
6.82±3.73
1.35±0.53
5.32
✓
9.25±5.56
9.90±4.81
1.47±0.40
6.87
Improvement
+1.47
+3.08
+0.12
+1.55
Table 5: Quantitative evaluation of InterSHAP on Multibench [ 27 ] datasets. × denotes the model without Lsyn , and ✓ denotes the full objective of HRIL.
Fusion
R
U
S
Average
Concat+Linear
99.9±0.1
65.4±10.73
50.0±0.00
71.77
HRIL
99.7±0.14
92.6±3.04
82.3±1.61
91.83
Table 6: Linear probing accuracy (in %) of multimodal interactions on the Trifeature dataset. Concat+Linear: concatenate the modality embeddings and pass through a linear layer.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
UR-FUNNY
MOSI
CoMM
7.77±3.05
3.82±1.90
HRIL without Lsyn
7.78±4.56
6.82±3.73
Multilinear-rank control
8.17±6.33
5.87±4.26
HRIL
9.25±5.56
9.90±4.81
Appendix
Table 7: Further comparison with InterSHAP on the UR-FUNNY and MOSI datasets.
Hyperparameter
Value
UR-FUNNY
MOSEI
Tucker rank rm
8
65.0±0.54
69.9±0.31
16
64.9±1.34
70.0±0.47
24
65.5±0.56
70.5±0.20
Projection dimension rp
8
64.8±0.59
70.2±0.26
16
64.9±0.66
70.2±0.20
24
65.5±0.56
70.5±0.20
Appendix
Table 8: Sensitivity analysis: accuracy (%)
Figure 3: Performance sensitivity to mini-batch size on MOSEI, UR-FUNNY, and MOSI.
Model
Modalities
weighted-F1 ↑
macro-F1 ↑
SimCLR ♣△ [ 7 ]
V
40.35±0.23
27.99±0.33
CLIP ♣ [ 36 ]
V
51.5
40.8
L
51.0
43.0
V+L
58.9
50.9
SLIP ♣△ [ 32 ]
V+L
56.54±0.19
47.35±0.27
CLIP ♣△ [ 36 ]
V+L
54.49±0.19
44.94±0.30
Appendix
Table 9: Linear probing F1 scores (%) on MM-IMDb. △ indicates further training on unlabeled data. ♣ denotes results from [ 11 ] .
Model
V&T CP ↑
Cross
86.3±0.25
Cross+Self
87.6±0.26
CoMM
87.0±1.77
HRIL (ours)
88.6±0.14
Appendix
Table 10: Linear probing accuracy (%) on Vision&Touch contact prediction (V&T CP).
Dataset
Model
Training Efficiency Metrics
Step Time (ms)
Optimizer Time (ms)
Backward / Step Ratio (%)
UR-FUNNY
CoMM
90.73
3.86
53.82
HRIL
155.65
3.47
46.23
MOSI
CoMM
90.00
3.90
54.07
HRIL
156.95
3.47
46.96
MOSEI
CoMM
90.57
4.06
53.99
Appendix
Table 11: Training efficiency comparison between CoMM [ 11 ] and HRIL on trimodal Multibench [ 27 ] datasets. Both methods use the same multimodal backbone. We report the average wall-clock step time, optimizer update time, and the percentage of each step spent in backward computation.
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing · Beijing Key Laboratory of Research on Large Models and Intelligent Governance · Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE +2