cs.AIOct 8, 2026

HRIL: Learning Multimodal Synergy via Higher-Order Tensor Modeling

Authors: Qun Dai, Liangjian Wen, Jiang Duan, Yong Dai, Dongkai Wang, Maolin Wang, Mingjie Wang, Jianzhuang Liu, +2 more

Organizations: School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics · Artificial Intelligence and Digital Finance Key Laboratory of Sichuan Province · Chengdu Everimaging Science and Technology Co., Ltd · X-Humanoid · Hong Kong Institute of AI for Science, City University of Hong Kong · Zhejiang Sci-Tech University · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · University of Electronic Science and Technology of China

Abstract

Self-supervised multimodal representation learning has achieved remarkable success across diverse domains, yet capturing synergistic information remains challenging due to the complexity of cross-modal interactions. Unlike the shared information across individual modalities, synergy arises when task-relevant signals emerge only from the joint configuration of multiple modalities and cannot be recovered from any modality in isolation. This work focuses on how to preserve the information capacity for such synergistic signals in multimodal representations. The key observation is that synergistic information is reflected in higher-order statistical dependence among modalities, which provides a principled target for explicitly modeling joint interactions. Motivated by this insight, we propose Higher-order Representation and Information Learning (HRIL), which constructs an empirical cross-moment tensor over modality embeddings to represent multi-way interactions. HRIL employs Tucker decomposition to obtain a core tensor, complemented by a synergy-aware regularizer that prevents energy concentration and preserves higher-order coupling capacity for synergistic information capture. Experiments on the controlled synergy task and real-world benchmarks demonstrate consistent improvements over existing multimodal contrastive methods, with notable gains on tasks dominated by synergistic interactions. Code is released at https://github.com/brightest66/HRIL.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning

    May 12, 2026Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen +4Cross-Modal LearningMultimodal Learning

  2. Structured Latent Modeling for Supervised Multimodal Information Decomposition

    Sep 28, 2026Wanting Huang, Sanvesh Srivastava, Weiran WangMultimodal Disentangled Representation LearningMultimodal Learning

  3. Information-Theoretic Decomposition for Multimodal Interaction Learning

    Jun 10, 2026Zequn Yang, Yake Wei, Haotian Ni +2Multimodal FusionMultimodal Learning