cs.LGAug 16, 2026

Fusion Anything: A Generalized Multimodal Foundation Model

Authors: Huizi Cui, Zongbo Han, Chenggong Ding, Naichuan Xiao, Jialong Yang, Jingdong Chen, Guangyu Wang, Qinghua Hu, +1 more

Organizations: Tianjin University · Beijing University of Posts and Telecommunications

Abstract

Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single task, making it difficult to quickly adapt to new downstream applications. Therefore, a natural yet aggressive question arises - whether there exists a general multimodal fusion model that can be applied to arbitrary modality combinations and arbitrary prediction tasks. We argue that a unified multimodal fusion model should not depend on specific modalities and should instead encode transferable patterns of multimodal correlation. To this end, we propose a simple and effective learning paradigm based on training on large-scale synthetic multimodal datasets generated with Structural Multimodal Causal Models (SMCMs), which formally characterizes the generative processes of real-world multimodal data. Building on this framework, we propose the Fusion Anything Model (FAM), a foundation model for generalized multimodal data fusion. By constructing large-scale synthetic multimodal data with diverse correlation patterns, our model encodes transferable multimodal correlations during training and activates appropriate associations through in-context examples during inference. Extensive experiments on 18 real-world datasets spanning 12 modalities and 11 prediction tasks demonstrate that our model achieves competitive performance with specialized models without task-specific adaptation.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

    Jul 28, 2026Mingqiao Ye, Zhaochong An, Zhitong Gao +11Modality-Specific EncodersComplementary Modalities

  2. Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis

    Jul 14, 2026Alexios Filippakopoulos, Elias Kallioras, Nikolaos Xiros +2Multimodal FusionCross-Modal

  3. SMA: Submodular Modality Aligner For Data Efficient Multimodal Learning

    May 13, 2026Truong Pham, Anay Majee, Rishabh IyerMultimodal AlignmentModality Alignment