cs.LGSep 24, 2026

ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models

Authors: Xunkai Li, Xu Wang, Yinlin Zhu, Xiong Yongfu, Yi Liu, Rong-Hua Li, Guoren Wang

Organizations: Department of Computer Science, Beijing Institute of Technology, Beijing, China · School of Airspace Science and Engineering, Shandong University, WeiHai, China · School of Computer Science and Engineering, Sun Yat-sen University, GuangZhou, China

Abstract

Multimodal attributed graphs connect entities, visual content, language, and observed relations. Learning one foundation across such graphs requires more than compressing each node into a fused Euclidean vector. The representation must preserve entity semantics, construct interaction state from graph neighborhoods, and expose that state to prediction units with different geometry. Our empirical study shows why these requirements are inseparable. Higher-grade channels recover pair relations across the foundation graphs, specialized queries reveal information hidden by a generic readout, and rigid blade isolation removes cross-grade capacity. We therefore introduce ICE (Interaction-aware Clifford Encoder), a multimodal graph foundation model built on a node-indexed Clifford latent field. Topology, text, and images enter explicit Cl(3) addresses. Edge-aware geometric products transform these directions into scalar, bivector, and trivector relations over observed neighborhoods. A protected Grade-1 route preserves entity semantics, while the full grade and depth bank remains available to fresh node and link heads. We establish exact cross-grade reachability, node-permutation equivariance, and a bound on the task residual around the semantic score. Experiments span one shared foundation over eleven graphs, six node-classification datasets, three link-prediction datasets, and matched few-shot tasks. ICE ranks first in all 30 reported supervised and few-shot comparisons. Core removals reduce every task summary, and mechanism controls connect the gains to higher-order transport, retained multidepth structure, semantic protection, and direct field access.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FedLAB: Traceable Semantic Codebooks for Federated Multimodal Graph Foundation Learning

    Jun 30, 2026Zekai Chen, Kairui Yang, Xuaner Chen +4Federated Graph LearningGraph Foundation Model

  2. CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer

    Jul 28, 2026Ankang Yang, Jitao Zhao, Di Jin +2Graph Foundation ModelMultimodal Foundation Model

  3. Toward Federated Multimodal Graph Foundation Models: A Topology-Aware Multimodal Alignment Framework

    Jul 17, 2026Xunkai Li, Guohao Fu, Yuming Ai +4Federated Graph LearningGraph Foundation Model