ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models
Authors: Xunkai Li, Xu Wang, Yinlin Zhu, Xiong Yongfu, Yi Liu, Rong-Hua Li, Guoren Wang
Organizations: Department of Computer Science, Beijing Institute of Technology, Beijing, China · School of Airspace Science and Engineering, Shandong University, WeiHai, China · School of Computer Science and Engineering, Sun Yat-sen University, GuangZhou, China
Multimodal attributed graphs connect entities, visual content, language, and observed relations. Learning one foundation across such graphs requires more than compressing each node into a fused Euclidean vector. The representation must preserve entity semantics, construct interaction state from graph neighborhoods, and expose that state to prediction units with different geometry. Our empirical study shows why these requirements are inseparable. Higher-grade channels recover pair relations across the foundation graphs, specialized queries reveal information hidden by a generic readout, and rigid blade isolation removes cross-grade capacity. We therefore introduce ICE (Interaction-aware Clifford Encoder), a multimodal graph foundation model built on a node-indexed Clifford latent field. Topology, text, and images enter explicit Cl(3) addresses. Edge-aware geometric products transform these directions into scalar, bivector, and trivector relations over observed neighborhoods. A protected Grade-1 route preserves entity semantics, while the full grade and depth bank remains available to fresh node and link heads. We establish exact cross-grade reachability, node-permutation equivariance, and a bound on the task residual around the semantic score. Experiments span one shared foundation over eleven graphs, six node-classification datasets, three link-prediction datasets, and matched few-shot tasks. ICE ranks first in all 30 reported supervised and few-shot comparisons. Core removals reduce every task summary, and mechanism controls connect the gains to higher-order transport, retained multidepth structure, semantic protection, and direct field access.
Figures & tables
Figure 1: Empirical study. (a) Pair-relation AUC after grade expansion and a second transport layer. (b) Accuracy and Macro-F1 gains (percentage points) from direct grade–depth access over a generic readout. (c) Rigid blade isolation reports the retained active share (“Act.”), deepest e12 energy (“ e12 ”), and matched validation wins (“Win.”).
Figure 2: ICE framework. One dataset-agnostic Clifford field connects multimodal input, edge-aware graph transport, foundation learning, and fresh node and link heads.
Method
RedditS
Movies
Grocery
Toys
Ele-F.
Books-NC
Acc.
F1
Acc.
F1
Acc.
F1
Acc.
F1
Acc.
F1
Acc.
F1
Traditional multimodal
MMGCN
90.27 ± 0.34
84.22 ± 0.73
53.41 ± 1.03
41.66 ± 2.03
82.56 ± 0.49
73.83 ± 0.93
80.02 ± 0.64
76.36 ± 1.23
86.59 ± 0.08
68.85 ± 0.35
83.22 ± 0.10
71.28 ± 0.22
MGAT
92.78 ± 0.50
87.27 ± 0.53
53.87 ± 0.50
44.09 ± 1.60
83.74 ± 0.62
74.77 ± 1.11
79.61 ± 0.74
77.09 ± 0.87
84.84 ± 0.08
69.62 ± 0.21
82.91 ± 0.04
71.45 ± 0.11
Self-supervised graph models
GRACE
93.01 ± 0.53
88.39 ± 1.12
48.09 ± 0.97
37.18 ± 1.33
70.83 ± 0.81
60.69 ± 1.05
72.82 ± 0.66
69.09 ± 0.63
83.58 ± 0.11
70.09 ± 0.47
74.96 ± 0.06
70.09 ± 0.11
Table 1: Node classification. Accuracy and Macro-F1 for each dataset over eight seeds. Best means are bold and second-best means are underlined.
Method
Link prediction
Few-shot link classification
Sports
Cloth
Books-LP
Sports 3
Sports 5
Sports 10
Cloth 3
Cloth 5
Cloth 10
Traditional multimodal
MMGCN
23.44 ± 0.43
17.74 ± 0.38
20.73 ± 0.48
55.86 ± 2.96
53.70 ± 1.20
56.66 ± 1.60
64.36 ± 4.07
68.61 ± 5.38
67.27 ± 4.20
MGAT
21.74 ± 0.96
15.47 ± 0.32
21.82 ± 0.53
54.92 ± 2.40
56.55 ± 2.69
57.92 ± 1.80
67.05 ± 1.45
70.34 ± 3.36
68.66 ± 2.94
Self-supervised graph models
GRACE
25.31 ± 0.16
18.27 ± 0.15
19.30 ± 0.27
56.42 ± 1.36
57.94 ± 2.58
58.50 ± 1.43
62.36 ± 2.75
64.94 ± 1.09
64.96 ± 1.86
Table 2: Link prediction and few-shot link classification. MRR for supervised link prediction and Accuracy for balanced 2-way classification over eight seeds.
Method
Grocery 5-way
Ele-Fashion 5-way
Books-NC 5-way
3
5
10
3
5
10
3
5
10
GFT
61.45 ± 2.22
64.60 ± 3.54
63.12 ± 4.07
59.08 ± 3.95
61.38 ± 3.41
62.28 ± 2.62
48.70 ± 4.96
50.95 ± 4.35
51.62 ± 4.73
UniGraph2
60.05 ± 4.13
61.25 ± 3.09
66.27 ± 3.11
53.98 ± 4.07
58.73 ± 2.95
60.38 ± 2.08
59.67 ± 4.23
61.55 ± 3.97
63.80 ± 4.02
PLANET
77.85 ± 4.06
79.88 ± 3.27
81.93 ± 3.48
70.50 ± 4.37
72.97 ± 3.55
74.85 ± 3.73
63.62 ± 4.11
67.59 ± 4.68
69.18 ± 3.90
ICE
78.96 ± 3.78
81.02 ± 3.08
83.10 ± 3.21
71.61 ± 4.01
74.07 ± 3.33
76.02 ± 3.42
64.83 ± 3.86
68.78 ± 4.32
70.42 ± 3.61
Table 3: Few-shot node classification. Focused 5-way comparison across 3, 5, and 10 shots; the complete matrix appears in Appendix Table 8 .
Figure 3: Mechanism behavior. (a–b) NC and LP access mass by grade and depth. (c–d) Accuracy (Acc.), Macro-F1 (F1), and MRR drops (percentage points) after edge rolling and protected Grade-1 route removal, respectively.
Configuration
NC Acc.
NC F1
LP MRR
ICE
83.21 ± 0.31
74.90 ± 0.38
26.24 ± 0.18
w/o Clifford product
82.00 ± 0.46
73.54 ± 0.52
24.73 ± 0.29
w/o layer 2
82.71 ± 0.37
74.34 ± 0.44
25.58 ± 0.21
w/o grade-depth bank
82.52 ± 0.39
74.12 ± 0.45
25.26 ± 0.24
w/o Grade-1 route
82.08 ± 0.43
73.50 ± 0.49
25.69 ± 0.25
w/o task access
82.29 ± 0.41
73.86 ± 0.47
25.16 ± 0.26
Table 4: Core ablation (eight-seed mean).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Graph
Nodes
Edges
Role
Classes
Exposure
Movies
16,672
218,390
NC
20
5.0
Toys
20,695
126,886
NC
18
5.0
Grocery
17,074
171,340
NC
20
5.0
Sports
50,250
356,202
LP
–
1.0
Cloth
125,839
951,271
LP
–
1.0
Ele-Fashion
97,766
199,602
NC
12
1.0
Appendix
Table 5: Foundation graph inventory and pretraining exposure.
Stage
Configuration
Value
Foundation
Input width per modality
3,584
Foundation
Hidden width and blades
256 and 8
Foundation
Transport layers
2
Foundation
RWPE steps
4
Foundation
Root batch and fanout
128 and (5,10)
Foundation
Epochs and optimizer
5 and AdamW
Appendix
Table 6: Foundation and downstream configuration.
Measure
Value
Foundation state
Foundation parameters
79.99 M
Foundation checkpoint
305.2 MiB
Fresh NC adaptation
Head parameters
22.7 to 41.3 K
Peak GPU memory
0.52 to 0.95 GiB
Appendix
Table 7: Model and adaptation scale. Architecture size and completed NC adaptation cost.
Method
Grocery 5-way
Ele-Fashion 5-way
Books-NC 5-way
3
5
10
3
5
10
3
5
10
MMGCN
48.13 ± 2.65
50.30 ± 3.13
53.73 ± 3.51
54.37 ± 2.99
57.05 ± 2.42
60.87 ± 3.08
52.93 ± 2.94
54.10 ± 2.90
56.42 ± 2.84
MGAT
50.83 ± 2.88
54.07 ± 3.18
55.53 ± 2.98
60.38 ± 2.57
61.10 ± 2.20
62.82 ± 2.18
53.15 ± 3.09
57.05 ± 2.35
58.22 ± 3.40
GRACE
57.28 ± 4.91
59.22 ± 3.50
62.22 ± 2.89
55.49 ± 4.10
59.90 ± 2.73
64.72 ± 3.04
56.84 ± 4.65
61.53 ± 3.14
59.20 ± 2.00
GraphMAE2
51.90 ± 5.76
54.60 ± 3.83
58.00 ± 4.26
57.35 ± 3.05
60.00 ± 3.53
63.65 ± 3.48
43.20 ± 2.70
46.30 ± 2.95
49.83 ± 2.60
GFM
RiemannGFM
63.63 ± 2.48
66.23 ± 2.73
67.14 ± 2.02
59.13 ± 3.71
60.24 ± 3.55
61.73 ± 2.90
54.29 ± 4.10
57.17 ± 3.83
60.54 ± 3.92
Appendix
Table 8: Few-shot node classification. Accuracy for 5-way 3-shot, 5-shot, and 10-shot tasks.
Figure 4: Component necessity. Absolute performance changes for the matched removals.
Representation
AUC
Hidden roll
0.5496
L1 Vector
0.7714
L1 V+B
0.7833
L1+L2 V+B
0.8090
Appendix
Table 9: Relation probe. Pair-relation AUC across the transport sequence.
Dataset
Accuracy
Macro-F1
Movies
10.888
5.539
Toys
3.213
1.648
Grocery
4.130
3.880
RedditS
0.472
0.302
Ele-Fash.
2.460
5.062
Books-NC
2.995
6.752
Appendix
Table 10: Task-access probe. Improvement over the generic query in percentage points.
Multimodal graph foundation models aim to learn reusable knowledge from graphs enriched with text, images, attributes, and relational topology, thereby supporting diverse graph-centric and modality-centric tasks. In practice, however, such multimodal graphs are often distributed across decentralized clients, where raw contents and local structures cannot be centrally shared due to privacy constraints. This motivates federated multimodal graph foundation learning, which requires not only transferable representation learning but also intrinsic semantic traceability under strict data isolation. Existing methods usually exchange or store knowledge through parameters, prototypes, embeddings, or compact codebooks, which support optimization and transfer but do not explicitly expose how modality evidence, node semantics, and topology context jointly support predictions. To bridge this gap, we propose FedLAB, a traceable semantic codebook framework that organizes multimodal graph knowledge into typed hierarchical codebooks for modality evidence, node semantics, and topology context. FedLAB further refines these trace units through federated semantic barycenter pre-training while keeping raw multimodal contents and graph structures local. Extensive experiments on 10 benchmarks and 6 downstream tasks show that FedLAB improves over state-of-the-art baselines by up to 7.53%, while preserving a native semantic trace interface.
Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks. Real-world graphs associate nodes with text, images, and other modalities, making multimodal graphs essential for representing complex entities and relations. Moreover, collecting labels and adapting models for every new graph domain is costly and often infeasible, motivating zero-shot transfer. Unfortunately, zero-shot transfer on multimodal graphs remains underexplored. Existing GNN-based graph foundation models typically require downstream adaptation, whereas LLM-based graph methods mainly address unimodal graphs or tasks within a single domain. This setting presents two key challenges. First, models must generalize knowledge from individual modalities while capturing transferable cross-modal relations. Second, without target-domain fine-tuning, node representations remain entangled with domain-specific structures and modality-specific characteristics, obscuring shared concepts in unseen domains. To address these challenges, we propose CHARM, a multimodal graph foundation model with hierarchical context modeling for zero-shot transfer. CHARM replaces isolated raw nodes with hierarchical graph contexts that capture multimodal semantics and cross-modal relations. These contexts map domain-specific node patterns to shared high-level concepts, reducing reliance on target-domain supervision or adaptation. A modality-aware graph context encoder integrates multimodal information with graph structure and converts the resulting representations into graph tokens for a large language model . Experiments show consistent improvements on zero-shot multimodal graph tasks.
Ankang Yang, Jitao Zhao, Di Jin +2
School of Computer Science and Technology, College of Intelligence and Computing, Tianjin University · Data Science Program, Columbian College of Arts and Sciences, The George Washington University
Multimodal-attributed graphs (MAGs), whose nodes carry modalities such as images and text alongside topological structure, now pervade applications including social platforms, e-commerce, and biomedical networks, offering richer semantic signals than single-modality graphs. In practice, such graphs are fragmented across privacy-restricted silos owned by different platforms and institutions, so learning a broadly transferable model over them demands collaborative training that never exposes raw data. This places the task at the intersection of multimodal graph learning and federated learning, yet existing methods cover only one side of it. To address the challenges from these two perspectives, we propose FedGAMMA, casting federated multimodal graph foundation learning as a two-stage semantic-structural alignment problem of federated pre-training and prompt-based fine-tuning. During pre-training, a shared-private semantic enhancer disentangles cross-modal commonality from modality-specific information, aligning it through optimal transport, a topology-aware graph fusion module decouples semantic and structural views via semantic residual graphs and dual positional encodings, and a dual-channel affinity-aware aggregation mechanism estimates client similarity from feature and graph centroids without exposing raw data. During fine-tuning, FedGAMMA adapts the pretrained encoder through lightweight graph-aware prompts, a shared prompt pool with controlled exploration, and channel-wise prompt synchronization. Experiments on twelve multimodal graph datasets show FedGAMMA consistently surpassing a broad range of baselines across downstream tasks, with gains of up to 12.96%. FedGAMMA further outperforms competitive baselines accross multi-domain datasets on multiple tasks with up to 5.71% under few-shot learning scenario.
Xunkai Li, Guohao Fu, Yuming Ai +4
Beijing Institute of Technology, Beijing, 100811, China