GraphSelect for Budgeted Representation Selection in Multimodal Graph Inference
Authors: Xu Wang, Xunkai Li, Yinlin Zhu, Rong-Hua Li
Organizations: School of Airspace Science and Engineering, Shandong University, Weihai, China · Department of Computer Science, Beijing Institute of Technology, Beijing, China · School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Multimodal graph predictors combine text, images, and relations to classify connected entities. How much of this input is needed to preserve their predictions? We study budgeted representation selection, which chooses a subset of candidate text and image vectors under a separate capacity for each modality. Predictions from the complete candidate input define the classes to preserve. The challenge is that a representation's contribution depends on the other selected inputs, while graph propagation extends its effects across nodes. Our empirical study shows that candidate rankings change with the selected input, while predicted probabilities remain informative after the class stops changing. Updating scores improves selection, and exchanging inputs can improve a subset whose capacity is already filled. These findings lead to GraphSelect, which starts from individual candidate gains and refines the subset through jointly evaluated exchanges. It screens promising removals and additions, accepts an exchange when it reduces the prediction loss, and updates the scores. Experiments on six graphs show higher mean objective recovery than six attribution and explanation methods adapted to the selection task. Across nine trained architectures on two graphs, retaining 20% of the candidate representations per modality gives a mean accuracy drop of 0.10 percentage points relative to full candidate input, preserving classification performance with substantially fewer text and image representations.
Figures & tables
Figure 1: A neighboring image can help preserve a graph prediction. Each box represents a product. The two jicama products have an also-viewed relation. Full input provides the reference prediction. With a limited budget, omitting the neighbor’s image lowers confidence in Produce. Exchanging the fern’s image for the neighbor’s restores most of that confidence at the same budget. The target’s own inputs remain available. Gray inputs are not selected and use modality means.
Figure 2: Three design probes. (a) Agreement between deletion rankings and first-addition gains. (b) Recovery under three prediction objectives. (c) Improvements from updating addition scores at each step.
Figure 3: GraphSelect overview. The graph predictor supplies the full-input target classes and evaluates each proposed subset. Singleton gains initialize selection, followed by screening and joint exchange evaluation. Accepted exchanges preserve both modality quotas and trigger score updates.
Method
Recovery at 20% capacity
Mean over 18 settings
SemArt
Grocery
Movies
Toys
RedditS
EleFashion
5%
10%
20%
Attribution and explanation adapters
Shapley Sampling
44.7 ± 1.2
37.3 ± 1.5
53.7 ± 1.1
36.1 ± 0.2
32.1 ± 1.3
47.5 ± 4.2
14.3 ± 3.4
24.9 ± 5.1
41.9 ± 7.7
Integrated Gradients
44.2 ± 2.3
37.1 ± 1.1
51.7 ± 1.5
35.2 ± 0.1
30.2 ± 0.7
43.3 ± 4.1
13.9 ± 3.7
23.9 ± 5.2
40.3 ± 7.3
Feature Ablation
47.5 ± 0.7
37.7 ± 1.7
55.2 ± 1.9
37.0 ± 0.2
33.7 ± 1.2
52.5 ± 2.6
15.2 ± 3.7
26.0 ± 5.4
43.9 ± 8.3
GNNExplainer
22.5 ± 1.2
28.0 ± 1.6
29.0 ± 1.8
24.9 ± 0.8
18.0 ± 0.4
28.3 ± 3.9
7.3 ± 2.3
13.9 ± 3.5
25.1 ± 4.4
Table 1: Objective recovery (%). Dataset columns summarize three seeds, and overall means summarize 18 graph and seed settings. Values are mean ± standard deviation. GraphSelect uses the budgeted search setting.
Figure 4: Selection quality. (a) Mean recovery at three capacities. (b) Dataset results at 20% capacity. (c) Paired recovery gains over Forward. All methods use the same candidate pool and modality quotas.
Figure 5: Capacity and prediction quality. Objective recovery and downstream performance gaps at four capacities on Grocery and RedditS. Red dashed lines indicate the corresponding full-input results.
Capacity (quota)
Objective recovery (%)
Accuracy drop
Macro-F1 drop
CE increase
2.5% (6)
24.81 ± 11.64
0.237 ± 0.270
0.202 ± 0.227
9.14 ± 16.80
5.0% (11)
37.56 ± 15.50
0.207 ± 0.256
0.171 ± 0.201
8.14 ± 15.62
10.0% (22)
56.02 ± 17.91
0.156 ± 0.237
0.119 ± 0.179
6.53 ± 13.12
15.0% (33)
68.38 ± 18.27
0.124 ± 0.225
0.084 ± 0.172
5.40 ± 11.31
20.0% (44)
77.61 ± 17.92
0.100 ± 0.212
0.065 ± 0.166
4.52 ± 10.15
30.0% (66)
90.39 ± 16.62
0.069 ± 0.176
0.040 ± 0.145
3.19 ± 7.50
Table 2: GraphSelect across six capacities. Values summarize 54 settings. Accuracy and Macro-F1 drops are in percentage points. The increase in cross-entropy to ground-truth labels is multiplied by 103 .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Entry
Operation in this paper
Published method adapters
Shapley Sampling, Integrated Gradients, Feature Ablation, GNNExplainer, and GOAt
Published scoring or explanation methods adapted to rank the same candidate representations
Zorro
A published sufficient-subset method adapted to the same candidates and modality quotas
Study controls
CE and TV deletion
Remove one candidate from the full input and measure cross entropy or total variation
KL divergence and prediction agreement
Build from the reference input by matching probabilities or predicted classes
Appendix
Table 3: Selection methods and controls. The rows summarize how each method selects representations.
Control
5%
10%
20%
Input state
CE deletion
17.3 ± 3.9
28.9 ± 6.0
47.0 ± 8.4
TV deletion
14.9 ± 3.4
25.1 ± 4.5
42.0 ± 6.4
Prediction target
KL divergence
13.9 ± 3.2
23.9 ± 5.0
40.7 ± 6.6
Prediction agreement
9.9 ± 2.8
15.7 ± 4.2
25.3 ± 4.7
Appendix
Table 4: Design controls. Objective recovery is mean ± standard deviation.
Figure 6: Selection across capacities. Mean objective recovery and downstream performance gaps relative to full input for four selectors on Grocery and RedditS at six capacities.
Figure 7: Variation across architectures. Full-input prediction quality and changes in recovery and accuracy gap from 20% to 30% capacity for nine architectures.
Architecture
Grocery
RedditS
Accuracy
Macro-F1
Accuracy
Macro-F1
Graph architectures
GCN
73.47 ± 2.15
65.01 ± 2.19
91.16 ± 1.00
85.73 ± 1.93
GraphSAGE
71.36 ± 2.23
62.03 ± 3.17
91.32 ± 0.41
85.96 ± 0.36
GAT
67.54 ± 2.17
57.43 ± 2.27
90.46 ± 0.32
85.12 ± 0.77
Multimodal graph architectures
Appendix
Table 5: Full input quality. Values summarize three seeds for nine architectures.
Architecture
20% capacity
30% capacity
Recovery
Accuracy drop
Recovery
Accuracy drop
Graph architectures
GCN
91.0 ± 4.7
0.021 ± 0.034
103.1 ± 5.8
0.006 ± 0.028
GraphSAGE
93.8 ± 15.8
0.015 ± 0.029
104.2 ± 14.4
0.000 ± 0.018
GAT
97.9 ± 9.4
0.030 ± 0.043
108.9 ± 10.7
0.000 ± 0.046
Multimodal graph architectures
Appendix
Table 6: Architecture details. Values summarize two graphs and three seeds.
Variant
Modification
Rvariant−RGraphSelect Loss ×104
GraphSelect W/T/L
Construction alternative
Forward construction
Repeated additions with rescoring
0.81 ± 3.75
57/68/1
Component removals
w/o exchange refinement
Return the initial set
4.60 ± 10.67
112/14/0
w/o current-state updates
Keep initial scores during exchange
0.76 ± 3.85
62/62/2
Appendix
Table 7: Search ablations over 126 settings excluded from tuning.
Figure 8: Search ablations across architectures. Loss differences between GraphSelect and its search controls, scaled by 104 . Positive values favor GraphSelect.
Dataset
Graph size
Classes
Nodes
Adjacency nonzeros
SemArt
21,382
1,173,014
10
Grocery
17,074
142,262
20
Movies
16,672
160,802
20
Toys
20,695
113,402
18
RedditS
15,894
283,080
20
Appendix
Table 8: Multimodal graph datasets. Numbers of nodes, nonzero adjacency entries, and classes in the processed graphs used for evaluation.
Dataset
Full graph
No candidate propagation
No graph terms
Affected targets
SemArt
84.11±0.09
84.09±0.11
83.42±0.91
277.3
Grocery
75.01±0.61
75.01±0.59
60.56±5.47
183.9
Movies
47.55±1.26
47.56±1.24
38.33±0.43
108.7
Toys
76.85±0.64
76.86±0.61
67.42±0.97
55.4
RedditS
94.69±0.30
94.67±0.28
90.53±0.36
17.4
EleFashion
80.06±0.38
80.06±0.38
82.21±0.25
120.6
Appendix
Table 9: Full-input test accuracy in the graph ablation study. Entries are mean ± population standard deviation over three seeds, in percent. The last column gives the mean number of target nodes affected by one candidate representation in the full graph predictor. Both ablations reduce this number to one.
Predictor
Capacity
R(∅)
R(E)
Initial ranking minus GraphSelect
Forward minus GraphSelect
Full graph
5%
0.522555
0.516625
1.599
0.352
10%
0.522555
0.516625
4.026
0.776
20%
0.522555
0.516625
6.567
1.756
No candidate propagation
5%
0.522529
0.517774
2.317
0.557
10%
0.522529
0.517774
3.728
0.983
20%
0.522529
0.517774
5.499
1.254
Appendix
Table 10: Selection losses in the graph ablation study. Each entry averages 18 graph and seed settings. The first two losses use each condition's own full-input target classes. The two rightmost columns report the control loss minus the GraphSelect loss, multiplied by 105 . Positive values favor GraphSelect .
Department of Computer Science, Beijing Institute of Technology, Beijing, China · School of Airspace Science and Engineering, Shandong University, WeiHai, China · School of Computer Science and Engineering, Sun Yat-sen University, GuangZhou, China