Organizations: School of Computer Science, Peking University · Key Laboratory of High Confidence Software Technologies, Peking University, Ministry of Education · School of Computer Science, Wuhan University · School of Information Science and Technology, Tibetan Language Intelligence National Key Laboratory, Tibet University
Molecular property prediction is central to drug development and materials discovery, but experiments are costly and labeled data are scarce. Context-aware methods use auxiliary assay labels to support few-shot prediction, and recent work supervises property relations with label agreement. However, label agreement is sensitive to class marginals and does not directly capture dependence between properties. We propose CalibHyper, a chance-corrected relational hypergraph method based on the joint label distribution. CalibHyper subtracts an independence baseline from the ordered four-state label distribution and shrinks the residual according to the number of joint observations. A swap-equivariant relation head estimates these residuals, which choose the auxiliary properties for each molecule and set the sign and weight of their hyperedge messages. On thirteen datasets from five benchmarks, in both 1-shot and 10-shot settings, CalibHyper and its ablation settings achieve ROC-AUC competitive with the strongest reported results.
Figures & tables
Figure 1: Motivation. (A) Few-shot prediction with auxiliary assays. (B) Chance agreement under rare positives. (C) Binary targets merge states 01 and 10 . (D) Structure-dependent association in Tox21. (E) Similar estimates from unequal numbers of joint observations. B, C, and E are schematic.
Figure 2: Label statistics (Appendix H ). (a) Observed and independence agreement and (b) mean Cohen’s κ per dataset (TC: ToxCast). (c) MUV-713 and MUV-733. (d) A CEETOX increase and decrease pair. (e) κ with and without phosphorus. (f) ∣Δκ∣ after one 11→10 relabeling.
Figure 3: Architecture. (A) Context graph and encoder. (B) C4RC: chance-corrected four-state relation head. (C) MSRHA: signed messages on top- k hyperedges. (D) Prediction and training.
Tox21
SIDER
MUV
ToxCast
PCBA
Method
10-shot
1-shot
10-shot
1-shot
10-shot
1-shot
10-shot
1-shot
10-shot
1-shot
Siamese
80.40 ± 0.35
65.00 ± 1.58
71.10 ± 4.32
51.43 ± 3.31
59.96 ± 5.13
50.00 ± 0.17
–
–
–
–
ProtoNet
74.98 ± 0.32
65.58 ± 1.72
64.54 ± 0.89
57.50 ± 2.34
65.88 ± 4.11
58.31 ± 3.18
63.70 ± 1.26
56.36 ± 1.54
64.93 ± 1.94
55.79 ± 1.45
MAML
80.21 ± 0.24
75.74 ± 0.48
70.43 ± 0.76
67.81 ± 1.12
63.90 ± 2.28
60.51 ± 3.12
66.79 ± 0.85
65.97 ± 5.04
66.22 ± 1.31
62.04 ± 1.73
TPN
76.05 ± 0.24
60.16 ± 1.18
67.84 ± 0.95
62.90 ± 1.38
65.22 ± 5.82
50.00 ± 0.51
62.74 ± 1.45
50.01 ± 0.05
–
–
EGNN
81.21 ± 0.16
79.44 ± 0.22
72.87 ± 0.73
70.79 ± 0.95
65.20 ± 2.08
62.18 ± 1.76
63.65 ± 1.57
61.02 ± 1.94
69.92 ± 1.85
62.14 ± 1.58
Table 1: ROC-AUC (%). CalibHyper: mean ± sample standard deviation over three seeds; baselines from Wang et al. (2026) . ToxCast averages its nine subsets; a dash marks an unreported result.
Ablation
Tox21
SIDER
MUV
CalibHyper (full)
96.82 ± 1.70
96.89 ± 2.33
97.19 ± 1.21
w/o C4RC
90.08 ± 2.96
90.78 ± 3.33
93.82 ± 2.19
w/o MSRHA
87.48 ± 3.52
86.46 ± 1.92
67.23 ± 1.94
w/o C4RC and MSRHA
93.97 ± 2.80
89.20 ± 3.90
87.15 ± 3.53
w/o chance correction
93.00 ± 3.13
88.39 ± 1.67
87.47 ± 3.97
w/o sign separation
95.81 ± 1.12
88.75 ± 1.93
67.43 ± 2.52
Table 2: Component ablation (10-shot ROC-AUC, %, three runs).
Figure 4: Single-run ablations, peak 10-shot ROC-AUC; details in Appendix C . (a) Masked auxiliary labels. (b) Auxiliary-property count N ( 12∗ : 11 during MUV training). (c) Routing count k .
Case
Molecule
Dataset and target
GT
ReCoG
Pin-Tuning
CalibHyper
C1
Tox21 SR-HSE
Neg.
0.99 Pos.
0.59 Pos.
7.39×10−6 Neg.
C2
Tox21 SR-HSE
Neg.
1.00 Pos.
0.57 Pos.
3.77×10−8 Neg.
C3
CEETOX ESTRONE_up
Neg.
1.00 Pos.
0.83 Pos.
7.98×10−5 Neg.
C4
CEETOX ESTRADIOL_up
Neg.
0.98 Pos.
0.87 Pos.
0.02 Neg.
Table 3: Molecular cases (10-shot): positive-class score and the label predicted at threshold 0.5 (bold: correct; GT: ground truth). CEETOX target names omit the prefix CEETOX_H295R_.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
M , T
set of molecules, set of properties
Ttrain , Ttest
disjoint meta-training and meta-test properties
yi,p∈{0,1,⋆}
label of molecule i for property p ; ⋆ means unmeasured
oi,p
observation indicator 1[yi,p=⋆]
Eτ=(Sτ,Qτ)
episode with target τ , its support set and query set
Aτ , N
auxiliary properties of the episode and their number
Appendix
Table B.1: Notation of the main text.
Symbol
Meaning
Ipq , Ppq
joint observation set and the empirical law of the pair on it (Definition E.2 )
πp∣q
unsmoothed positive rate of p on jointly observed molecules (Definition E.2 )
bpq∗ , ri,p,q∗
baseline and ideal target built from exact unsmoothed marginals (Definitions E.4 and E.6 )
χ∅ , χp , χq , χpq
Walsh basis of R4 , with χpq=χp⊙χq (Definition E.5 )
sp , sq
labels in ±1 form, sp=1−2Yp (Definition E.5 )
Pswap , S
swap acting on output coordinates and on inputs (Definition E.7 )
Appendix
Table B.2: Additional symbols of Appendices E to G .
GIN layers / width d
5 / 300
Training episodes
2,000
Laplace constant α
1
Query batch (train/test)
16 / 64
Routing count k
5
Shrinkage constant n0
5
Appendix
Table C.1: Hyperparameters of the main experiments.
ROC-AUC
AP
Masked
Ablation setting
Peak
Last-5
Final
Peak
Final
Tox21
50%
C4RC, unsigned MSRHA
94.87
73.73
86.32
85.53
40.31
50%
w/o MSRHA
94.89
79.70
90.07
66.13
49.88
50%
Head switched to binary
98.80
88.44
80.66
93.30
30.17
70%
C4RC, unsigned MSRHA
85.62
72.15
57.30
41.84
22.39
Appendix
Table C.2: Missing auxiliary labels in the 10-shot setting (%).
ROC-AUC
AP
N
Ablation setting
Peak
Last-5
Final
Peak
Final
Tox21
1
C4RC, unsigned MSRHA
84.46
82.58
82.68
38.22
34.26
1
w/o MSRHA
85.39
82.44
81.59
38.17
32.28
2
C4RC, unsigned MSRHA
88.57
83.13
82.84
47.35
36.14
2
w/o MSRHA
81.46
74.68
75.29
32.45
24.34
Appendix
Table C.3: Number of auxiliary properties in the 10-shot setting (%).
k
Tox21
MUV
CEETOX
1
93.01
89.43
76.38
5
99.47
92.32
91.89
10
97.87
86.42
80.69
Appendix
Table C.4: Peak ROC-AUC (%) for different routing counts k in the 10-shot setting.
Method
APR
ATG
BSK
CEETOX
CLD
NVS
OT
TOX21
Tanguay
10-shot
ProtoNet
73.58
59.26
70.15
66.12
78.12
65.85
64.90
68.26
73.61
MAML
72.66
62.09
66.42
64.08
74.57
66.56
64.07
68.04
77.12
EGNN
80.33
66.17
73.43
66.51
78.85
71.05
68.21
76.40
85.23
Pre-PAR
86.09
72.72
82.45
72.12
83.43
74.94
71.96
82.81
88.20
Pre-GS-Meta
90.15
82.54
88.21
74.19
86.34
76.29
74.47
90.63
91.47
Appendix
Table C.5: ROC-AUC (%) on the nine ToxCast subsets. CalibHyper: mean over three runs. Baselines are the means reported in Tables 8 and 9 of Wang et al. (2026) ; method groups follow Table 1 . † Ablation setting with binary relation targets and the MSRHA gate fixed at γ=0 (Proposition 5 ).
State
Observed count
Expected under independence
Both negative
3117
3117.0
Only MUV-733 positive
1
1.00
Only MUV-713 positive
3
3.00
Both positive
0
0.001
Appendix
Table H.1: Joint label states of MUV-713 and MUV-733.
dn negative
dn positive
up negative
230
151
up positive
119
0
Appendix
Table H.2: Joint label states of the OHPROG increase (up) and decrease (dn) assays.
Molecules
Joint observations
Cohen’s κ
All jointly observed
6657
+0.018
Containing phosphorus
213
+0.240
Other
6444
+0.013
Appendix
Table H.3: Cohen’s κ between NR-AR and SR-p53 on subsets of Tox21.
Property pair
Joint observations
Joint positives
Original κ^
After relabeling
∣Δκ^∣
MUV-652, MUV-712
3066
1
+0.2211
−0.0013
2.2×10−1
PCBA-1460, PCBA-485364
204894
1151
+0.2231
+0.2230
1.9×10−4
Appendix
Table H.4: Effect of relabeling one joint positive as (1,0) .
Machine learning is transforming molecular sciences by accelerating property prediction, simulation, and the discovery of new molecules and materials. Acquiring labeled data in these domains is often costly and time-consuming, whereas large collections of unlabeled molecular data are readily available. Standard semi-supervised learning methods often rely on label-preserving augmentations, which are challenging to design in the molecular domain, where minor changes can drastically alter properties. In this work, we show that semi-supervised methods that rely on an ensemble consensus can boost predictive accuracy across a diverse range of molecular datasets, task types, and graph neural network architectures. We find that training with an ensemble consensus objective increases robustness in models and exhibits an effect similar to knowledge distillation; an individual member of an ensemble trained this way outperforms a full ensemble trained in a traditional supervised fashion in almost all cases. In addition, this type of semi-supervised training reduces calibration error.
Motivation: Noisy labels are a common challenge in molecular property prediction because molecular annotations are often obtained from assays, curated databases, or weak annotation pipelines rather than directly observed clean biological states. Treating recorded labels as reliable supervision can cause models to memorize corrupted observations and learn misleading molecular evidence. In multimodal molecular representation learning, this issue can be amplified by graph-text fusion or alignment, which may propagate label-induced errors across modalities. Results: We propose MOLAR, a noise-aware framework for learning multimodal molecular representations from noisy labels. MOLAR separates latent clean-property inference from recorded-label observation: graph and text views contribute residual evidence to a clean-property distribution, and a categorical label-observation channel maps this distribution to recorded labels for training. This formulation derives posterior label reliability and modality-specific molecular evidence from the model. Experiments on naturally noisy molecular benchmarks and controlled label-flipping benchmarks show that MOLAR consistently outperforms representative baselines. Visualization analyses further show that MOLAR provides interpretable reliability and modality-evidence diagnostics.
Yingxu Wang, Kunyu Zhang, Nan Yin +2
Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence, AI Diyafah St, 7909, Abu Dhabi, United Arab Emirates · Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong, China · International College, Zhengzhou University, Daxue North Road, 450000, Henan, China +2
Accurate molecular property prediction requires both statistical reliability and chemical reasoning. Graph neural networks can be calibrated directly on labeled assays but remain limited by the coverage of their training data. Large language models (LLMs) can compare molecular evidence and articulate chemical rationales, yet are unreliable as standalone quantitative predictors. The central challenge is therefore to determine when an LLM should influence a calibrated model and by how much. Here we present CoMPASS, a retrieval-calibrated framework for small-large model collaboration. CoMPASS retains a graph attention network (GAT) as the predictive anchor, retrieves locally relevant training molecules, provides attention-grounded evidence to an LLM, and converts its proposal into a bounded correction through an agreement-aware gate. Across six classification and two regression benchmarks, CoMPASS improves the GAT anchor in regions of correctable uncertainty while limiting LLM intervention in high-confidence regimes. Ablations show that the gains arise from validation-calibrated retrieval and bounded fusion rather than prompting alone. These results suggest that generative reasoning should augment calibrated prediction through evidence-grounded, controlled corrections rather than direct output replacement. Code is available at https://github.com/littlepeachs/CoMPASS.
Wentao Li, Jiangjie Qiu, Yijun Li +2
Beijing Key Laboratory of Artificial Intelligence for Advanced Chemical Engineering Materials · State Key Laboratory of Chemical Engineering and Low-Carbon Technology · Department of Chemical Engineering, Tsinghua University