Molecular Property Prediction under Structural Shift with Tabular Foundation Models
Organizations: Nums AI
Abstract
Predicting molecular properties for compounds that differ structurally from labeled training molecules is important for drug discovery and materials design. Tabular foundation models (TFMs) offer a promising approach through in-context learning, but their performance under structural shifts and the value of molecular comparisons in this setting remain underexplored. We study structural generalization in molecular property prediction and introduce MolPAIR (Molecular Pair-Augmented In-context Refinement), a framework that combines molecule-level and molecular-pair contexts without task-specific parameter updates. A global tabular foundation model (TFM) first predicts a query's property from labeled molecular examples. A second frozen TFM predicts differences in prediction errors between the query and labeled reference molecules, using these comparisons to refine the initial prediction. Across 58 MoleculeACE and Polaris tasks, CheMeleon representations combined with TabPFN-3 already outperform each evaluated baseline on a majority of tasks. MOLPAIR further improves this predictor on 46 of 58 tasks, with gains across four molecular representations and three TFM backbones. These results show that explicit molecular comparisons can strengthen tabular in-context learning for structural generalization while keeping the molecular encoder and pretrained model weights fixed. The code and datasets are available at https://github.com/nums-ai/MolPAIR.
Figures & tables
| Method | MACE RMSE | MACE rank | Polaris cls. rank | Polaris reg. rank | Top one | Elo |
|---|---|---|---|---|---|---|
| MolPAIR | 0.885 0.100 | 2.00 | 2.18 | 2.59 | 29 | 1360 [1289, 1455] |
| Global baseline | 0.893 0.097 | 2.93 | 2.36 | 4.00 | 7 | 1246 [1196, 1310] |
| MiniMol (tuned) | 0.931 0.107 | 5.30 | 3.91 | 3.12 | 14 | 1122 [1060, 1190] |
| ExtraTrees (tuned) | 0.921 0.102 | 4.90 | 5.36 | 5.12 | 2 | 1062 [1007, 1123] |
| CheMeleon (fine-tuned) | 0.930 0.089 | 5.40 | 6.91 | 4.24 | 3 | 1035 [976, 1097] |
| XGBoost (tuned) | 0.924 0.101 | 5.03 | 6.36 | 5.59 | 2 | 1026 [969, 1088] |
| Cutoff | Global RMSE | MolPAIR RMSE | W/L |
|---|---|---|---|
| 0.40 | 0.9939 0.1745 | 0.9926 0.1792 | 38/14 |
| 0.50 | 0.9570 0.1318 | 0.9521 0.1302 | 41/13 |
| 0.60 | 0.8932 0.0966 | 0.8849 0.1004 | 46/12 |
| 0.70 | 0.8357 0.0791 | 0.8327 0.0790 | 42/16 |
| Control | Pooled Elo [95% CI] | W/L | RMSE reduction [95% CI] |
|---|---|---|---|
| MolPAIR (full) | Reference | – | Reference |
| Global baseline | 295.7 [209.6, 393.9] | 46/12 | 0.00819 [0.00235, 0.01388] |
| Two-global-TFM ensemble | 209.6 [111.1, 321.3] | 44/14 | 0.00547 [ 0.00032, 0.01093] |
| NN residual | 112.0 [37.9, 191.7] | 41/17 | 0.00402 [ 0.00110, 0.00930] |
| Single-molecule residual | 160.9 [78.7, 253.9] | 43/15 | 0.00768 [0.00251, 0.01295] |
| Shuffled relation labels | 88.9 [14.0, 168.8] | 41/17 | 0.00405 [ 0.00100, 0.00947] |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Value |
|---|---|
| Global representation | CheMeleon, 2,048 input features before filtering |
| Relation representation | RDKit2D, up to 996 selected features |
| Global and relation backbone | TabPFN-3, eight estimators |
| Cross-fitting | Five folds within the outer training split |
| Training anchors | Up to eight non-self anchors within the same OOF fold |
| Relation context budget | At most 8,192 rows, deterministic subsampling |
| Suite | Endpoints | Planned seeds | Task types | Primary metrics |
|---|---|---|---|---|
| MoleculeACE | 30 | 5 | regression | RMSE |
| Polaris | 28 | 5 | mixed | MAE , MSE , Pearson, Spearman, PR-AUC , ROC-AUC |
| Total | 58 | 5 | mixed | released metric |
| Max. similarity | Valid splits | Query–seed predictions | Endpoints | Direct Elo |
|---|---|---|---|---|
| 0.40 | 219 / 290 | 26,488 | 52 | 171.9 |
| 0.50 | 268 / 290 | 42,159 | 54 | 197.7 |
| 0.60 | 286 / 290 | 55,733 | 58 | 231.1 |
| 0.70 | 289 / 290 | 72,474 | 58 | 166.3 |
| Representation | Dimension | Source or use |
| RDKit2D | data dependent | deterministic descriptors |
| Morgan | 1,024 to 4,096 | radius two or three, bit or count |
| CheMeleon | 2,048 | frozen directed message-passing fingerprint |
| MiniMol | 512 | frozen pretrained GINE fingerprint |
| Monroe | released dimension | frozen pretrained molecular fingerprint |
| Model and trials | Search space |
|---|---|
| Random Forest, 30 | max_features in ; max_depth in ; min_samples_leaf from 1 to 16 (log); min_samples_split from 2 to 32 (log); bootstrap on or off; and, for classification, class weight in {none, balanced, balanced subsample}. |
| ExtraTrees, 30 | The same space as Random Forest. |
| XGBoost, 50 | Learning rate from 0.005 to 0.2 (log), depth 3–12, child weight 0.01–20 (log), row subsampling 0.5–1.0, column subsampling 0.4–1.0, gamma 0–5, L1 –10 (log), L2 –100 (log), and maximum bins in . |
| LightGBM, 50 | Learning rate 0.005–0.2 (log), leaves 15–255 (log), depth in , minimum child rows 5–100 (log), row subsampling 0.5–1.0, column subsampling 0.4–1.0, L1 –10 (log), L2 –100 (log), and maximum bins in . |
| Model and trials | Search space |
|---|---|
| Chemprop, 20 | Message width in , message-passing depth 3–6, mean or sum aggregation, feed-forward width in , one to three feed-forward layers, dropout 0–0.4 in steps of 0.1, batch size in , one to five warmup epochs, maximum learning rate – (log), initial-rate ratio 0.02–0.2 (log), final-rate ratio 0.001–0.05 (log), and class balancing on or off for classification. |
| MiniMol head, 20 | Hidden-width multiplier in , one to three hidden layers, dropout 0–0.3 in steps of 0.1, learning rate – (log), weight decay – (log), batch size in , epochs in , and warmup epochs in . |
| Method | MACE RMSE | RMSE reduction [95% CI] | W/L | Direct Elo [95% CI] |
|---|---|---|---|---|
| Regression on 47 tasks | ||||
| MolPAIR | 0.8849 | Reference | – | – |
| Global baseline | 0.8931 | 0.00828 [0.00246, 0.01427] | 40/7 | 298.1 [184.0, 451.9] |
| PADRE (RF) | 1.0012 | 0.11632 [0.08687, 0.14614] | 44/3 | 451.9 [298.1, 809.9] |
| DeepDelta | 1.0588 | 0.17393 [0.14058, 0.20757] | 44/3 | 451.9 [298.1, 809.9] |
| iMoLD | 1.0796 | 0.19478 [0.15370, 0.24499] | 47/0 | 809.9 |
| Endpoint | Metric | MolPAIR | Global | MiniMol | XGBoost | Chemprop |
|---|---|---|---|---|---|---|
| 1862 Ki | RMSE | |||||
| 1871 Ki | RMSE | |||||
| 2034 Ki | RMSE | |||||
| 2047 EC50 | RMSE | |||||
| 204 Ki | RMSE | |||||
| 2147 Ki | RMSE |
| Endpoint | Metric | MolPAIR | Global | MiniMol | XGBoost | Chemprop |
|---|---|---|---|---|---|---|
| adme fang hclint 1 | Pearson | |||||
| adme fang hppb 1 | Pearson | |||||
| adme fang perm 1 | Pearson | |||||
| adme fang rclint 1 | Pearson | |||||
| adme fang rppb 1 | Pearson | |||||
| adme fang solu 1 | Pearson |
| Backbone | Representation | Global RMSE | MolPAIR RMSE | Elo |
|---|---|---|---|---|
| TabPFN-3 | RDKit2D | +197.2 | ||
| TabPFN-3 | CheMeleon | +231.1 | ||
| TabPFN-3 | MiniMol | +290.8 | ||
| TabPFN-3 | Monroe | +166.3 | ||
| TabICLv2 | RDKit2D | +166.3 | ||
| TabICLv2 | CheMeleon | +269.4 |
| Control | Regression W/L | Classification W/L | RMSE reduction [95% CI] | Direct Elo [95% CI] |
|---|---|---|---|---|
| Gated single residual | 37/10 | 6/5 | 0.00768 [0.00260, 0.01292] | 181.4 [85.0, 290.8] |
| Row-matched single residual | 41/6 | 6/5 | 0.00479 [0.00110, 0.00826] | 249.7 [151.7, 368.8] |
| Control | W/L | RMSE reduction [95% CI] |
|---|---|---|
| In-sample residuals | 44/14 | 0.01547 [0.00839, 0.02274] |
| Random anchors | 41/17 | 0.00373 [-0.00015, 0.00763] |
| No bidirectional symmetry | 33/25 | 0.00042 [-0.00015, 0.00105] |
| No uncertainty weighting | 28/30 | 0.00003 [-0.00038, 0.00045] |
| Variant | MACE RMSE | Full W/T/L | Full variant Elo [95% CI] |
| Full reference | 0.88496 | – | – |
| No global-output features | 0.88577 | 35/0/23 | +72.4 [ 11.9, 166.3] |
| Pair predictor, similarity-only weights | 0.88505 | 24/0/34 | 60.1 [ 151.7, 23.8] |
| NN residual, similarity-only weights | 0.88899 | 41/0/17 | +151.7 [60.1, 269.4] |
| Uniform weights | 0.88639 | 41/0/17 | +151.7 [60.1, 249.7] |
| 8 query anchors | 0.88436 | 27/0/31 | 23.8 [ 110.7, 60.1] |
| Weighting | Pair RMSE | NN RMSE | W/L | Elo [95% CI] |
|---|---|---|---|---|
| Similarity-only | 0.88505 | 0.88899 | 40/18 | +137.7 [47.9, 249.7] |
| Gate | MACE RMSE | Adaptive W/L | Direct Elo [95% CI] |
|---|---|---|---|
| Adaptive | 0.885112 | – | – |
| Fixed 0.10 | 0.889325 | 44/14 | [97.7, 314.1] |
| Fixed 0.20 | 0.886889 | 32/26 | [ , 124.0] |
| Fixed 0.25 | 0.886215 | 27/31 | [ , 60.1] |
| Fixed 0.2674 | 0.886066 | 26/32 | [ , 47.9] |
| Fixed 0.30 | 0.885905 | 24/34 | [ , 23.8] |
| Method | Hit identification | Lead optimization |
|---|---|---|
| PR-AUC | Cluster Spearman | |
| NN with ECFP4 | 0.4606 | 0.1579 |
| Gradient boosting with ECFP4 | 0.4594 | 0.2140 |
| SVM with ECFP4 | 0.4065 | 0.3124 |
| Chemprop | 0.5040 | 0.2172 |
| CheMeleon with TabPFN-3 | 0.5551 | 0.3004 |
| Method | MoleOOD | ERM | MixUp | Global TFM | MolPAIR | GroupDRO |
|---|---|---|---|---|---|---|
| ROC-AUC | 0.6946 | 0.6739 | 0.6714 | 0.6704 | 0.6687 | 0.6423 |
| Regression | Classification | |||||
|---|---|---|---|---|---|---|
| Similarity | Improved (%) | Worsened (%) | Improved (%) | Worsened (%) | ||
| 6,824 | 57.1 | 41.9 | 2,494 | 64.8 | 33.0 | |
| 12,415 | 57.2 | 42.8 | 3,624 | 67.4 | 32.6 | |
| 10,596 | 56.5 | 43.5 | 3,757 | 70.0 | 30.0 | |
| 11,298 | 54.8 | 45.2 | 3,071 | 74.0 | 26.0 | |
| 419 | 53.2 | 46.8 | 104 | 71.2 | 28.8 | |
| Diagnostic | Regression | Classification | All |
|---|---|---|---|
| Endpoints | 47 | 11 | 58 |
| Task–seed runs | 231 | 55 | 286 |
| Query–seed predictions | 42,026 | 13,707 | 55,733 |
| No-anchor predictions | 474 | 657 | 1,131 |
| No-anchor fallback (%) | 1.13 | 4.79 | 2.03 |
| Active correction (%) | 98.71 | 94.81 | 97.75 |
| Scenario | Global | MolPAIR | MiniMol head |
|---|---|---|---|
| Cold end to end | 5.17 | 46.11 | 13.91 |
| Warm task adaptation | 0.17 | 13.41 | 8.60 |
| Update with 64 labels | 0.15 | 4.49 | 2.09 |
| Warm query with 1 molecule | 1.10 | 1.64 | 0.002 |
| Warm query with 128 molecules | 1.14 | 3.91 | 0.003 |
| Warm query with 10,000 molecules | 7.16 | 228.95 | 0.004 |