RelICL: Training-free Relational Learning with Tabular Foundation Models
Organizations: Data and Web Science Group University of Mannheim Mannheim, Germany
Abstract
Tabular foundation models achieve state-of-the-art performance on single-table tasks without any training. Recent work suggests that they are also well-suited for relational learning via deep feature synthesis (DFS), which flattens a relational schema into a single table by adding aggregates of the other tables' columns as features. This approach is appealing because it directly benefits from improvements to or customization of the underlying tabular foundation model. In this paper, we identify two key problems with DFS: feature explosion and interaction blindness. The first problem arises because the number of DFS features grows quickly as the schema becomes more complex, limiting scalability and performance. The second problem arises because column-wise aggregates do not account for feature interactions, limiting performance. We propose and explore an alternative method termed RelICL, which keeps the benefits of DFS but alleviates these two problems. At its heart, RelICL propagates and fuses information step by step through the schema graph, using the same tabular foundation model that is eventually used for prediction to do so. In our experimental study using RelBench tasks, RelICL was on par with the strongest approach based on deep feature synthesis.
Figures & tables
| Dataset | AUROC | AUROC | |
|---|---|---|---|
| bioresponse | 87.7 1.0 | 87.5 0.8 | -0.3 0.3 |
| churn | 92.8 0.8 | 92.6 0.6 | -0.2 0.2 |
| compass | 72.3 0.7 | 72.4 0.7 | +0.1 0.1 |
| credit-g | 82.2 1.8 | 83.1 0.4 | +0.9 1.4 |
| jasmine | 89.8 1.1 | 86.8 0.4 | -3.0 0.6 |
| Pretrained (rel.) | Pretrained (tab.) | ||||
|---|---|---|---|---|---|
| Model | RT | PluRel | RDBLearn | TabPFN-Rel | RelICL |
| RT | – | 0 6 / 18 ∗ | 0 0 / 8 | 0 2 / 10 | 0 7 / 18 ∗ |
| PluRel | 11 / 18 ∗ | – | 0 0 / 8 | 0 2 / 10 | 0 5 / 18 ∗ |
| RDBLearn | 0 8 / 8 0 | 0 8 / 8 0 | – | 0 1 / 15 | 0 5 / 15 |
| TabPFN-Rel | 0 8 / 10 | 0 8 / 10 | 13 / 15 | – | 10 / 21 |
| RelICL | 11 / 18 ∗ | 13 / 18 ∗ | 10 / 15 | 11 / 21 | – |
| Table 5: Regression, MAE (lower is better) / (%, higher is better). The table gives the per-task numbers behind the regression part of Tab. 4 . It shows both metrics because RT and PluRel report only , while RDBLearn and TabPFN-Rel report only MAE. RelICL reports both, so it can be compared with either group. Bold marks the best value per task and metric. | Table 6: RelICL with the TabPFN backbone. |
| Pretrained (rel.) Pretrained (tab.) Task RT PluRel RDBLearn TabPFN-Rel RelICL item-ltv – / 32.5 – / 40.5 48.56 / – 47.80 / – 43.97 / 28.6 user-ltv – / 36.9 – / 18.5 14.54 / – 14.40 / – 14.78 / 21.0 ad-ctr – / 4.5 – / 4.9 0.03 / – 0.03 / – 0.04 / 18.3 user-attendance – – – 0.24 / – 0.24 / 5.1 driver-position – / 52.4 – / 55.5 – 3.76 / – 3.68 / 25.2 item-sales – / 14.0 – / 20.0 0.06 / – 0.06 / – 0.04 / 69.2 post-votes – / 33.9 – / 25.5 0.07 / – 0.07 / – 0.06 / 25.6 site-success – / 4.5 – / 38.6 0.42 / – 0.39 / – 0.40 / -22.0 study-adverse – / 2.6 – / 1.6 43.91 / – 42.60 / – 32.42 / 54.0 | Pretrained (tab.) RelICL (TabPFN) 43.41 / 29.2 15.48 / 16.5 0 0.05 / 0.1 0 0.30 / 17.2 0 4.00 / 13.4 0 0.04 / 50.1 0 0.08 / 17.0 0 0.43 / -7.1 29.97 / 56.9 |
| – Not reported. | |
| Table 7: Classification (AUROC (%), higher is better). The table gives the per-task numbers behind the classification part of Tab. 4 . Bold marks the best value per task. | Table 8: RelICL with the TabPFN backbone. |
| Pretrained (rel.) Pretrained (tab.) Task RT PluRel RDBLearn TabPFN-Rel RelICL item-churn 70.9 72.5 82.1 82.8 78.5 user-churn (ama) 64.0 65.0 67.6 70.2 65.4 user-clicks 59.5 47.9 69.0 61.5 64.3 user-visits 61.8 63.4 65.5 66.9 66.2 user-ignore – – – 70.1 88.5 user-repeat – – – 76.9 79.3 driver-dnf 81.2 81.0 – 71.5 73.6 driver-top3 89.3 88.4 – 79.3 81.4 user-churn (hm) 62.8 66.0 68.0 70.6 68.4 user-badge 80.1 82.0 85.3 86.4 86.3 user-engagement 75.7 86.2 89.4 90.6 90.0 study-outcome 51.8 51.8 71.6 73.1 73.4 | Pretrained (tab.) RelICL (TabPFN) 78.4 64.6 58.7 65.7 87.9 77.7 74.2 78.4 67.8 84.9 89.5 73.6 |
| – Not reported. | |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | OpenML ID | Rows | Features | Positive class (%) |
|---|---|---|---|---|
| bioresponse | 4134 | 3751 | 1776 | 54.2 |
| churn | 40701 | 5000 | 20 | 14.1 |
| compass | 42193 | 5278 | 13 | 47.0 |
| credit-g | 31 | 1000 | 20 | 70.0 |
| jasmine | 41143 | 2984 | 144 | 50.0 |
| kr-vs-kp | 3 | 3196 | 36 | 52.2 |
| Dataset | AUROC | AUROC | AUROC | ||
|---|---|---|---|---|---|
| bioresponse | 87.7 1.0 | 87.5 0.8 | -0.3 0.3 | 86.1 1.0 | -1.6 0.4 |
| churn | 92.8 0.8 | 92.6 0.6 | -0.2 0.2 | 91.4 0.6 | -1.4 0.2 |
| compass | 72.3 0.7 | 72.4 0.7 | +0.1 0.1 | 72.3 0.7 | 0.0 0.1 |
| credit-g | 82.2 1.8 | 83.1 0.4 | +0.9 1.4 | 79.4 0.4 | -2.8 1.5 |
| jasmine | 89.8 1.1 | 86.8 0.4 | -3.0 0.6 | 86.5 0.8 | -3.3 0.4 |
| Dataset | Tables | Rows | Columns |
|---|---|---|---|
| rel-amazon | 3 | 23,218,245 | 15 |
| rel-avito | 8 | 24,653,915 | 42 |
| rel-event | 5 | 44,822,992 | 131 |
| rel-f1 | 9 | 97,606 | 67 |
| rel-hm | 3 | 16,931,173 | 37 |
| rel-stack | 7 | 5,399,818 | 51 |
| Dataset | Task | Train | Val | Test | Pos. rate (%) |
|---|---|---|---|---|---|
| rel-amazon | item-churn | 2,536,014 | 177,689 | 166,842 | 36.9 |
| rel-amazon | user-churn | 4,708,383 | 409,792 | 351,885 | 60.6 |
| rel-avito | user-clicks | 59,454 | 21,183 | 47,996 | 1.5 |
| rel-avito | user-visits | 86,619 | 29,979 | 36,129 | 85.1 |
| rel-event | user-ignore | 19,239 | 2,013 | 1,958 | 13.0 |
| rel-event | user-repeat | 3,842 | 268 | 246 | 44.7 |
| Dataset | Task | Train | Val | Test | Target (mean SD) | |
|---|---|---|---|---|---|---|
| rel-amazon | item-ltv | 2,707,679 | 166,978 | 178,334 | 77.13 | 664.61 |
| rel-amazon | user-ltv | 4,708,383 | 409,792 | 351,885 | 16.78 | 55.44 |
| rel-avito | ad-ctr | 5,100 | 1,766 | 1,816 | 0.05 | 0.11 |
| rel-event | user-attendance | 19,239 | 2,013 | 1,958 | 0.26 | 0.64 |
| rel-f1 | driver-position | 7,453 | 499 | 760 | 11.93 | 5.21 |
| rel-hm | item-sales | 5,488,184 | 105,542 | 105,542 | 0.08 | 0.58 |
| Regression | Classification |
|---|---|
| rel-f1/driver-position | rel-f1/driver-top3 |
| rel-event/user-attendance | rel-event/user-repeat |
| rel-trial/study-adverse | rel-event/user-ignore |
| rel-trial/site-success | rel-trial/study-outcome |
| rel-avito/ad-ctr | rel-avito/user-clicks |
| rel-hm/item-sales | rel-hm/user-churn |
| Hyperparameter | Search space | Selected (default values) |
|---|---|---|
| Anchor budget fraction a | ||
| Train sample size | ||
| Context table sample size | ||
| Count features after optimization b | true | |
| Count feature windows (days) b | none | |
| Relative time encoding | true |
| Kind | Model | item-ltv | user-ltv | ad-ctr | user-attendance | driver-position | item-sales | post-votes | site-success | study-adverse | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Trained | GraphSAGE | 50.05 / | – | 14.31 / | – | 0.04 / | – | 0.26 / | – | 4.02 / | – | 0.06 / | – | 0.07 / | – | 0.40 / | – | 44.47 / | – |
| Trained | RelGNN | 48.77 / | – | 14.23 / | – | 0.04 / | – | 0.24 / | – | 3.80 / | – | 0.05 / | – | 0.07 / | – | 0.30 / | – | 44.46 / | – |
| Trained | RelGT | 48.92 / | – | 14.27 / | – | 0.04 / | – | 0.25 / | – | 3.92 / | – | 0.05 / | – | 0.07 / | – | 0.33 / | – | 43.99 / | – |
| Trained | DS+LightGBM | 41.12 / | – | 13.93 / | – | 0.04 / | – | 0.28 / | – | 3.96 / | – | 0.04 / | – | 0.07 / | – | 0.41 / | – | 40.58 / | – |
| Pretrained (rel.) | RT | – / | 32.5 | – / | 36.9 | – / | 4.5 | – | – / | 52.4 | – / | 14.0 | – / | 33.9 | – / | 4.5 | – / | 2.6 | |
| Pretrained (rel.) | RT (target DB) | – / | 32.2 | – / | 38.3 | – / | 8.0 | – | – / | 58.7 | – / | 30.9 | – / | 35.0 | – / | 5.1 | – / | 3.1 | |
| Kind | Model | item-churn | user-churn (ama) | user-clicks | user-visits | user-ignore | user-repeat | driver-dnf | driver-top3 | user-churn (hm) | user-badge | user-engagement | study-outcome |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Trained | GraphSAGE | 82.8 | 70.4 | 65.9 | 66.2 | 81.6 | 76.9 | 72.6 | 75.5 | 69.9 | 88.9 | 90.6 | 68.6 |
| Trained | RelGNN | 82.6 | 71.0 | 68.2 | 66.2 | 86.2 | 79.6 | 75.3 | 85.7 | 70.9 | 89.0 | 90.8 | 71.2 |
| Trained | RelGT | 82.5 | 70.4 | 68.3 | 66.8 | 81.6 | 76.1 | 75.9 | 83.5 | 69.3 | 86.3 | 90.5 | 68.6 |
| Trained | DS+LightGBM | 81.8 | 67.6 | 64.3 | 64.5 | 84.2 | 75.4 | 69.8 | 82.4 | 69.0 | 86.2 | 90.3 | 72.0 |
| Pretrained (rel.) | RT | 70.9 | 64.0 | 59.5 | 61.8 | – | – | 81.2 | 89.3 | 62.8 | 80.1 | 75.7 | 51.8 |
| Pretrained (rel.) | PluRel | 72.5 | 65.0 | 47.9 | 63.4 | – | – | 81.0 | 88.4 | 66.0 | 82.0 | 86.2 | 51.8 |
| Pretrained (rel.) | Pretrained (tab.) | |||||
|---|---|---|---|---|---|---|
| Kind | Model | RT | PluRel | RDBLearn | TabPFN-Rel | RelICL |
| Pretrained (rel.) | RT | – | 3 / 8 ∗ | – | – | 5 / 8 ∗ |
| Pretrained (rel.) | PluRel | 5 / 8 ∗ | – | – | – | 3 / 8 ∗ |
| Pretrained (tab.) | RDBLearn | – | – | – | 0 / 7 | 2 / 7 |
| Pretrained (tab.) | TabPFN-Rel | – | – | 6 / 7 | – | 4 / 9 |
| Pretrained (tab.) | RelICL | 3 / 8 ∗ | 5 / 8 ∗ | 5 / 7 | 5 / 9 | – |
| Pretrained (rel.) | Pretrained (tab.) | |||||
|---|---|---|---|---|---|---|
| Kind | Model | RT | PluRel | RDBLearn | TabPFN-Rel | RelICL |
| Pretrained (rel.) | RT | – | 0 3 / 10 | 0 0 / 8 | 0 2 / 10 | 0 2 / 10 |
| Pretrained (rel.) | PluRel | 0 6 / 10 | – | 0 0 / 8 | 0 2 / 10 | 0 2 / 10 |
| Pretrained (tab.) | RDBLearn | 0 8 / 8 0 | 0 8 / 8 0 | – | 0 1 / 8 | 0 3 / 8 |
| Pretrained (tab.) | TabPFN-Rel | 0 8 / 10 | 0 8 / 10 | 0 7 / 8 0 | – | 0 6 / 12 |
| Pretrained (tab.) | RelICL | 0 8 / 10 | 0 8 / 10 | 0 5 / 8 0 | 0 6 / 12 | – |
| Trained | Pretrained (rel.) | Pretrained (tab.) | |||||||||||
| Kind | Model | GraphSAGE | RelGNN | RelGT | DS+LightGBM | RT | RT (target DB) | PluRel | KumoRFM-2 | RDBLearn | TabPFN-Rel (API) | TabPFN-Rel | RelICL |
| Trained | GraphSAGE | – | 0 0 / 9 | 0 0 / 9 | 0 3 / 9 | – | – | – | 0 1 / 9 | 0 4 / 7 0 | 0 4 / 9 | 0 3 / 9 | 0 1 / 9 |
| Trained | RelGNN | 0 8 / 9 0 | – | 0 5 / 9 0 | 0 4 / 9 0 | – | – | – | 0 1 / 9 | 0 4 / 7 0 | 0 5 / 9 0 | 0 5 / 9 0 | 0 4 / 9 |
| Trained | RelGT | 0 8 / 9 0 | 0 2 / 9 | – | 0 4 / 9 0 | – | – | – | 0 1 / 9 | 0 4 / 7 0 | 0 4 / 9 | 0 4 / 9 | 0 3 / 9 |
| Trained | DS+LightGBM | 0 5 / 9 0 | 0 4 / 9 0 | 0 4 / 9 0 | – | – | – | – | 0 3 / 9 | 0 6 / 7 0 | 0 5 / 9 0 | 0 5 / 9 0 | 0 3 / 9 |
| Pretrained (rel.) | RT | – | – | – | – | – | 0 1 / 8 0 ∗ | 0 3 / 8 0 ∗ | – | – | – | – | 0 5 / 8 0 ∗ |
| Trained | Pretrained (rel.) | Pretrained (tab.) | ||||||||||
| Kind | Model | GraphSAGE | RelGNN | RelGT | DS+LightGBM | RT | PluRel | KumoRFM-2 | RDBLearn | TabPFN-Rel (API) | TabPFN-Rel | RelICL |
| Trained | GraphSAGE | – | 0 2 / 12 | 0 7 / 12 | 0 9 / 12 | 0 8 / 10 | 0 8 / 10 | 0 5 / 12 | 0 6 / 8 0 | 0 3 / 12 | 0 7 / 12 | 0 6 / 12 |
| Trained | RelGNN | 10 / 12 | – | 0 9 / 12 | 11 / 12 | 0 8 / 10 | 0 8 / 10 | 0 6 / 12 | 0 6 / 8 0 | 0 8 / 12 | 0 9 / 12 | 0 9 / 12 |
| Trained | RelGT | 0 5 / 12 | 0 3 / 12 | – | 10 / 12 | 0 8 / 10 | 0 8 / 10 | 0 4 / 12 | 0 6 / 8 0 | 0 4 / 12 | 0 5 / 12 | 0 8 / 12 |
| Trained | DS+LightGBM | 0 3 / 12 | 0 1 / 12 | 0 2 / 12 | – | 0 8 / 10 | 0 8 / 10 | 0 1 / 12 | 0 5 / 8 0 | 0 1 / 12 | 0 3 / 12 | 0 5 / 12 |
| Pretrained (rel.) | RT | 0 2 / 10 | 0 2 / 10 | 0 2 / 10 | 0 2 / 10 | – | 0 3 / 10 | 0 0 / 10 | 0 0 / 8 | 0 2 / 10 | 0 2 / 10 | 0 2 / 10 |