Relational foundation models (RFMs) are pretrained once on a collection of relational databases and prediction tasks, and then applied zero-shot to previously unseen databases and tasks. To make a prediction for a target row, an RFM samples a neighborhood of rows linked to that row through foreign keys and uses this neighborhood as its inference context. Lowering inference cost is an important goal for any foundation model, and for RFMs this cost grows with the size of the context. The simplest ways to shrink the context is to drop some of the sampled rows, but this ignores the semantics of the database schema, so it is as likely to discard informative rows as uninformative ones. We propose STEER, a sampling approach that shrinks the inference context by concentrating it on the tables most relevant to the prediction task at hand. STEER obtains relevance information by prompting a large language model to rank the foreign-key edges of the database schema into relevance tiers for the given task, and then maps each tier to a probability of following that edge during traversal. Because the ranking uses only the schema, it is computed once per task and reused across all subsequent predictions, amortizing its cost. We evaluate STEER on three state-of-the-art RFMs (RT, RT-J, and Griffin) and show that it reduces inference context size by about 40% on average while maintaining, and in some cases improving, prediction accuracy.
Figures & tables
Figure 1. RFM inference on a relational database. (1) The schema, with primary (PK) and foreign keys (FK), and a prediction task defined in English, in PQL, and as a task table whose foreign key references the entity being predicted for. (2) Starting from the task row (seed), the sampler follows F→P and P→F links to collect related rows (node colors match the tables in (1)); rows dated after the prediction timestamp t are excluded. (3) The sampled rows form the context given to the RFM, which outputs the prediction. A three-part figure. Part 1 shows a relational schema for Formula 1 with primary and foreign key links into a driver table, and the prediction task written in English, in PQL, and as a task table. Part 2 shows a graph of sampled rows expanding from the seed row, with rows dated after time t crossed out. Part 3 shows the sampled rows forming a context that is passed to a relational foundation model, which outputs a prediction of 1 or 0.
Figure 2. The STEER pipeline. A prediction request is a task table whose rows are prediction targets. STEER looks up the configuration of the task in the Configs Storage. If the storage returns None , the Configs Generator generates a configuration with two LLM calls, one per traversal direction, that rank the schema edges into four relevance tiers. STEER saves this configuration in the Configs Storage, so that future invocations of this task will not result in LLM calls. The STEER Sampler then produces a sample using filtered traversal based on the configuration ( Algorithm 1 ). The resulting sample is given to the RFM, which uses it for inference and outputs a prediction. A pipeline diagram. On the left, a task table for the driver-dnf task has several rows, each marked with a question mark. An arrow labeled with a lookup of the task name leads to a Configs Storage. If the storage returns None, an arrow leads down to a Configs Generator, which contains two LLM calls, one for each traversal direction, that produce a tier for each schema edge, and the configuration returns to the storage. A configuration is passed to the STEER Sampler, which sends a sample to the RFM, which outputs a prediction of 1 or 0. An expanded view of the STEER Sampler shows, for a visited row, candidate edges with a number line from 0 to 1 each: an edge is followed when the random number is below its probability and rejected otherwise.
Figure 3. Classification results: average sample size and test accuracy (AUROC). Two side-by-side bar charts comparing Baseline, Naive Pruning, STEER with Claude, and STEER with GPT across RT, RT-J, and Griffin. The left panel shows average sample size in tokens and the right panel shows average AUROC.
Figure 4. Regression results: average sample size and test accuracy ( R2 ). Two side-by-side bar charts comparing Baseline, Naive Pruning, STEER with Claude, and STEER with GPT across RT, RT-J, and Griffin. The left panel shows average sample size in tokens and the right panel shows average R-squared.
Figure 5. Classification results for direction of traversal ablation: average sample size and test accuracy (AUROC). Two side-by-side bar charts comparing Baseline, P-to-F only, F-to-P only, and STEER across RT and RT-J. The left panel shows average sample size in tokens and the right panel shows average AUROC.
Figure 6. Regression results for direction of traversal ablation: average sample size and test accuracy ( R2 ). Two side-by-side bar charts comparing Baseline, P-to-F only, F-to-P only, and STEER across RT and RT-J. The left panel shows average sample size in tokens and the right panel shows average R-squared.
RFM
Base.
Claude High
Claude Default
GPT High
GPT Default
Simple
STEER
Simple
STEER
Simple
STEER
Simple
STEER
Classification (AUROC, %)
RT
69.98
71.21
71.46
70.58
71.30
69.64
71.44
70.04
71.27
RT-J
71.72
70.83
71.64
71.34
71.52
70.60
72.37
70.77
71.92
Regression ( R2 , %)
RT
17.33
17.59
21.21
18.05
21.51
16.35
20.31
17.14
20.60
Table 1. Prompting strategy ablation: average test accuracy for classification and regression, comparing simple prompting with STEER’s smart prompting. Results are shown for Claude and GPT with high and default reasoning effort. Within each simple/STEER pair, bold indicates the higher accuracy.
RFM
Baseline
STEER
STEER + Col. Filter
Test ↑
Tok. ↓
Test ↑
Tok. ↓
Test ↑
Tok. ↓
Classification (AUROC, %)
RT
69.98
915.6
71.46
546.6
70.82
467.4
RT-J
71.72
912.6
71.64
517.7
72.20
450.6
Regression ( R2 , %)
RT
17.33
954.5
21.21
489.1
26.60
409.2
Table 2. Column filter ablation: average test accuracy and sample size for classification and regression, comparing baseline sampling, STEER without column filter, and STEER with column filter. Bold indicates the highest test score in a row and the highlighted cell indicates the lowest token count.
Database
Task
RFM
Baseline
STEER
STEER + Prob. Tuning
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-churn
RT
68.20
1023.2
71.06
230.2
71.06
230.2
RT-J
72.26
1018.9
74.21
118.7
74.21
118.7
rel-amazon
user-churn
RT
64.14
1024.0
65.29
82.5
64.55
71.2
RT-J
65.45
1021.3
65.53
68.1
65.13
58.6
rel-arxiv
paper-citation
RT
79.85
989.6
79.99
863.2
79.75
919.2
Table 3. Probability tuning ablation: classification tasks (AUROC %, higher is better). For each task and model, we report test accuracy and number of tokens with baseline sampling, STEER with the default probability profile, and STEER with a probability profile tuned for the specific task. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
STEER
STEER + Prob. Tuning
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-ltv
RT
4.59
957.2
4.26
203.5
36.05
174.3
RT-J
59.31
951.5
60.17
108.4
60.17
108.4
rel-amazon
user-ltv
RT
15.20
1023.2
26.26
100.8
25.73
70.3
RT-J
61.52
1021.3
57.19
86.9
54.41
58.6
rel-arxiv
author-publication
RT
22.77
740.5
24.75
491.4
22.57
46.9
Table 4. Probability tuning ablation: regression tasks ( R2 %, higher is better). For each task and model, we report test accuracy and number of tokens with baseline sampling, STEER with the default probability profile, and STEER with a probability profile tuned for the specific task. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Database
Task
RFM
Baseline
Naive Pruning
STEER (Claude)
STEER (GPT)
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-churn
RT
68.20
1023.2
63.90
511.6
71.06
230.2
71.07
236.4
RT-J
72.26
1018.9
66.58
509.9
74.21
118.7
74.41
122.6
Griffin
69.04
10885.8
65.64
5523.0
69.10
8601.3
69.07
9057.0
rel-amazon
user-churn
RT
64.14
1024.0
60.39
512.0
65.29
82.5
64.55
71.2
RT-J
65.45
1021.3
62.09
511.2
65.53
68.1
65.05
58.6
Appendix
Table 5. Main results: classification tasks (AUROC %, higher is better). Used in Figure 3 . For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), naive pruning (full-budget sampling, then half the resulting sample discarded at random), STEER with Claude-based ranking, and STEER with GPT-based ranking. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
Naive Pruning
STEER (Claude)
STEER (GPT)
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-ltv
RT
4.59
957.2
3.03
478.7
4.26
203.5
4.26
203.5
RT-J
59.31
951.5
24.29
476.2
60.17
108.4
60.17
108.4
Griffin
15.70
10175.6
13.27
5167.8
15.38
8079.4
15.56
7241.0
rel-amazon
user-ltv
RT
15.20
1023.2
3.42
511.6
26.26
100.8
25.29
101.0
RT-J
61.52
1021.3
11.26
511.1
57.19
86.9
56.01
86.7
Appendix
Table 6. Main results: regression tasks ( R2 %, higher is better). Used in Figure 4 . For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), naive pruning (full-budget sampling, then half the resulting sample discarded at random), STEER with Claude-based ranking, and STEER with GPT-based ranking. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
P→F only
F→P only
STEER
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-churn
RT
68.20
1023.2
71.25
259.9
71.06
230.2
71.06
230.2
RT-J
72.26
1018.9
74.59
137.8
74.29
118.7
74.21
118.7
rel-amazon
user-churn
RT
64.14
1024.0
65.41
124.0
63.19
738.6
65.29
82.5
RT-J
65.45
1021.3
65.30
105.5
63.50
742.2
65.53
68.1
rel-arxiv
paper-citation
RT
79.85
989.6
80.30
937.0
79.92
856.8
79.99
863.2
Appendix
Table 7. Direction of traversal ablation: classification tasks (AUROC %, higher is better). Used in Figure 5 . Ablation results use only the RT and RT-J models and the Claude LLM. For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), P→F -only traversal, F→P -only traversal, and the STEER traversal filter that combines P→F and F→P . Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
P→F only
F→P only
STEER
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-ltv
RT
4.59
957.2
4.46
954.1
4.26
203.5
4.26
203.5
RT-J
59.31
951.5
58.89
945.4
60.00
108.4
60.17
108.4
rel-amazon
user-ltv
RT
15.20
1023.2
24.13
119.5
25.99
851.9
26.26
100.8
RT-J
61.52
1021.3
53.86
105.5
53.66
602.3
57.19
86.9
rel-arxiv
author-publication
RT
22.77
740.5
24.75
491.4
22.64
680.6
24.75
491.4
Appendix
Table 8. Direction of traversal ablation: regression tasks ( R2 %, higher is better). Used in Figure 6 . Ablation results use only the RT and RT-J models and the Claude LLM. For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), P→F -only traversal, F→P -only traversal, and the STEER traversal filter that combines P→F and F→P . Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Base.
Claude High
Claude Default
GPT High
GPT Default
Simple
STEER
Simple
STEER
Simple
STEER
Simple
STEER
rel-amazon
item-churn
RT
68.20
71.25
71.06
71.25
71.07
71.06
71.07
71.06
71.06
RT-J
72.26
74.76
74.21
74.76
74.41
74.34
74.41
74.34
74.21
rel-amazon
user-churn
RT
64.14
65.41
65.29
65.41
65.29
64.55
64.55
65.41
65.29
RT-J
65.45
65.54
65.53
65.54
65.53
65.05
65.05
65.54
65.53
rel-arxiv
paper-citation
RT
79.85
79.85
79.99
79.85
79.83
74.21
79.39
74.21
79.12
Appendix
Table 9. Prompting strategy ablation: classification tasks (AUROC %, higher is better). Used in Table 1 . Ablation results use only the RT and RT-J models. For each task and model, we report test accuracy with the baseline (random sampling). We also report simple prompting and STEER’s smart prompting with four LLM variants: Claude and GPT with high and default reasoning effort. Within each simple/STEER pair, bold indicates the higher accuracy (ties favor STEER). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference).
Database
Task
RFM
Base.
Claude High
Claude Default
GPT High
GPT Default
Simple
STEER
Simple
STEER
Simple
STEER
Simple
STEER
rel-amazon
item-ltv
RT
4.59
4.00
4.26
4.00
4.26
4.00
4.26
4.26
4.22
RT-J
59.31
59.05
60.17
59.05
60.17
59.05
60.17
60.17
59.84
rel-amazon
user-ltv
RT
15.20
24.13
26.26
24.13
25.29
25.73
25.29
25.73
23.90
RT-J
61.52
54.60
57.19
54.60
56.01
54.41
56.01
54.41
55.79
rel-arxiv
author-publication
RT
22.77
22.99
24.75
22.80
24.30
23.14
23.99
22.57
23.15
Appendix
Table 10. Prompting strategy ablation: regression tasks ( R2 %, higher is better). Used in Table 1 . Ablation results use only the RT and RT-J models. For each task and model, we report test accuracy with the baseline (random sampling). We also report simple prompting and STEER’s smart prompting with four LLM variants: Claude and GPT with high and default reasoning effort. Within each simple/STEER pair, bold indicates the higher accuracy (ties favor STEER). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference).
Database
Task
RFM
Baseline
STEER
STEER + Col. Filter
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-churn
RT
68.20
1023.2
71.06
230.2
69.98
91.3
RT-J
72.26
1018.9
74.21
118.7
72.67
50.8
rel-amazon
user-churn
RT
64.14
1024.0
65.29
82.5
63.75
35.5
RT-J
65.45
1021.3
65.53
68.1
65.38
30.2
rel-arxiv
paper-citation
RT
79.85
989.6
79.99
863.2
78.37
862.0
Appendix
Table 11. Column filter ablation: classification tasks (AUROC %, higher is better). Used in Table 2 . Ablation results use only the RT and RT-J models and the Claude LLM. For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), STEER without column filter, and STEER with column filter. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
STEER
STEER + Col. Filter
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-ltv
RT
4.59
957.2
4.26
203.5
33.25
42.9
RT-J
59.31
951.5
60.17
108.4
59.64
30.2
rel-amazon
user-ltv
RT
15.20
1023.2
26.26
100.8
26.95
50.1
RT-J
61.52
1021.3
57.19
86.9
22.33
41.7
rel-arxiv
author-publication
RT
22.77
740.5
24.75
491.4
11.63
480.0
Appendix
Table 12. Column filter ablation: regression tasks ( R2 %, higher is better). Used in Table 2 . Ablation results use only the RT and RT-J models and the Claude LLM. For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), STEER without column filter, and STEER with column filter. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Relational Foundation Models (RFMs) promise a single pre-trained predictor that, given any relational database, returns predictions in one forward pass via relational in-context learning (ICL). Yet a substantial gap separates open RFMs from their commercial counterparts, and the origin of this gap has not been systematically understood. We dissect a representative framework, the Relational Transformer (RT), from two perspectives. Model side: we show that RT performs relation-level ICL, and a kernel regression view shows it fails when sparse label-cell coverage yields an underdetermined regression. Data side: we ablate RT's pre-training source and find that existing synthetic-only pre-training and in-distribution pre-training drive the same architecture into different regimes, lazy vs. feature-learning. Probing this gap reveals that the missing ingredient is a support-identifiable relational latent in the label-generation process. These two diagnoses translate into (1) a dual-stage ICL architecture that combines the relational backbone with a batch-level ICL layer lifted from a pre-trained tabular foundation model to overcome relation-level label scarcity, and (2) a homophily-aware synthetic plus continual real-data pre-training mixture, augmented with a prototype-based regularization. These choices define OpenRFM, a simple yet effective RFM that improves average task performance by approximately 30% over the RT backbone and surpasses the commercial model KumoRFMv1 on a large set of evaluation tasks.
Zhikai Chen, Junyu Yin, Jialiang Gu +5
Michigan State University · George Mason University · Georgia Institute of Technology +1
Relational Foundation Models (RFMs) require large-scale synthetic relational databases for pretraining, but existing approaches tightly couple data generation with the model training pipeline. We study whether PluRel, a general-purpose synthetic relational database generator, can serve as an external data source for RDB-PFN, a relational in-context learner originally pretrained with a 600K-task single-table warm-up followed by an approximately 1.8M-task adaptation stage. We build a conversion pipeline that maps PluRel-generated databases, including externally constructed binary prediction tasks, into the RDB-PFN training format and evaluate three curriculum strategies: SCHEMA-GUIDED FIRST (real-world schema then fully synthetic), FULLY SYNTHETIC (diverse synthetic schemas throughout), and SCHEMA-GUIDED LAST (fully synthetic then real-world schema). Using only approximately 5,500 relational databases (approximately 33K tasks), roughly 55x fewer tasks than the original protocol, and no single-table warm-up, our best curriculum (SCHEMA-GUIDED FIRST) achieves 0.6346 average ROC-AUC across 19 real benchmark tasks at 1024-shot context, recovering 87.6% of the published RDB-PFN performance (0.7245). At 64-shot context, the gap narrows to 93.8% (0.6116 vs. 0.6517). Our results demonstrate that external synthetic generators can provide useful pretraining signals for RFMs when combined with appropriate curriculum design and that exposure to a real-world schema early in training is substantially more effective than late-stage schema adaptation.
Relational databases store much of the world's structured information, and they are essential for driving complex predictive applications. However, deep learning progress on relational data remains limited, as conventional approaches flatten databases into single tables via manual feature engineering, discarding relational context. Relational deep learning (RDL) addresses this by modeling databases as relational entity graphs (REGs) for graph neural networks (GNNs), but remains task- and database-specific. To combine the strengths of both paradigms, we propose a hybrid architecture combining a fine-tuned BART encoder to capture intra-row semantics with a GraphSAGE-based GNN over REGs to inject relational context. Experiments on RelBench show that the GNN substantially enriches BART's row embeddings, achieving a ROC-AUC of 67.40 on the driver-dnf task from the rel-f1 dataset. This performance is competitive with supervised baselines such as LightGBM (68.86) and narrows the gap to RDL (72.62) to within 5.22 points, though a substantial gap remains to state-of-the-art foundation models such as KumoRFM (82.63). These results suggest that lightweight hybrid LM-GNN architectures offer a promising and resource-efficient path towards foundation models for relational databases.