Relational foundation models (RFMs) are pretrained once on a collection of relational databases and prediction tasks, and then applied zero-shot to previously unseen databases and tasks. To make a prediction for a target row, an RFM samples a neighborhood of rows linked to that row through foreign keys and uses this neighborhood as its inference context. Lowering inference cost is an important goal for any foundation model, and for RFMs this cost grows with the size of the context. The simplest ways to shrink the context is to drop some of the sampled rows, but this ignores the semantics of the database schema, so it is as likely to discard informative rows as uninformative ones. We propose STEER, a sampling approach that shrinks the inference context by concentrating it on the tables most relevant to the prediction task at hand. STEER obtains relevance information by prompting a large language model to rank the foreign-key edges of the database schema into relevance tiers for the given task, and then maps each tier to a probability of following that edge during traversal. Because the ranking uses only the schema, it is computed once per task and reused across all subsequent predictions, amortizing its cost. We evaluate STEER on three state-of-the-art RFMs (RT, RT-J, and Griffin) and show that it reduces inference context size by about 40% on average while maintaining, and in some cases improving, prediction accuracy.
Figures & tables
Figure 1. RFM inference on a relational database. (1) The schema, with primary (PK) and foreign keys (FK), and a prediction task defined in English, in PQL, and as a task table whose foreign key references the entity being predicted for. (2) Starting from the task row (seed), the sampler follows F→P and P→F links to collect related rows (node colors match the tables in (1)); rows dated after the prediction timestamp t are excluded. (3) The sampled rows form the context given to the RFM, which outputs the prediction. A three-part figure. Part 1 shows a relational schema for Formula 1 with primary and foreign key links into a driver table, and the prediction task written in English, in PQL, and as a task table. Part 2 shows a graph of sampled rows expanding from the seed row, with rows dated after time t crossed out. Part 3 shows the sampled rows forming a context that is passed to a relational foundation model, which outputs a prediction of 1 or 0.
Figure 2. The STEER pipeline. A prediction request is a task table whose rows are prediction targets. STEER looks up the configuration of the task in the Configs Storage. If the storage returns None , the Configs Generator generates a configuration with two LLM calls, one per traversal direction, that rank the schema edges into four relevance tiers. STEER saves this configuration in the Configs Storage, so that future invocations of this task will not result in LLM calls. The STEER Sampler then produces a sample using filtered traversal based on the configuration ( Algorithm 1 ). The resulting sample is given to the RFM, which uses it for inference and outputs a prediction. A pipeline diagram. On the left, a task table for the driver-dnf task has several rows, each marked with a question mark. An arrow labeled with a lookup of the task name leads to a Configs Storage. If the storage returns None, an arrow leads down to a Configs Generator, which contains two LLM calls, one for each traversal direction, that produce a tier for each schema edge, and the configuration returns to the storage. A configuration is passed to the STEER Sampler, which sends a sample to the RFM, which outputs a prediction of 1 or 0. An expanded view of the STEER Sampler shows, for a visited row, candidate edges with a number line from 0 to 1 each: an edge is followed when the random number is below its probability and rejected otherwise.
Figure 3. Classification results: average sample size and test accuracy (AUROC). Two side-by-side bar charts comparing Baseline, Naive Pruning, STEER with Claude, and STEER with GPT across RT, RT-J, and Griffin. The left panel shows average sample size in tokens and the right panel shows average AUROC.
Figure 4. Regression results: average sample size and test accuracy ( R2 ). Two side-by-side bar charts comparing Baseline, Naive Pruning, STEER with Claude, and STEER with GPT across RT, RT-J, and Griffin. The left panel shows average sample size in tokens and the right panel shows average R-squared.
Figure 5. Classification results for direction of traversal ablation: average sample size and test accuracy (AUROC). Two side-by-side bar charts comparing Baseline, P-to-F only, F-to-P only, and STEER across RT and RT-J. The left panel shows average sample size in tokens and the right panel shows average AUROC.
Figure 6. Regression results for direction of traversal ablation: average sample size and test accuracy ( R2 ). Two side-by-side bar charts comparing Baseline, P-to-F only, F-to-P only, and STEER across RT and RT-J. The left panel shows average sample size in tokens and the right panel shows average R-squared.
RFM
Base.
Claude High
Claude Default
GPT High
GPT Default
Simple
STEER
Simple
STEER
Simple
STEER
Simple
STEER
Classification (AUROC, %)
RT
69.98
71.21
71.46
70.58
71.30
69.64
71.44
70.04
71.27
RT-J
71.72
70.83
71.64
71.34
71.52
70.60
72.37
70.77
71.92
Regression ( R2 , %)
RT
17.33
17.59
21.21
18.05
21.51
16.35
20.31
17.14
20.60
Table 1. Prompting strategy ablation: average test accuracy for classification and regression, comparing simple prompting with STEER’s smart prompting. Results are shown for Claude and GPT with high and default reasoning effort. Within each simple/STEER pair, bold indicates the higher accuracy.
RFM
Baseline
STEER
STEER + Col. Filter
Test ↑
Tok. ↓
Test ↑
Tok. ↓
Test ↑
Tok. ↓
Classification (AUROC, %)
RT
69.98
915.6
71.46
546.6
70.82
467.4
RT-J
71.72
912.6
71.64
517.7
72.20
450.6
Regression ( R2 , %)
RT
17.33
954.5
21.21
489.1
26.60
409.2
Table 2. Column filter ablation: average test accuracy and sample size for classification and regression, comparing baseline sampling, STEER without column filter, and STEER with column filter. Bold indicates the highest test score in a row and the highlighted cell indicates the lowest token count.
Database
Task
RFM
Baseline
STEER
STEER + Prob. Tuning
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-churn
RT
68.20
1023.2
71.06
230.2
71.06
230.2
RT-J
72.26
1018.9
74.21
118.7
74.21
118.7
rel-amazon
user-churn
RT
64.14
1024.0
65.29
82.5
64.55
71.2
RT-J
65.45
1021.3
65.53
68.1
65.13
58.6
rel-arxiv
paper-citation
RT
79.85
989.6
79.99
863.2
79.75
919.2
Table 3. Probability tuning ablation: classification tasks (AUROC %, higher is better). For each task and model, we report test accuracy and number of tokens with baseline sampling, STEER with the default probability profile, and STEER with a probability profile tuned for the specific task. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
STEER
STEER + Prob. Tuning
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-ltv
RT
4.59
957.2
4.26
203.5
36.05
174.3
RT-J
59.31
951.5
60.17
108.4
60.17
108.4
rel-amazon
user-ltv
RT
15.20
1023.2
26.26
100.8
25.73
70.3
RT-J
61.52
1021.3
57.19
86.9
54.41
58.6
rel-arxiv
author-publication
RT
22.77
740.5
24.75
491.4
22.57
46.9
Table 4. Probability tuning ablation: regression tasks ( R2 %, higher is better). For each task and model, we report test accuracy and number of tokens with baseline sampling, STEER with the default probability profile, and STEER with a probability profile tuned for the specific task. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Database
Task
RFM
Baseline
Naive Pruning
STEER (Claude)
STEER (GPT)
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-churn
RT
68.20
1023.2
63.90
511.6
71.06
230.2
71.07
236.4
RT-J
72.26
1018.9
66.58
509.9
74.21
118.7
74.41
122.6
Griffin
69.04
10885.8
65.64
5523.0
69.10
8601.3
69.07
9057.0
rel-amazon
user-churn
RT
64.14
1024.0
60.39
512.0
65.29
82.5
64.55
71.2
RT-J
65.45
1021.3
62.09
511.2
65.53
68.1
65.05
58.6
Appendix
Table 5. Main results: classification tasks (AUROC %, higher is better). Used in Figure 3 . For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), naive pruning (full-budget sampling, then half the resulting sample discarded at random), STEER with Claude-based ranking, and STEER with GPT-based ranking. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
Naive Pruning
STEER (Claude)
STEER (GPT)
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-ltv
RT
4.59
957.2
3.03
478.7
4.26
203.5
4.26
203.5
RT-J
59.31
951.5
24.29
476.2
60.17
108.4
60.17
108.4
Griffin
15.70
10175.6
13.27
5167.8
15.38
8079.4
15.56
7241.0
rel-amazon
user-ltv
RT
15.20
1023.2
3.42
511.6
26.26
100.8
25.29
101.0
RT-J
61.52
1021.3
11.26
511.1
57.19
86.9
56.01
86.7
Appendix
Table 6. Main results: regression tasks ( R2 %, higher is better). Used in Figure 4 . For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), naive pruning (full-budget sampling, then half the resulting sample discarded at random), STEER with Claude-based ranking, and STEER with GPT-based ranking. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
P→F only
F→P only
STEER
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-churn
RT
68.20
1023.2
71.25
259.9
71.06
230.2
71.06
230.2
RT-J
72.26
1018.9
74.59
137.8
74.29
118.7
74.21
118.7
rel-amazon
user-churn
RT
64.14
1024.0
65.41
124.0
63.19
738.6
65.29
82.5
RT-J
65.45
1021.3
65.30
105.5
63.50
742.2
65.53
68.1
rel-arxiv
paper-citation
RT
79.85
989.6
80.30
937.0
79.92
856.8
79.99
863.2
Appendix
Table 7. Direction of traversal ablation: classification tasks (AUROC %, higher is better). Used in Figure 5 . Ablation results use only the RT and RT-J models and the Claude LLM. For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), P→F -only traversal, F→P -only traversal, and the STEER traversal filter that combines P→F and F→P . Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
P→F only
F→P only
STEER
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-ltv
RT
4.59
957.2
4.46
954.1
4.26
203.5
4.26
203.5
RT-J
59.31
951.5
58.89
945.4
60.00
108.4
60.17
108.4
rel-amazon
user-ltv
RT
15.20
1023.2
24.13
119.5
25.99
851.9
26.26
100.8
RT-J
61.52
1021.3
53.86
105.5
53.66
602.3
57.19
86.9
rel-arxiv
author-publication
RT
22.77
740.5
24.75
491.4
22.64
680.6
24.75
491.4
Appendix
Table 8. Direction of traversal ablation: regression tasks ( R2 %, higher is better). Used in Figure 6 . Ablation results use only the RT and RT-J models and the Claude LLM. For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), P→F -only traversal, F→P -only traversal, and the STEER traversal filter that combines P→F and F→P . Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Base.
Claude High
Claude Default
GPT High
GPT Default
Simple
STEER
Simple
STEER
Simple
STEER
Simple
STEER
rel-amazon
item-churn
RT
68.20
71.25
71.06
71.25
71.07
71.06
71.07
71.06
71.06
RT-J
72.26
74.76
74.21
74.76
74.41
74.34
74.41
74.34
74.21
rel-amazon
user-churn
RT
64.14
65.41
65.29
65.41
65.29
64.55
64.55
65.41
65.29
RT-J
65.45
65.54
65.53
65.54
65.53
65.05
65.05
65.54
65.53
rel-arxiv
paper-citation
RT
79.85
79.85
79.99
79.85
79.83
74.21
79.39
74.21
79.12
Appendix
Table 9. Prompting strategy ablation: classification tasks (AUROC %, higher is better). Used in Table 1 . Ablation results use only the RT and RT-J models. For each task and model, we report test accuracy with the baseline (random sampling). We also report simple prompting and STEER’s smart prompting with four LLM variants: Claude and GPT with high and default reasoning effort. Within each simple/STEER pair, bold indicates the higher accuracy (ties favor STEER). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference).
Database
Task
RFM
Base.
Claude High
Claude Default
GPT High
GPT Default
Simple
STEER
Simple
STEER
Simple
STEER
Simple
STEER
rel-amazon
item-ltv
RT
4.59
4.00
4.26
4.00
4.26
4.00
4.26
4.26
4.22
RT-J
59.31
59.05
60.17
59.05
60.17
59.05
60.17
60.17
59.84
rel-amazon
user-ltv
RT
15.20
24.13
26.26
24.13
25.29
25.73
25.29
25.73
23.90
RT-J
61.52
54.60
57.19
54.60
56.01
54.41
56.01
54.41
55.79
rel-arxiv
author-publication
RT
22.77
22.99
24.75
22.80
24.30
23.14
23.99
22.57
23.15
Appendix
Table 10. Prompting strategy ablation: regression tasks ( R2 %, higher is better). Used in Table 1 . Ablation results use only the RT and RT-J models. For each task and model, we report test accuracy with the baseline (random sampling). We also report simple prompting and STEER’s smart prompting with four LLM variants: Claude and GPT with high and default reasoning effort. Within each simple/STEER pair, bold indicates the higher accuracy (ties favor STEER). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference).
Database
Task
RFM
Baseline
STEER
STEER + Col. Filter
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-churn
RT
68.20
1023.2
71.06
230.2
69.98
91.3
RT-J
72.26
1018.9
74.21
118.7
72.67
50.8
rel-amazon
user-churn
RT
64.14
1024.0
65.29
82.5
63.75
35.5
RT-J
65.45
1021.3
65.53
68.1
65.38
30.2
rel-arxiv
paper-citation
RT
79.85
989.6
79.99
863.2
78.37
862.0
Appendix
Table 11. Column filter ablation: classification tasks (AUROC %, higher is better). Used in Table 2 . Ablation results use only the RT and RT-J models and the Claude LLM. For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), STEER without column filter, and STEER with column filter. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 14 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).
Database
Task
RFM
Baseline
STEER
STEER + Col. Filter
Test ↑
Tokens ↓
Test ↑
Tokens ↓
Test ↑
Tokens ↓
rel-amazon
item-ltv
RT
4.59
957.2
4.26
203.5
33.25
42.9
RT-J
59.31
951.5
60.17
108.4
59.64
30.2
rel-amazon
user-ltv
RT
15.20
1023.2
26.26
100.8
26.95
50.1
RT-J
61.52
1021.3
57.19
86.9
22.33
41.7
rel-arxiv
author-publication
RT
22.77
740.5
24.75
491.4
11.63
480.0
Appendix
Table 12. Column filter ablation: regression tasks ( R2 %, higher is better). Used in Table 2 . Ablation results use only the RT and RT-J models and the Claude LLM. For each task and model, we report test accuracy and number of tokens with the baseline (random sampling), STEER without column filter, and STEER with column filter. Within each row, bold indicates the highest test score and the highlighted cell indicates the lowest token count (ties are broken in favor of one non-baseline condition). The Mean block at the bottom of the table reports the per-model average across all 10 tasks. The Δ block reports the change relative to the baseline (point difference for test accuracy, and percentage reduction for tokens).