Relational in-context learning (ICL) uses labeled support examples and their linked relational context to predict labels for new queries. This creates a failure mode when target-derived features are present in the support context but unavailable for the query. We study this setting as support-set target leakage. We construct 20 controlled target-derived features that vary in signal fidelity, representation, semantic transparency, coverage, and zero-, one-, and two-hop relational placement, and evaluate them across 13 RelBench tasks and five relational ICL configurations that vary the ICL head, message-passing depth, pretraining cohort, or relational encoder architecture. We evaluate matched 0-hop, 1-hop, and 2-hop leakage settings, together with a Full leakage condition containing all 20 leaker columns. Within the tested configurations, target-table (0-hop) and Full leakage produce the largest aggregate deviations from clean evaluation, while higher-hop effects are often weaker, consistent with differences in effective exposure associated with temporal reachability, sampling, and aggregation fidelity. Leakage effects are strongly task- and model-dependent and can reverse relative conclusions between model variants even when aggregate changes are small. For leaker detection, we compare an Integrated Gradients (IG)-based screening method with mutual information (MI) and leave-one-column-out (LOCO) on a common Baseline subset. Ranking quality is strongest in the high-impact 0-hop and Full leakage conditions, but detector-based removal does not consistently restore the clean evaluation. A four-task rel-salt case study further shows the same evaluation concern with native-schema leakage candidates from the original relational schema. These results identify the support/query information boundary as an important component of reliable relational ICL evaluation.
Figures & tables
Figure 1: Support-set leakage is structured by task, model, relational exposure, and leaker type. (A) Signed macro-F1 change under leaky support/clean query across 13 tasks, five configurations, and 0-hop, 1-hop, 2-hop, and Full leakage. (B1) Median ∣ΔF1∣ by model and hop; (B2) by dataset and hop. (C) Baseline IG leaker-family q90 detection and associated recovery: det is the fraction of evaluated true-leaker columns flagged; n is the number of runs with at least one family member flagged; and success is the fraction of those runs with valid recovery satisfying ∣F1post−F1clean∣+10−6<∣F1leaky−F1clean∣ . Multiple families may be removed together, so these are associations rather than causal family effects. Full task-, leaker-, and detector-level results are provided in Appendix D , E , F .
Figure 2: Support-set leakage can change model-selection conclusions and is not reliably repaired by automatic column removal. (A) Variant-versus-Baseline macro-F1 advantages under clean and support-only leakage; axis crossings indicate changes in model ordering. (B) Strict ranking reversals among non-tied comparisons, reported as count/denominator and percentage. (C1) Mean AUPRC for IG, MI, and LOCO on the common Baseline detector subset (C2) Mean Baseline macro-F1 before leakage, under leakage, and after detector-based removal. MI and LOCO are restricted to this subset because of their computational cost. (C3) Distance from clean before versus after removal; points below the diagonal move toward the clean evaluation. (D) rel-salt case study averaged over four Baseline tasks, comparing clean, target-field-contaminated, and post-removal performance.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Overview of the evaluation and detection–mitigation protocol. The primary support-set leakage condition compares (Sleaky,Qclean) with the clean evaluation (Sclean,Qclean) , while (Sleaky,Qleaky) is used only as a matched-availability diagnostic. For detection, Ridge-IG provides a screening signal; columns flagged at the q90 operating point are removed from both support and query, relational embeddings are recomputed, and the corresponding ICL configuration is reevaluated.
Figure 4: Support-set target leakage protocol and controlled stress test. (A) Evaluation conditions: clean support/clean query, leaky support/clean query (the primary availability-mismatch setting), and leaky support/leaky query (diagnostic control). (B) The 20 synthetic constructions vary relational distance, modality, semantic transparency, signal fidelity, and coverage; the full catalog is provided in Appendix C . (C) Median macro-F1 change relative to clean under leaky support / clean query and matched-availability (leaky support / leaky query) conditions.
Setting
Encoder
MP depth
Pretraining
ICL head
Change from Baseline
Baseline
Griffin
4
CTU–RelBench
TabPFN
Baseline reference configuration
MP-8
Griffin
8
CTU–RelBench
TabPFN
Message-passing depth
PluRel
Griffin
4
PluRel
TabPFN
Pretraining cohort
TabICL
Griffin
4
CTU–RelBench
TabICLv2
ICL head
Bimodal
Bimodal
4
CTU–RelBench
TabPFN
Relational encoder
Appendix
Table 1: Model configurations. Each variant changes the indicated component relative to the Baseline while retaining the remaining evaluation configuration.
Table 2: CTU–RelBench pretraining cohort. Databases used to pretrain the CTU–RelBench Griffin checkpoints. The four RelBench databases used for downstream evaluation were held out from this pretraining cohort.
Hop
Leaker type
Encoding
Name sem.
Value sem.
0
Perfect deterministic
Numerical
No
No
0
Noisy 80%
Numerical
No
No
0
Rank/ordinal
Numerical
No
No
0
Jittered numerical proxy
Numerical
No
No
0
Semantic name, opaque values
Text
Yes
No
0
Non-semantic name, semantic values
Text
No
Yes
Appendix
Table 3: Synthetic leaker catalog. Name semantics indicates whether the column name identifies its relationship to the target; value semantics indicates whether the stored representation has human-interpretable class meaning.
Task
First-hop exposure
Relevant task/proxy diagnostic
paper-citation
∼ 16% train; ∼ 68% test
repeated paper labels
searchinfo-isuserloggedon
∼ 0% in audit
path rarely available
searchstream-click
0%
first edge unavailable
user-clicks
∼ 60% test
1.54% positive in RelBench test
user-visits
∼ 69% test
85.06% positive in RelBench test
event-interest tasks
∼ 4%
event often occurs later
Appendix
Table 4: Selected task diagnostics underlying the higher-hop analysis. Reachability denotes the measured availability of the selected first-hop path under the task timestamps and sampling protocol; values are approximate where reported from the empirical path audit. The class-balance statistics for the Avito tasks refer to the corresponding RelBench test targets. The study-outcome 2-hop entry reports agreement between the rounded aggregate proxy and the queried label.
Method
Condition
AUPRC
AUROC
Random AUPRC
IG
0-hop
0.596
0.906
0.052
IG
1-hop
0.128
0.568
0.052
IG
2-hop
0.051
0.282
0.052
IG
Full
0.612
0.763
0.245
MI
0-hop
0.419
0.689
0.052
MI
1-hop
0.120
0.574
0.052
Appendix
Table 5: Detector ranking on the common Baseline subset. All three methods are evaluated on the same task–condition and candidate-column sets within each leakage condition. Random AUPRC is the mean leaker prevalence in these candidate sets.
Condition
Rule
Precision
Recall
F1
Hit rate
0-hop
q80
0.291
0.863
0.403
0.882
q90
0.357
0.716
0.430
0.882
q95
0.469
0.627
0.489
0.882
Bonferroni-style ( α=0.05 )
0.199
0.431
0.224
0.471
BH-style ( q=0.05 )
0.195
0.490
0.244
0.529
1-hop
q80
0.032
0.216
0.053
0.294
Appendix
Table 6: IG threshold sensitivity. Mean thresholded detection metrics across runs under alternative operating rules. Hit rate denotes the fraction of runs in which at least one active leaker is flagged. Bonferroni-style and BH-style thresholds are included as alternative score-based operating rules and are not interpreted as providing formal family-wise error or false-discovery-rate control.
Condition
k
Precision@ k
Recall@ k
Hit rate@ k
0-hop
1
0.647
0.206
0.647
3
0.510
0.520
0.882
5
0.400
0.667
0.882
1-hop
1
0.000
0.000
0.000
3
0.020
0.020
0.059
5
0.035
0.059
0.118
Appendix
Table 7: Top- k IG verification. Mean verification metrics across runs when only the k∈{1,3,5} highest-ranked columns are inspected. Precision@ k is the fraction of inspected columns that are active leakers, Recall@ k is the fraction of active leakers recovered, and Hit rate@ k is the fraction of runs in which at least one active leaker appears among the top- k columns.
Model
Task / condition
TP
FP
FN
Interpretation
TabICL
user-repeat / 0-hop
3
10
0
original-feature ablation
PluRel
study-outcome / Full
6
8
14
residual + ablation
TabICL
eligibilities-child / Full
11
3
9
residual + ablation
Appendix
Table 8: Representative mitigation failure modes. Counts refer to active synthetic leakers and original-column false positives removed by the detector.
Figure 5: Individual-leaker effects under support-only leakage. Signed change in clean-query macro-F1 when one synthetic leaker at a time is added to the Baseline support context. Rows denote the 13 RelBench tasks and columns denote the 20 synthetic constructions, grouped by relational placement. Negative and positive values indicate degradation and improvement, respectively, relative to the corresponding clean Baseline evaluation.
Model
Dataset
Task
Clean F1
0-hop delta F1
1-hop delta F1
2-hop delta F1
Full delta F1
Baseline
rel-arxiv
paper-citation
0.6719
-0.0729
0.0000
-0.0103
-0.1828
Baseline
rel-avito
searchinfo-isuserloggedon
0.4174
0.0146
-0.0070
-0.0070
-0.1711
Baseline
rel-avito
searchstream-click
1.0000
-0.6452
0.0000
0.0000
-0.7987
Baseline
rel-avito
user-clicks
0.4960
0.0000
0.0000
0.0000
0.0000
Baseline
rel-avito
user-visits
0.4698
0.0000
0.0000
0.0000
0.0000
Baseline
rel-event
event_interest-interested
0.4370
-0.2209
0.0000
0.0000
-0.0445
Appendix
Table 9: Complete task-level support-only leakage effects. Each Δ F1 is leaky-support/clean-query macro-F1 minus the corresponding clean macro-F1.
Table 11: Task-level strict ranking reversals. The two difference columns are variant minus Baseline macro-F1 under the corresponding evaluation condition.
Minimum margin ϵ
Reversals retained
0
32
0.005
24
0.01
20
0.02
12
Appendix
Table 12: Ranking-reversal margin sensitivity. Number of strict reversals retained when both the clean and support-only variant–Baseline performance margins must exceed ϵ in absolute value.
Task
Δ Included
Δ Post
TP
FP
item-incoterms
-0.008
-0.008
2
1
item-plant
-0.033
-0.010
1
2
sales-incoterms
-0.235
-0.252
2
1
sales-office
0.000
0.000
3
0
Mean
-0.069
-0.068
–
–
Appendix
Table 13: rel-salt target-field case study. Δ Included and Δ Post report macro-F1 change relative to this task-specific reference when the other seven target-designated fields are included and after detector-based removal, respectively. TP counts detector-flagged target-designated candidate fields, while FP counts flagged other original features; these labels are defined relative to the candidate set used in this case study and do not imply that every target-designated field is a confirmed deployment-time leaker.
Relational in-context learning (ICL) conditions predictions on the labeled support examples and their linked tables, creating a failure mode when the support set contains target-derived features that are unavailable for the query. We formulate this problem as support-set target leakage, distinct from leakage during dataset construction, temporal splitting, or representation learning. Here, the target-derived (leaker) columns are present only in the labeled support set during relational in-context inference, while queries remain clean. We construct 14 synthetic leaker types, corresponding to 20 columns, spanning proxies with different noise levels, coverage, modalities, semantic transparency, and relational distances. We evaluate a frozen relational encoder with an ICL head on held-out RelBench databases and use Integrated Gradients (IG) to rank and remove suspicious columns. Our results show that the effect of support-set leakage varies across tasks and relational distances. Target-table leakers cause the clearest degradation, while one- and two-hop leakers are not consistently used by the model. IG ranks target-table leakers highly across datasets and partially recovers performance in settings where leakage has the largest effect.
Roshan Reddy Upendra, Alexandre Dorais, Joe Meyer +5
Tabular foundation models achieve state-of-the-art performance on single-table tasks without any training. Recent work suggests that they are also well-suited for relational learning via deep feature synthesis (DFS), which flattens a relational schema into a single table by adding aggregates of the other tables' columns as features. This approach is appealing because it directly benefits from improvements to or customization of the underlying tabular foundation model. In this paper, we identify two key problems with DFS: feature explosion and interaction blindness. The first problem arises because the number of DFS features grows quickly as the schema becomes more complex, limiting scalability and performance. The second problem arises because column-wise aggregates do not account for feature interactions, limiting performance. We propose and explore an alternative method termed RelICL, which keeps the benefits of DFS but alleviates these two problems. At its heart, RelICL propagates and fuses information step by step through the schema graph, using the same tabular foundation model that is eventually used for prediction to do so. In our experimental study using RelBench tasks, RelICL was on par with the strongest approach based on deep feature synthesis.
Simon Forbat, Rainer Gemulla
Data and Web Science Group University of Mannheim Mannheim, Germany
Relational Foundation Models (RFMs) promise a single pre-trained predictor that, given any relational database, returns predictions in one forward pass via relational in-context learning (ICL). Yet a substantial gap separates open RFMs from their commercial counterparts, and the origin of this gap has not been systematically understood. We dissect a representative framework, the Relational Transformer (RT), from two perspectives. Model side: we show that RT performs relation-level ICL, and a kernel regression view shows it fails when sparse label-cell coverage yields an underdetermined regression. Data side: we ablate RT's pre-training source and find that existing synthetic-only pre-training and in-distribution pre-training drive the same architecture into different regimes, lazy vs. feature-learning. Probing this gap reveals that the missing ingredient is a support-identifiable relational latent in the label-generation process. These two diagnoses translate into (1) a dual-stage ICL architecture that combines the relational backbone with a batch-level ICL layer lifted from a pre-trained tabular foundation model to overcome relation-level label scarcity, and (2) a homophily-aware synthetic plus continual real-data pre-training mixture, augmented with a prototype-based regularization. These choices define OpenRFM, a simple yet effective RFM that improves average task performance by approximately 30% over the RT backbone and surpasses the commercial model KumoRFMv1 on a large set of evaluation tasks.
Zhikai Chen, Junyu Yin, Jialiang Gu +5
Michigan State University · George Mason University · Georgia Institute of Technology +1