Tabular foundation models (TFMs) achieve strong predictive performance through in-context learning, yet repeatedly conditioning on labeled data makes inference expensive. Knowledge distillation can reduce this cost by transferring their predictive ability to lightweight, dataset-specific students. However, the dependence of TFM predictions on both a labeled context and a query introduces two design questions: how to construct teacher supervision and whether expanding query coverage improves distillation. We examine these questions across two TFMs and both neural and tree-based students, and derive an effective distillation recipe. The recipe uses the full labeled training set as teacher context and trains students solely on teacher predictions for observed and synthetic queries. On TabArena, the resulting students outperform their supervised trained tuned-and-ensembled counterparts by 57-98 Elo points. Applied unchanged to TALENT, the same recipe improves matched default students on 236-258 of 300 datasets and reduces median primary error by 4.0-6.4%. The distilled students also achieve median inference speedups of 3.0-21.6 times over their teachers, offering a practical trade-off between predictive performance and repeated inference cost. Code is available at https://github.com/nums-ai/TFM_Distillation .
Figures & tables
Figure 1: Constructing supervision for TFM distillation. (Left) We study how to construct teacher supervision for ditillation (RQ1) and whether and how expanding teacher queries improves distillation (RQ2). We find that training solely on teacher predictions generated with the full labeled training set as context achieves the best performance. (Right) Elo versus median inference time per 1K samples on TabArena for default lightweight students, their distilled counterparts, and the TFM teachers, with Elo calibrated jointly over the same 76 configurations as in Section 4 .
Figure 2: Teacher-only supervision with the full labeled context achieves the best aggregate performance. Full-context targets reach the highest Elo at α=1 in every teacher–student pair, exceeding the best OOF–label mixture at α=0.6 .
Figure 3: Simple synthetic queries improve distillation across all four teacher–student pairs. Each of the four query generators yields higher Elo than distillation using only observed queries, with CutMix achieving the highest Elo in every pair.
Figure 4: Effect of the synthetic-query budget for CutMix . Each teacher group pairs mean Elo with teacher prediction generation time. Increasing the synthetic-to-real query ratio improves mean Elo across both teachers and student families while increasing generation time. The dashed line marks the 2:1 ratio used in our fixed recipe, chosen as a common operating point on this performance–generation-cost trade-off rather than tuned per dataset.
Model
Δ Elo ↑
KD gain (%) ↑
W/T/L
TabICL v2 → TabM
+56.7 [3.9, 107.4]
1.30 [0.61, 1.93]
33/0/18
TabICL v2 → XGBoost
+71.4 [14.3, 128.7]
1.69 [0.38, 2.21]
35/0/16
TabPFN v3 → TabM
+84.5 [34.1, 135.1]
1.38 [0.71, 2.36]
36/0/15
TabPFN v3 → XGBoost
+97.9 [41.9, 156.9]
1.91 [0.94, 3.15]
37/0/14
Table 1: KD improves each student beyond supervised tuning and ensembling. Each KD model is compared with the tuned-and-ensembled version of the same architecture on 51 TabArena datasets. KD gain is the median relative primary-error reduction, W/T/L counts dataset-level wins, ties, and losses. Elo is calibrated jointly over all 76 configurations; differences are computed before rounding. Brackets show paired dataset-bootstrap 95% intervals.
Model
Elo ↑
Improvability (%) ↓
Win rate (%) ↑
TabPFN v3 → TabM
1510.7
11.21
76.4
RealMLP (Tuned+Ensemble)
1489.9
11.09
74.5
TabICL v2 → TabM
1483.0
11.44
73.9
TabPFN v3 → XGBoost
1457.8
12.41
71.4
TabICL v2 → XGBoost
1431.3
12.62
68.8
TabM (Tuned+Ensemble)
1426.2
12.50
68.3
Table 2: Distilled students and leading supervised models on TabArena. Results cover 51 datasets and 816 dataset–split pairs using a joint pool of 72 official baseline configurations and four KD configurations. Win rate averages pairwise outcomes against the other 75 methods in this pool. Distilled students use default student configurations and r=2 , with no student hyperparameter search. Light-blue row shading identifies distilled students. Among the displayed non-teacher models, 1st , 2nd , and 3rd best results in each column are highlighted. Teachers are reported separately as references and excluded from these markings.
Figure 5: Transfer of teacher gains to distilled students across 51 TabArena datasets. Teacher and KD gains are measured as relative error reduction over the raw student, 100(1−E/Eraw) ; the dashed line denotes full recovery of the teacher gain. Annotations report Spearman correlations with 95% bootstrap CIs. Because both axes share the raw baseline, these correlations are descriptive.
Model
n
Gap recovery (%)
KD W/T/L
TabICL v2 → TabM
43
70.5 [57.4, 90.0]
42/0/1
TabICL v2 → XGBoost
43
69.4 [54.0, 80.6]
41/0/2
TabPFN v3 → TabM
46
70.9 [55.3, 91.6]
43/0/3
TabPFN v3 → XGBoost
48
67.9 [52.7, 78.1]
46/0/2
Table 3: Recovery of teacher gains on TabArena. We restrict to datasets where the teacher reduces default-student test error by more than 1% . Gap recovery is 100(Eraw−EKD)/(Eraw−ET) , using errors averaged over official splits; entries report the dataset median with bootstrap 95% CIs, and W/T/L against the raw student.
Model
Teacher (ms)
KD (ms)
Speedup
Δ Error (%)
Offline (s)
TabICL v2 → TabM
232.2
47.8
3.0×
1.72 [0.53, 2.94]
186.3 [93.1, 450.9]
TabICL v2 → XGBoost
232.2
29.1
8.7×
2.33 [1.07, 5.54]
48.0 [20.0, 146.5]
TabPFN v3 → TabM
566.9
46.7
11.1×
2.76 [0.70, 4.12]
192.4 [110.4, 563.4]
TabPFN v3 → XGBoost
566.9
29.7
21.6×
2.53 [1.33, 4.17]
59.6 [30.6, 173.9]
Table 4: Inference savings, prediction error, and distillation cost. Results on 51 TabArena datasets at outer split 0. Latencies are per 32-row request; speedup uses dataset-paired teacher/KD ratios, and Δ Error reports the relative error increase with paired-bootstrap 95% CIs. Offline time indicates end-to-end distillation tuime. Point estimates are dataset medians and brackets denote IQRs.
Context rows
<1 K ( n=10 )
1 K– <10 K ( n=26 )
≥10 K ( n=15 )
Model
Speedup
Nbreak-even
Speedup
Nbreak-even
Speedup
Nbreak-even
TabICL v2 → TabM
1.9 ×
3,260.5
2.1 ×
2,073.6
31.4 ×
468.3
TabICL v2 → XGBoost
3.0 ×
393.0
6.8 ×
211.2
23.9 ×
160.1
TabPFN v3 → TabM
10.2 ×
503.2
6.8 ×
388.9
53.7 ×
266.5
TabPFN v3 → XGBoost
13.1 ×
97.5
20.3 ×
97.6
32.3 ×
100.8
Table 5: Inference savings and break-even requests by teacher context size. Entries report dataset medians of inference speedup tT/tS and break-even 32-row requests. All 51 datasets are included; students are faster on all 15 datasets with at least 10K teacher context rows.
Median KD gain (%) ↑
Model
All
Binary
Multiclass
Regression
W/T/L
n=297
n=117
n=80
n=100
n=300
TabICL v2 → TabM
4.20
8.46
5.80
1.44
238/3/59
[3.14, 5.33]
[5.55, 10.41]
[2.81, 10.65]
[0.47, 2.99]
TabICL v2 → XGBoost
5.69
9.53
9.08
2.71
253/2/45
[4.60, 7.50]
[7.32, 12.23]
[7.21, 13.46]
[1.74, 3.49]
Table 6: The fixed recipe improves both students across TALENT task types. KD gains are median relative primary-error reductions against the same default supervised student, after averaging five matched seeds. Brackets show paired dataset-bootstrap 95% intervals. W/T/L counts wins, ties, and losses on all 300 datasets; gain summaries exclude three zero-error baselines.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Elo ↑
Improvability (%) ↓
Win rate (%) ↑
AutoGluon 1.5 (extreme, 4h)
1668.8
5.57
87.9
TabPFN v3 (Default)
1657.8
6.83
87.2
TabPFN v2.6 (Default)
1605.4
8.62
83.9
RealTabPFN v2.5 (Tuned+Ensemble)
1585.5
8.26
82.4
TabICL v2 (Default)
1582.7
7.60
82.2
RealTabPFN v2.5 (Tuned)
1543.2
9.02
79.2
Appendix
Table 7: Full TabArena results on 51 datasets and 816 dataset–split pairs. All metrics are computed jointly over the 76 displayed configurations: 72 official baselines and four KD students using default student configurations and r=2 . Shaded rows denote KD students.
Model
Δ Elo ↑
KD gain (%) ↑
W/T/L
TabICL v2 → TabM
+56.4 [4.2, 106.6]
1.37 [0.45, 2.03]
34/0/17
TabICL v2 → XGBoost
+72.1 [14.3, 130.3]
1.64 [0.50, 2.34]
33/0/18
TabPFN v3 → TabM
+86.0 [34.0, 138.5]
1.55 [0.71, 2.28]
35/0/16
TabPFN v3 → XGBoost
+101.6 [44.3, 162.0]
1.84 [0.87, 3.78]
36/0/15
Appendix
Table 8: KD versus tuned-and-ensembled students after excluding design splits. Results use 51 datasets and 663 dataset–split pairs, omitting splits 0–2. KD gain is the median relative primary-error reduction against the matched tuned ensemble; W/T/L counts dataset-level wins, ties, and losses. Elo is jointly recalibrated over all 76 configurations. Brackets show paired dataset-bootstrap 95% intervals (20,000 resamples).
Model
Elo ↑
Improvability (%) ↓
Win rate (%) ↑
TabPFN v3 → TabM
1513.5
11.18
76.5
RealMLP (Tuned+Ensemble)
1490.8
11.11
74.5
TabICL v2 → TabM
1484.0
11.43
73.9
TabPFN v3 → XGBoost
1460.6
12.36
71.6
TabICL v2 → XGBoost
1431.1
12.63
68.6
TabM (Tuned+Ensemble)
1427.6
12.49
68.3
Appendix
Table 9: Broader TabArena comparison after excluding design splits. The displayed models match Table 2 ; all metrics use the 76-configuration pool on the remaining 663 dataset–split pairs. KD students use default configurations and r=2 . Shading marks KD students; 1st , 2nd , and 3rd results among displayed non-teacher models are highlighted.
Figure 6: Teacher advantage and KD gain after excluding design splits. Each point is one of the 51 TabArena datasets, using errors averaged over splits ≥3 . Teacher and KD gains are relative error reductions over the default student, 100(1−E/Eraw) ; the dashed line denotes full recovery. Annotations report Spearman correlations with paired dataset-bootstrap 95% intervals (20,000 resamples). These correlations are descriptive because both axes share the raw baseline.
Model
n
Gap recovery (%)
KD W/T/L
TabICL v2 → TabM
42
71.8 [56.0, 89.8]
41/0/1
TabICL v2 → XGBoost
43
69.0 [53.9, 83.1]
41/0/2
TabPFN v3 → TabM
45
73.4 [56.4, 86.2]
42/0/3
TabPFN v3 → XGBoost
48
67.8 [52.7, 78.8]
46/0/2
Appendix
Table 10: Recovery of teacher gains after excluding design splits. Errors are averaged over retained splits before selecting datasets where the teacher reduces default-student error by more than 1% . Gap recovery is 100(Eraw−EKD)/(Eraw−ET) . Entries report the median with dataset-bootstrap 95% intervals (20,000 resamples); W/T/L compares KD with the raw student on these n datasets.
Tabular foundation models (TFMs) achieve strong performance on health datasets, but their inference cost and infrastructure requirements limit practical use. We study whether their predictive behavior can be transferred to lightweight tabular models through knowledge distillation. Since in-context TFMs condition on the training set at inference time, naive distillation can introduce context leakage; we address this with stratified out-of-fold teacher labeling. Across 19 healthcare datasets, 6 TFM teachers, 4 student families, and several multi-teacher ensembles, we find that distilled students retain at least 90% of teacher AUC, outperforming teachers in some cases, while running at least 26× faster on CPU and preserving calibration and fairness critical for health applications. Moreover, multi-teacher averaging does not consistently improve over the best single teacher. Leakage-aware distillation is thus a viable route for bringing TFM-quality predictions into inference-constrained health settings.
A fraud scorer needs to answer in under 2 ms. The best tabular foundation models (TFMs) take 151-1,275 ms on GPU. We close this gap by distilling the TFM offline into an XGBoost or CatBoost student that runs natively on CPU. The central obstacle is specific to in-context learning (ICL) teachers: they leak labels when scoring their own training set, so the soft targets collapse to near-one-hot vectors with no inter-class structure left to distill. Stratified out-of-fold (OOF) teacher labeling prevents this. Across 153 classification datasets drawn from TALENT, OpenML-CC18, TabZilla, and TabArena, distilling TabICLv2 into XGBoost gives 0.882 macro-mean AUC (96.5% of teacher AUC) at 1.9 ms on CPU, a 38x to 860x speedup across teacher-student pairs with a statistically significant edge over a tuned CatBoost baseline (Wilcoxon p = 0.0008; 51% win rate). Four further findings: teacher rank transfers exactly to student rank; gains concentrate on low-dimensional data (< 21 features: +0.011 over CatBoost vs. >21 features: +0.001); multi-teacher averaging helps MLP students (+0.006, p = 0.003) but adds less than 0.001 for tree students; and on high-dimensional tasks where the teacher itself trails CatBoost, distillation makes things worse rather than better. The full pipeline is open-sourced as part of the TabTune library.
Tabular Foundation Models (TFMs) have demonstrated strong empirical performance as black-box inference engines through in-context learning. However, their use in transfer learning is limited by two obstacles: strict context-size constraints and sensitivity to distribution shifts between source and target tasks. Directly pooling heterogeneous source data can therefore lead to negative transfer. To address these challenges, we propose Context-Constrained Transfer Learning via ANchoring and DIstillation (TL-ANDI), a posterior-aware distillation framework for TFMs. TL-ANDI constructs a compact source context by solving a budget-constrained optimal transport problem whose cost jointly measures target covariate coverage and posterior compatibility. The selected anchor samples are then equipped with locally distilled labels and combined with a residual calibration step using target data.