TAFFY: A Task-Adaptive Tabular Foundation Model with In-Context Diversity
Authors: Zijian Li, Xiangchen Song, Gongxu Luo, Jie Qiao, Ruichu Cai, Zhenhao Chen, Xinshuai Dong, Fan Feng, +2 more
Organizations: Carnegie Mellon University · Mohamed bin Zayed University of Artificial Intelligence · Guangdong University of Technology · University of California, San Diego
Recent progress in tabular foundation models suggests that training on synthetic tasks can substantially improve in-context learning capabilities, with overall performance largely depending on how well models can infer task-specific predictive relationships from the available context during inference. In this paper, we introduce TAFFY, a tabular foundation model with an In-Context Diversity Prior and a Task-Conditioned Looped Transformer that strengthen this ability. Specifically, to construct each synthetic pretraining context, the In-Context Diversity Prior samples from multiple related environments derived via controlled interventions and distribution shifts on a shared causal process. This in-context diversity encourages the model to learn a more comprehensive and task-specific representation. Moreover, the Task-Conditioned Looped Transformer iteratively and selectively applies a shared group of Transformer blocks to refine contextual representations, with a task-conditioned gate modulating the final hidden-state update. This enables task-adaptive iterative refinement. Together, these components encourage the model to identify predictive relationships from contextual contrasts during pretraining and dynamically modulate context integration for each task. Across six classification and five regression benchmark datasets, TAFFY attains the lowest average rank.
Figures & tables
Figure 1: Predictive performance and pretraining-data efficiency. (a–c) Average accuracy ranks on OpenML-CC18, BCCO, and TALENT using the ranking pools of Table 1 (lower is better; rank axes start at 2). (d) Accuracies on TALENT classification datasets versus cumulative maximum feature elements. Markers denote measured checkpoints connected by lines; the grey line marks TabICLv2’s best measured accuracy.
Figure 2: The pretraining pipeline of Taffy combines the in-context diversity prior with the task-conditioned looped transformer. (a) In-Context Diversity Prior proceeds through three steps: base-table generation, multiple-environment construction, and context assembly. (b) Column and row encoders produce representations for a looped Transformer, whose final update is modulated by a support-derived task gate before query prediction.
Model
Classification
Regression
BCCO
OpenML
PFN
TALENT
TabArena
TabZilla
BCCO
CTR23
PFN
TALENT
TabArena
Ensemble methods
AutoGluon
10.62
8.29
9.88
9.96
9.98
8.52
5.96
5.85
5.96
5.89
6.15
Tree-based methods
CatBoost
10.29
9.88
9.30
10.70
10.78
11.02
7.20
7.39
7.25
7.02
7.77
XGBoost
11.66
10.52
9.95
11.81
11.03
11.07
11.30
10.55
10.93
10.39
10.92
Table 1: Average ranks across all evaluated classification and regression datasets (lower is better). The classification row uses Taffy-4L on every suite; regression uses the separately trained model.
Figure 3: Component analyses: (a) in-context versus cross-context diversity; (b) task-specific versus global gating; (c) matched training/inference depth; (d) extra inference loops after two-loop training. In (a,b), each variant is ranked by accuracy against four fixed baselines (TabICLv1, TabICLv2, LimiX-2M, and LimiX-16M) on each dataset, with average ranks for ties before averaging across datasets. In (c), ranks are computed on all datasets of each benchmark. Lower ranks are better; thin lines show checkpoints, thick lines show smoothed trends.
Figure 4: Average-rank ablation on TALENT classification datasets. Each variant is ranked separately against TabICLv1, TabICLv2, LimiX-2M, and LimiX-16M.
Property
2L
3L
4L
Categorical fraction
−.533∗
−.453∗
−.452∗
Excess kurtosis
+.230∗
+.326∗
+.323∗
Label entropy
+.084
+.010
−.011
Missing ratio
+.021
+.042
+.031
Table 2: Spearman correlations between α and task attributes. ∗ indicates corrected q<0.05 .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Operator
Implementation
Preserved
Changed
Strength
Distribution shifts (row sampling from M0 )
Label shift
Reweight rows by ae(y)=exp(τeue(y))
p(x∣y)
p(y) , hence p(y∣x)
KL =0.1
Covariate shift
Reweight rows by be(x)=exp(−τe∥x−μe∥2)
p(y∣x)
p(x)
KL =0.1
Conditional shift
Reweight rows within each class by ce(x,y)=exp(−τe∥x−μe,y∥2)
p(y)
p(x∣y) , hence p(y∣x)
KL =0.1
Interventions (mechanism changes in M0 )
Hard intervention
Set zj:=ve,j with ve,j∼U[ℓj,uj] for j∈Ie
Other mechanisms
zj and its descendants
5% of non-label nodes
Appendix
Table 3: Environment operators of the In-Context Diversity Prior. The three distribution shifts reweight rows sampled from the unchanged base model, with strength fixed by KL(pe∥p0)=0.1 . The two interventions modify the mechanisms of 5% of the non-label nodes and regenerate the data.
Category
Statistics
Dim.
Size and validity
log(1+d) , log(1+ns) , log(1+N) , ns/N , d/H , 1−d/H , and the fraction of finite support values over valid features.
7
Support labels
Observed class count divided by Kmax ; label entropy divided by logKmax ; largest and smallest proportions among observed classes, with Kmax=20 .
4
Support features
Per-feature mean, standard deviation, mean absolute value, minimum, maximum, and range.
6×4
Per-feature fractions of zero, near-zero, near-integer, and near-binary values.
4×4
Total
51
Appendix
Table 4: Statistics used by the task-specific support gate. The descriptor contains 51 entries. Here, H is the input feature width before grouping, including padding, and d≤H is the number of valid features. Feature summaries are aggregated across valid features by mean, standard deviation, minimum, and maximum.
Model
Time/step (s)
Relative
25k steps (h)
GPU-hours
Single-pass backbone
7.01
1.00×
48.66 (2.03 days)
3,114
Taffy-2L
8.79
1.25×
61.03 (2.54 days)
3,906
Taffy-3L
10.62
1.51×
73.74 (3.07 days)
4,719
Taffy-4L
12.62
1.80×
87.67 (3.65 days)
5,611
Appendix
Table 5: Pretraining cost on 64 AMD MI210 GPUs. Time per step is measured; total time and GPU-hours are projected for the 25,000-step schedule. Relative is the time per step divided by that of the single-pass backbone.
Method
Avg. rank ↓
Elo ↑
RF
8.42
1000.00
TabICLv1
5.17
1595.23
LimiX-2M
5.50
1552.44
TabPFN2
5.17
1595.23
Mitra
8.25
1044.05
LimiX-16M
5.12
1600.49
Appendix
Table 6: Results on TALENT classification datasets with more than ten classes. Only methods with observed results on all such datasets are included.
Model
Classification
Regression
Average
BCCO
OpenML
PFN
TALENT
TabArena
TabZilla
BCCO
CTR23
PFN
TALENT
TabArena
Ensemble methods
AutoGluon
10.53
8.29
9.88
10.00
9.98
8.52
5.86
5.85
5.96
6.00
6.15
7.91
Tree-based methods
CatBoost
10.47
9.88
9.30
10.72
10.78
11.02
7.12
7.39
7.25
6.99
7.77
8.97
XGBoost
11.68
10.52
9.96
11.71
11.03
11.09
11.28
10.55
10.93
10.29
10.92
10.91
Appendix
Table 7: Average ranks on the LimiX-2M-eligible, sample-compatible evaluation pool (lower is better). Rank aggregation uses the complete applicable-method subset within each suite; dashes denote incomplete task-type coverage.
Classification
Regression
Method
TALENT
BCCO
OpenML
PFN
TabArena
TabZilla
TALENT
BCCO
CTR23
PFN
TabArena
Mean
AutoGluon
1139.13
1001.45
1217.39
1064.14
1022.49
1187.80
1241.40
1295.80
1310.03
1257.72
1332.51
1188.17
CatBoost
1105.99
1015.31
1150.20
1088.56
986.67
1086.68
1171.15
1210.85
1219.76
1178.34
1204.78
1128.94
XGBoost
1056.39
957.92
1123.70
1061.09
975.45
1084.46
981.73
956.97
1050.70
964.64
970.83
1016.72
ET
986.81
934.06
1080.91
1017.65
800.16
1053.96
1016.33
1027.80
1078.09
1002.07
1074.19
1006.55
RF
1000.00
1000.00
1000.00
1000.00
1000.00
1000.00
1000.00
1000.00
1000.00
1000.00
1000.00
1000.00
Appendix
Table 8: Corresponding observed-only Elo ratings under the operational development-set reconstruction (higher is better).
Model
Classification
Regression
Average
BCCO
OpenML
PFN
TALENT
TabArena
TabZilla
BCCO
CTR23
PFN
TALENT
TabArena
Ensemble methods
AutoGluon
1011.38
1217.39
1064.14
1138.25
1021.54
1187.80
1303.99
1310.03
1257.72
1236.25
1332.51
1189.18
Tree-based methods
CatBoost
1013.58
1150.20
1088.56
1106.17
973.79
1086.68
1216.33
1219.76
1178.34
1174.86
1204.78
1128.46
XGBoost
962.97
1123.70
1061.09
1061.68
976.48
1084.46
957.85
1050.70
964.64
986.98
970.83
1018.31
Appendix
Table 9: Elo ratings on the LimiX-2M-eligible, sample-compatible evaluation pool (higher is better; RF = 1000 per benchmark). Only observed pairwise outcomes are used; missing scores are not imputed, and dashes denote incomplete task-type coverage.
Descriptor
Model
Partial ρ [95% CI]
q
Categorical fraction
Taffy-2L
−0.533[−0.652,−0.394]
0.0003
Taffy-3L
−0.453[−0.580,−0.299]
0.0003
Taffy-4L
−0.452[−0.579,−0.296]
0.0003
Numerical excess kurtosis
Taffy-2L
0.230[0.057,0.383]
0.0338
Taffy-3L
0.326[0.167,0.466]
0.000703
Taffy-4L
0.323[0.167,0.463]
0.000703
Appendix
Table 10: Adjusted associations between α and TALENT dataset properties. Categorical fraction and numerical excess kurtosis use separate multiple-testing corrections.
Support rows
Effective columns
Model
ρ
q
ρ
q
Taffy-2L
0.192
0.0374
0.138
0.199
Taffy-3L
0.297
2.31×10−4
0.053
0.660
Taffy-4L
0.397
3.44×10−7
0.022
0.881
Appendix
Table 11: Unadjusted associations between α and support size on TALENT datasets. All q -values use Benjamini–Hochberg correction over the original 216 tests.
Inference condition
Update coefficient
Average rank ↓
Task-specific gate
αS
1.57
Support-independent gate
tanh(a)
1.88
Gate bypass (full update)
1
2.55
Appendix
Table 12: Frozen-checkpoint gate intervention on matched TALENT classification datasets. Mean rank among the three inference conditions; lower is better.
Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. We explore how predictions emerge across depth, identify distinct stages of inference and reveal latent-space dynamics that differ from those of language models. Our findings indicate substantial depthwise redundancy across multiple models, suggesting iterative refinement with overlapping computations during inference stages. Guided by these insights, we design a proof-of-concept, looped single-layer model that uses only 20% of the original model's parameters while achieving comparable performance. The code is available at https://github.com/amirbalef/is_one_layer_enough.
Amir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger
TU Dortmund University, Dortmund, Germany · Lamarr Institute for Machine Learning and Artificial Intelligence, Dortmund, Germany · University of Tübingen, Tübingen, Germany
Tabular foundation models, such as TabPFNv2 and TabICL, have recently dethroned gradient-boosted trees at the top of predictive benchmarks, demonstrating the value of in-context learning for tabular data. We introduce TabICLv2, a new state-of-the-art foundation model for regression and classification built on three pillars: (1) a novel synthetic data generation engine designed for high pretraining diversity; (2) various architectural innovations, including a new scalable softmax in attention improving generalization to larger datasets without prohibitive long-sequence pretraining; and (3) optimized pretraining protocols, notably replacing AdamW with the Muon optimizer. On the TabArena and TALENT benchmarks, TabICLv2 without any tuning surpasses the performance of the current state of the art, RealTabPFN-2.5 (hyperparameter-tuned, ensembled, and fine-tuned on real data). With only moderate pretraining compute, TabICLv2 generalizes effectively to million-scale datasets under 50 GB GPU memory while being markedly faster than RealTabPFN-2.5. We provide extensive ablation studies to quantify these contributions and foster open research by releasing code for inference, pretraining, and synthetic data generation at https://github.com/soda-inria/tabicl.
Jingang Qu, David Holzmüller, Gaël Varoquaux +1
SODA Team, INRIA Saclay, Palaiseau, France · Probabl, France
Prior-Data Fitted networks (PFNs) have been very successful in tabular contexts, handling prediction tasks in context. However, they are designed for single-task inference, meaning that predicting several target values within a context requires repeated forward calls and precludes inter-task information sharing. We propose TabPFN-MT, which is trained on an expanded multi-target synthetic prior to capture inter-task dependencies in context. This model uses an expanded y-encoder and a shared decoder head to enable multitask in-context learning and simultaneous inference. The model is uniquely specialized for small-to-medium datasets by relying on in-context learning rather than traditional gradient-based training. Within this regime (averaging fewer than 1,000 samples), extensive evaluations across 344 datasets demonstrate that TabPFN-MT establishes a new state-of-the-art for deep tabular multitask learning. Furthermore, despite the inherent compute asymmetry of joint optimization, our model remains highly competitive with the latest state-of-the-art single-task ensembles. Notably, on multitask datasets it achieves an overall Accuracy rank of 4.89, the highest average rank among all models tested. Crucially, TabPFN-MT delivers this highly competitive performance while reducing the inference cost for T tasks from O(T) to O(1) forward passes, offering a massive computational efficiency improvement for multi-target tabular applications.
Cormac Cureton, Narges Armanfard
McGill University · Mila - Quebec AI Institute · Montreal, QC, Canada