Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits
Organizations: School of Computing, University of Connecticut, Storrs, USA. · Department of Machine Learning Research, Morgan Stanley, New York, USA.
Abstract
Which forms of test-time compute improve the predictions of strong pretrained tabular foundation models (TFMs)? We systematically study this along three axes: adaptation, aggregation, and context construction. Our evaluation spans modern TFMs across the TabArena benchmark, supplemented by experiments on wide and large-scale tables from OpenML. For adaptation, we introduce DiagScale, a diagonal query-key similarity update. It trains only 0.003-0.03% of model parameters and achieves gains comparable to full fine-tuning across three independently pretrained backbones. For aggregation, both pool composition and selection strategy matter. TabPFN-3 already averages predictions from different preprocessing variants of the same data, and adding more such predictions yields diminishing returns. With a broader pool of 96 configurations, greedy selection reduces error by 2.4% relative to the default predictor, but uniform averaging increases error. For context construction, attention-guided retrieval improves TabPFN-3's predictions on some large tables and supports source pools beyond the full context memory limit. The context expansion methods we test yield no consistent improvement. Taken together, our results suggest that adaptation and selective aggregation yield consistent benchmark-level gains. The benefits of context construction depend more on the task and data regime. Adaptation and aggregation over the same backbone yield further gains when combined, but require substantially more computation than default inference. These trade-offs motivate choosing strategies according to the available computation budget. Code is available at https://github.com/kanghui-learning/test-time-compute-for-tabular-foundation-models.
Figures & tables
| Setting | Updated | Params | Elo | Norm. score |
|---|---|---|---|---|
| Full fine-tuning | all trainable weights | 58.3M | ||
| LoRA ( ) | low-rank updates | 3.51M | ||
| Value–output | , | 12.9M | ||
| Query–key | , | 12.9M | ||
| SoftScale | query-side scaling | 1.03M |
| Backbone | Backbone params | Trainable params (% of backbone) | Elo DiagScale | Elo Full FT |
|---|---|---|---|---|
| TabPFN-3 | 58.3M | 12.7K ( ) | ||
| TabICL v2 | 28.5M | 7.3K ( ) | ||
| TabFM | 1.65B | 53.8K ( ) |
Appendix figures & tables33 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Reference (TPU/JAX) | Ours (bf16, H100) | % |
|---|---|---|---|
| Binary — reported as AUC (30 datasets) | |||
| Amazon_employee_access | 0.1416 | 0.1419 | |
| APSFailure | 0.007387 | 0.005329 | |
| bank-marketing | 0.2309 | 0.2306 | |
| Bank_Customer_Churn | 0.1229 | 0.1233 | |
| Bioresponse | 0.1178 | 0.1136 | |
| Dataset | OpenML ID | Total rows | Predictors | Task | Largest pool |
|---|---|---|---|---|---|
| Higgs | 45570 | 11,000,000 | 28 | binary | 10,000,000 |
| US_Accidents | 46650 | 7,728,394 | 43 | 4-class | 7,678,394 |
| COMET_MC | 5889 | 7,619,400 | 4 | binary | 7,569,400 |
| Poker-Hand | 1567 | 1,025,009 | 10 | 10-class | 900,000 |
| Covertype | 1596 | 581,012 | 54 | 7-class | 500,000 |
| Reference entry | Imputed (%) |
|---|---|
| TabFM (frozen reproduction) | 0 |
| AutoGluon 1.5 (extreme, 4h) | 0 |
| TabPFN-3 (default) | 0 |
| TabPFN-2.6 (default) | 0 |
| TabICLv2 (default) | 0 |
| RealTabPFN-2.5 (tuned + ensembled) | 0 |
| Backbone | Task | Full checkpoint | DiagScale | Fraction (%) |
|---|---|---|---|---|
| TabPFN-3 | Classification | 53,153,144 | 13,056 | 0.02456 |
| Regression | 58,274,944 | 12,672 | 0.02175 | |
| TabICL v2 | Classification | 27,552,258 | 7,296 | 0.02648 |
| Regression | 28,544,991 | 7,296 | 0.02556 | |
| TabFM | Classification | 1,639,444,298 | 53,760 | 0.00328 |
| Regression | 1,647,782,989 | 53,760 | 0.00326 |
| Backbone | Method | Global | Five-fold (mean SD) |
|---|---|---|---|
| TabPFN-3 | Full FT | ||
| TabPFN-3 | DiagScale | ||
| TabICL v2 | Full FT | ||
| TabICL v2 | DiagScale |
| Setting | Updated | Params | Elo | Score | Error [95% CI] | Wins | Fallback |
|---|---|---|---|---|---|---|---|
| softscale | query scaling | 1.0M | 34/51 | ||||
| full | all trainable weights | 58.3M | 37/51 | ||||
| qk | query + key projections | 12.9M | 33/51 | ||||
| middle3 | middle ICL blocks 10-12 | 6.4M | 36/51 | ||||
| attn | all attention sublayers | 26.2M | 35/51 | ||||
| mlp | all MLP sublayers | 25.2M | 38/51 |
| Backbone | Error change | 95% interval | Wins |
|---|---|---|---|
| TabPFN-3 | 21/51 | ||
| TabICL v2 | 26/51 | ||
| TabFM | 28/51 |
| Backbone | Method | Median | Mean | Best | Wins | (size) |
|---|---|---|---|---|---|---|
| TabPFN-3 | full fine-tuning | 37/51 | ||||
| TabPFN-3 | DiagScale | 31/51 | ||||
| TabICL v2 | full fine-tuning | 35/51 | ||||
| TabICL v2 | DiagScale | 34/51 | ||||
| TabFM | full fine-tuning | 31/51 | ||||
| TabFM | DiagScale | 25/51 |
| Dataset | Train only | Train only |
|---|---|---|
| Houses | ||
| Protein | ||
| Bankruptcy |
| Reducer | error [95% CI] | Wins | ||
| 1 | 1 | inner uniform | 4/51 | |
| 1 | 2 | inner uniform | 9/51 | |
| 1 | 4 | inner uniform | 18/51 | |
| 1 | 8 | inner uniform | — (baseline) | |
| 1 | 16 | inner uniform | 33/51 | |
| 1 | 32 | inner uniform | 42/51 | |
| Type | Metric | Reducer | Median | Worst | Wins |
|---|---|---|---|---|---|
| binary (30) | ROC AUC | uniform | 9/30 | ||
| best-single | 13/30 | ||||
| greedy | 22/30 | ||||
| multiclass (8) | log loss | uniform | 1/8 | ||
| best-single | 7/8 | ||||
| greedy | 8/8 |
| 1 | ||||
|---|---|---|---|---|
| 8 | ||||
| 16 | ||||
| 32 | ||||
| 64 | ||||
| 96 |
| Total | Within memory | Beyond wall | |||
|---|---|---|---|---|---|
| Dataset | Eval. | Geo. | Model-aware | Model-aware | |
| Poker-Hand | 1.03M | 900k | — | ||
| Covertype | 0.58M | 500k | — | ||
| US_Accidents | 7.73M | 2.9M | at 7.68M | ||
| Higgs | 11M | 2.9M | at 6M, at 10M | ||
| COMET_MC | 7.62M | 2.9M | at 7.57M | ||
| Pool | Capped union | Recursive splitting | Time ( ) | |
|---|---|---|---|---|
| Capped | Split | |||
| 6M | ||||
| 8M | ||||
| 10M | ||||
| Dataset | Seed | LR | Rounds | Val. loss | Test loss |
|---|---|---|---|---|---|
| US_Accidents | 0 | 0.02 | 7,972 | 0.14011 | 0.13994 |
| 1 | 0.02 | 8,093 | 0.14328 | 0.14279 | |
| 2 | 0.02 | 8,056 | 0.13990 | 0.14081 | |
| 3 | 0.02 | 8,411 | 0.13998 | 0.14279 | |
| 4 | 0.02 | 7,479 | 0.14216 | 0.14276 | |
| Higgs | 0 | 0.02 | 72,769 | 0.45890 | 0.45808 |
| Dataset | attn. % | union | 1 view | 8 views | ||
|---|---|---|---|---|---|---|
| blood-transfusion-service-center | 499 | 4 | 100 | 499 | ||
| diabetes | 512 | 4 | 99 | 512 | ||
| maternal_health_risk | 676 | 4 | 91 | 634 | ||
| concrete_compressive_strength | 687 | 4 | 93 | 684 | ||
| airfoil_self_noise | 1,002 | 4 | 83 | 972 | ||
| students_dropout_and_academic_success | 2,950 | 4 | 50 | 1,814 |
| Dataset | Feed-all | Random ids | Attention | Attention vs. random | |
|---|---|---|---|---|---|
| Covertype | 500k | ||||
| Poker-Hand | 900k | ||||
| Higgs | 900k | ||||
| Higgs | 2.9M |
| Arm | Log loss | vs. feed-all | AUC | Rare-class share |
|---|---|---|---|---|
| Feed-all ( k) | — | |||
| Attention union | ||||
| Random identities, same size | ||||
| Class-marginal repair |
| Dataset | Draws | Union | Overlap | ||
|---|---|---|---|---|---|
| COMET_MC | 900k | ||||
| Covertype | 500k | ||||
| Poker-Hand | 900k | ||||
| Higgs | 900k | ||||
| Higgs | 2.9M |
| Default | Selected | controls | ||||
|---|---|---|---|---|---|---|
| Dataset | Domain | vs. 1 view | vs. default | Per view vs. fixed | Imp. vs. random | |
| Bioresponse | molecular | |||||
| hiva_agnostic | molecular | |||||
| QSAR-TID-11 | molecular | |||||
| CIFAR_10 | image | |||||
| micro-mass | mass-spec | |||||
| No selection | Importance, | Random, | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | Metric | 1 v. | 8 v. | gini | 1 v. | pin | p.v. | 1 v. | pin | p.v. | Imp. pin |
| Bioresponse | 1 AUC | ||||||||||
| hiva_agnostic | log loss | ||||||||||
| QSAR-TID-11 | RMSE | — | |||||||||
| CIFAR_10 | log loss | ||||||||||
| micro-mass | log loss | ||||||||||
| Family | Scope | Passes | Median | Example/control |
|---|---|---|---|---|
| Blind synthesis | 51 dataset operator | 3 | none | |
| Transduction | 57 dataset operator | 14 | anneal , airfoil | |
| Learned rows | 171 probe/dataset/dose | 10 | concrete improves over placebo | |
| Full-size joint rows | 5 datasets | 0 | no improvement observed | |
| Self-stacking | 19 datasets | 1 | Marketing_Campaign only | |
| OOF target encoding | 9 datasets | 2 | Marketing_Campaign |
| Operator | Family | Median | IQR | Improves | |
|---|---|---|---|---|---|
| Gaussian jitter | Blind synthesis | 19 | 2/19 | ||
| Row interpolation | Blind synthesis | 19 | 1/19 | ||
| Class-balanced interpolation | Blind synthesis | 13 | 0/13 | ||
| Whole-query pseudo-labels | Transduction | 19 | 3/19 | ||
| Cross-half pseudo-labels | Transduction | 19 | 5/19 | ||
| Cross-half, confidence | Transduction | 14 | 2/14 |
| Setting | Training / construction | Deployment | Labels / selection |
|---|---|---|---|
| Blind synthesis | Jitter, pair interpolation, or class-balanced interpolation of training rows; no optimization. | , with at most added rows. | Inherited training labels; interpolated targets for regression mixup. |
| Transduction | Predict the query batch from . For cross-half expansion, split queries into two parts. | plus pseudo-labeled rows from the other half; the whole-query control adds the whole batch. | Model predictions only; optional probability or quantile-width filter. |
| Row controls | Duplicate training rows or sample noise anchors; no optimization. | plus the constructed rows. | Training labels or target tails for duplication; labels for noise anchors sampled from training class proportions or fixed to the majority class. |
| Fold-rotated rows | Optimize free rows using real training folds plus as context; score a different training fold. | . | Fixed seed-row labels; training-side early stopping. |
| Free-row distillation | Optimize row features with alone as context; predict real training queries. | ; alone is a diagnostic arm. | Fixed seed-row labels; Nominal 10% training-side checkpoint slice; initialization details below. |
| MLP-generated rows | Same distillation objective; optimize a conditional MLP with a fixed latent bank. | plus generated rows. | Fixed conditioning labels from training rows; training-side checkpoint selection. |
| Backbone | Method | Top-5 share | Datasets improved |
|---|---|---|---|
| TabFM | DiagScale | 25/51 | |
| TabFM | full fine-tuning | 31/51 | |
| TabICL v2 | DiagScale | 34/51 | |
| TabICL v2 | full fine-tuning | 35/51 | |
| TabPFN-3 | DiagScale | 31/51 | |
| TabPFN-3 | aggregation | 41/51 |
| Outcome | ||||
|---|---|---|---|---|
| TabFM / DiagScale | 0.212 | 0.282 | 0.599 | 0.358 |
| TabFM / full FT | 0.284 | 0.906 | 0.528 | 0.941 |
| TabICL v2 / DiagScale | 0.426 | 0.039 | 0.130 | 0.123 |
| TabICL v2 / full FT | 0.412 | 0.056 | 0.130 | 0.138 |
| TabPFN-3 / DiagScale | 0.339 | 0.317 | 0.444 | 0.264 |
| TabPFN-3 / aggregation | 0.496 | 0.004 | 0.031 | 0.013 |
| Recipe | Elo | Elo | Fit (h) | Predict (h) | Total (h) | frozen |
|---|---|---|---|---|---|---|
| TabPFN-3 | ||||||
| Frozen | 1658 | |||||
| 32 native views | 1675 | |||||
| DiagScale | 1682 | |||||
| Full fine-tuning | 1681 | |||||
| Aggregation (96) | 1721 | |||||