cs.LGOct 8, 2026

Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits

Authors: Kanghui Ning, Marin Biloš, James T. Wilson, Yilang Zhang, Kashif Rasul, Dongjin Song, Anderson Schneider, Yuriy Nevmyvaka

Organizations: School of Computing, University of Connecticut, Storrs, USA. · Department of Machine Learning Research, Morgan Stanley, New York, USA.

Abstract

Which forms of test-time compute improve the predictions of strong pretrained tabular foundation models (TFMs)? We systematically study this along three axes: adaptation, aggregation, and context construction. Our evaluation spans modern TFMs across the TabArena benchmark, supplemented by experiments on wide and large-scale tables from OpenML. For adaptation, we introduce DiagScale, a diagonal query-key similarity update. It trains only 0.003-0.03% of model parameters and achieves gains comparable to full fine-tuning across three independently pretrained backbones. For aggregation, both pool composition and selection strategy matter. TabPFN-3 already averages predictions from different preprocessing variants of the same data, and adding more such predictions yields diminishing returns. With a broader pool of 96 configurations, greedy selection reduces error by 2.4% relative to the default predictor, but uniform averaging increases error. For context construction, attention-guided retrieval improves TabPFN-3's predictions on some large tables and supports source pools beyond the full context memory limit. The context expansion methods we test yield no consistent improvement. Taken together, our results suggest that adaptation and selective aggregation yield consistent benchmark-level gains. The benefits of context construction depend more on the task and data regime. Adaptation and aggregation over the same backbone yield further gains when combined, but require substantially more computation than default inference. These trade-offs motivate choosing strategies according to the available computation budget. Code is available at https://github.com/kanghui-learning/test-time-compute-for-tabular-foundation-models.

Figures & tables

Appendix figures & tables33 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models

    May 7, 2026Amir Rezaei Balef, Mykhailo Koshil, Katharina EggenspergerTabular Foundation ModelsTransformer Interpretability

  2. TabPFN-3.5: Technical Report

    Sep 15, 2026Benjamin Jäger, Nick Erickson, Léo Grinsztajn +44Tabular Foundation ModelsInference-Time Optimization

  3. TabSwift: An Efficient Tabular Foundation Model with Row-Wise Attention

    Jun 5, 2026Si-Yang Liu, Han-Jia YeTabular Foundation ModelsEfficient Inference