cs.LGMar 9, 2026

Distributional Regression with Tabular Foundation Models: Evaluating Probabilistic Predictions via Proper Scoring Rules

Authors: Jonas LandsgesellPascal KnollTizian Wenzel

Organizations: University of Stuttgart (Stuttgart, Germany) · Ludwig Maximilian University of Munich (Munich, Germany) · Munich Center for Machine Learning (Munich, Germany)

Abstract

Modern tabular foundation models such as TabPFN and TabICL naturally produce full predictive distributions, while the benchmarks used to evaluate them (TabArena, TALENT, and others) still rely almost exclusively on point-estimate metrics (RMSE, R2R^2). This mismatch implicitly rewards machine learning models or pipelines that elicit a good conditional mean while ignoring the quality of the predictive distribution. We make the case for using proper scoring rules for training, fine-tuning, and benchmarking (ranking) of tabular foundation models. Although all strictly proper scoring rules are theoretically equivalent at the population level, they may differ on finite data: We demonstrate analytically and empirically that different scoring rules can induce different inductive biases during finite-sample optimization, leading to different model performance. We validate this finding by running fine-tuning experiments with TabPFN and TabICL using different scoring rules for various data sets, revealing non-trivial interactions between training objectives and evaluation metrics. Our results show that practitioners can adapt tabular foundation models to task-specific scoring objectives, and that the choice of scoring rule can influence model behavior in practice.

Explore similar work

CardsList
  1. Towards Evaluating Data Priors for Tabular Foundation Models

    Jun 28, 2026Zeynep Türkmen, Kürşat Kaya, Alexander Pfefferle +1Tabular Foundation ModelsPrior Knowledge

  2. Do Tabular Foundation Models Agree with Themselves?

    Aug 6, 2026Christian Klötergens, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme +1Tabular Foundation ModelsPosterior