Tabular foundation models (TFMs) provide predictive distributions for regression, but their prediction regions can exhibit undercoverage or overcoverage even when point predictions are accurate. We introduce C-USIM (Conditionally-Uniformized Score Integration Method), a lightweight application of highest predictive density split conformal prediction that accommodates multimodal predictions. Given calibration and test outputs, it requires no additional training or model inference. It provides finite-sample marginal validity under our assumptions. We bound conditional-marginal coverage gaps using distribution-estimation error and score discreteness, and examine coverage heterogeneity through percentile rank-score plots. Experiments with TabPFN and TabICL show improved marginal coverage accuracy and lower average conditional and group coverage errors. Under a fixed data budget, allocating more observations to calibration can reduce marginal coverage error despite less accurate point predictions.
Figures & tables
Figure 1: C-USIM overview with ncal calibration scores and target coverage 1−α . Using a pretrained TFM, the procedure predicts densities, computes density-rank scores, calibrates a threshold, and constructs prediction regions without fitting an additional model. See Section 4 for details.
Figure 2: Percentile rank–score plots over test covariates (plot (a)), obtained using TabPFN on the sine with shifted exponential noise function ( B.2.1 ). Plot (b) approximates Pt(Xi) at the nominal score threshold t=1−α (blue) and the ideal C-USIM threshold t=q1−α (red), which is estimated by the empirical (1−α) -quantile of the pooled Monte Carlo scores. The histogram summarizes interpolated estimates of Pt(X1),…,Pt(Xntest) , approximating the distribution of Pt(X) . See Appendices B.2 and B.3 for details.
Figure 3: Synthetic and Journal SJR overview. (a,b) Synthetic seed means, averaged equally over the selected mechanisms. (c,d) Journal SJR seed means of marginal coverage and CEC-X, respectively; (d) uses a representative covariate grouping. See Table 1 for experiment settings and Appendix B.6 for detailed results.
Figure 4: Split-ratio sensitivity on B.2.2 ; settings follow Table 1 . Rows correspond to TabPFN and TabICL. Columns show mean-function RMSE, absolute marginal coverage error, and CCAD; lower values are better. Standard boxplots summarize variation across seeds; connected diamonds mark the means.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Percentile rank–score diagnostics for B.2.1 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 6: Percentile rank–score diagnostics for B.2.1 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 7: Percentile rank–score diagnostics for B.2.1 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 8: Percentile rank–score diagnostics for B.2.2 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 9: Percentile rank–score diagnostics for B.2.2 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 10: Percentile rank–score diagnostics for B.2.2 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Setting
Synthetic
Journal SJR
Split ratio
Dataset size
1,792 per run
27,803 cleaned rows
1,792 per run
Total label budget
1,536
1,536
1,536
Input dimensions
1, 5, 10, 20
7 (categorical)
1, 5, 10, 20
Run seeds
100–109
12100–12299
25600–25649
Number of seeds
10
200
50
Plug-in context observations
1,536
1,536
–
Appendix
Table 1: Experimental settings. Synthetic and split-ratio sample counts are per mechanism and run; conditional response draws are additional Monte Carlo samples. Split-ratio coverage and CCAD use the known conditional CDF rather than the stored response draws. A dash denotes an unused setting.
Table 2: Synthetic comparison under a fixed budget of 1,536 labels, averaged over ten seeds per model and case. Each arrow denotes Plug-in HPD(1536) → C-USIM(512+1024). Coverage is in percent and CCAD in percentage points (pp). Cases are defined in Appendix B.2 .
Figure 17: Synthetic coverage (a,b) and CCAD (c,d) across ten seeds for Plug-in HPD(1536) and C-USIM(512+1024); methods and label budgets follow Table 2 . Boxes show Q1–Q3 and medians, with 1.5-IQR whiskers and all outliers. Dashed lines mark 95% coverage. CCAD is in percentage points on model-specific scales. These seed distributions are not confidence intervals.
Figure 18: Journal SJR across 200 seeds for Plug-in HPD(1536) and C-USIM(512+1024); methods and label budgets follow Table 3 . (a,b) CEC-X, weighted by group size; (c,d) group coverage; (e,f) absolute group error. Errors are computed within each seed. Boxes show Q1–Q3 and medians, with 1.5-IQR whiskers and all outliers, not confidence intervals. Red bands mark higher mean absolute group error with C-USIM; dashed lines mark 95% coverage. Errors are in percentage points. CEC-X shares a scale across models; group plots use model-specific scales with visual padding below zero.
Measure
TabPFN
TabICL
Plug-in HPD(1536) → C-USIM
Marginal coverage (%)
91.50→94.83
92.93→94.87
Mean marginal gap (pp)
3.499→0.553
2.241→0.611
Mean group gap (pp)
2.979→1.760
2.458→1.856
CEC-X (pp)
3.698→1.662
2.782→1.707
Mean set length
1.149→1.453
1.255→1.540
Appendix
Table 3: Journal SJR under a fixed budget of 1,536 labels at 95% target coverage, averaged over 200 seeds. Each arrow denotes Plug-in HPD(1536) → C-USIM(512+1024). Absolute marginal and group gaps and CEC-X are computed within each seed before averaging, and are reported in percentage points (pp).
Figure 19: Additional real-world examples under a total budget of 1,536 labels. Points are means over 50 seeds, not confidence intervals. (a,b) JP Anime; (c,d) Allstate Claims Severity. Dashed lines mark 95% coverage. CEC-X uses the representative covariate grouping and is computed within each seed before averaging. Settings and detailed results are in Table 4 and Appendix B.6.3 .
Dataset
Rows
Features
Categorical
Test rows
Response scale
JP Anime
15,351
10
8
5,372
ln(Score)
Allstate
188,318
130
116
9,731
Claim loss
Appendix
Table 4: Additional real-world datasets.
Dataset
Measure
TabPFN
TabICL
Plug-in HPD (1536) → C-USIM
JP Anime
Marginal coverage (%)
88.54→94.96
91.40→94.96
Mean marginal gap (pp)
6.461→0.590
3.598→0.540
Mean group gap (pp)
6.224→2.770
4.235→2.837
CEC-X (pp)
6.598→2.429
4.211→2.521
Mean set length
0.447→0.565
0.494→0.539
Appendix
Table 5: Additional real-world datasets under a fixed total budget of 1,536 labels, averaged over 50 seeds. Each arrow denotes Plug-in HPD(1536) → C-USIM(512+1024). Marginal and group errors are computed within each seed before averaging; group measures use the representative grouping. Length uses the first 256 fixed test inputs.
TabPFN
TabICL
Case
Gap (pp)
CCAD (pp)
Gap (pp)
CCAD (pp)
1229:307 ( ≃8:2 ) → 512:1024 ( 1:2 )
B.2.1
1.033→0.554
2.110→2.152
0.911→0.440
2.830→3.381
B.2.1
0.964→0.497
1.998→2.093
1.013→0.561
2.370→2.541
B.2.1
0.857→0.546
1.231→1.379
1.098→0.552
2.486→3.104
B.2.2
1.026→0.583
2.271→2.218
1.119→0.650
2.569→2.634
Appendix
Table 6: Split-ratio sensitivity across all six selected mechanisms with 1,536 total labels. Entries are means over the same 50 seeds. Gap is the per-seed absolute deviation of marginal coverage from 95%; gap and CCAD use percentage points (pp).
Model
Example
Coverage (%)
CCAD (pp)
Mean length
Plug-in HPD(512) → C-USIM
TabPFN
B.2.1
92.95→95.28
2.947→2.159
0.618→0.717
B.2.1
90.24→94.74
5.074→2.186
0.805→1.290
B.2.1
92.11→94.91
3.268→1.708
0.516→0.580
B.2.2
88.49→95.33
6.569→2.201
0.841→2.100
B.2.2
88.42→94.81
6.714→1.642
2.281→3.313
Appendix
Table 7: Synthetic comparison using identical predictive densities from a 512-observation context. Each arrow denotes Plug-in HPD(512) → C-USIM(512+1024), averaged over ten seeds. C-USIM additionally uses 1,024 calibration responses.
Figure 20: Synthetic coverage, CCAD, and prediction-set length for all three configurations. Points are means over ten seeds, not confidence intervals. Horizontal lines separate mechanisms. The two 512-context methods share predictive densities. Dashed lines mark 95% coverage. All interval components contribute to length, excluding gaps.
Figure 21: Journal SJR with both plug-in baselines and C-USIM, using the representative grouping. (a,d) CEC-X over 200 seeds: boxes show quartiles and medians, whiskers extend to 1.5 IQR, all outliers are retained, and colored markers denote means. (b,e) Mean coverage in each group. (c,f) Mean within-seed absolute group error. Horizontal lines in the group plots separate groups. These summaries are not confidence intervals. Dashed lines mark 95% coverage; all ten groups are retained.
Measure
TabPFN
TabICL
Plug-in HPD(512) → C-USIM
Marginal coverage (%)
91.06→94.83
91.72→94.87
Mean marginal gap (pp)
3.964→0.553
3.431→0.611
Mean group gap (pp)
3.846→1.760
3.630→1.856
CEC-X (pp)
4.333→1.662
3.863→1.707
Mean set length
1.261→1.453
1.359→1.540
Appendix
Table 8: Journal SJR using identical predictive densities from a 512-observation context, averaged over 200 seeds. Each arrow denotes Plug-in HPD(512) → C-USIM(512+1024). Both configurations use the same joint calibration/test query table; only C-USIM uses the 1,024 calibration responses. Group errors use the representative grouping. Coverage uses all 9,731 test rows; length uses the first 256 inputs in the fixed test order.
Dataset
Measure
TabPFN
TabICL
Plug-in HPD (512) → C-USIM
JP Anime
Marginal coverage (%)
91.05→94.96
94.01→94.96
Mean marginal gap (pp)
3.974→0.590
1.853→0.540
Mean group gap (pp)
4.697→2.770
3.330→2.837
CEC-X (pp)
4.681→2.429
3.103→2.521
Mean set length
0.486→0.565
0.522→0.539
Appendix
Table 9: Additional real-world datasets using identical predictive densities from a 512-observation context, averaged over 50 seeds. Each arrow denotes Plug-in HPD(512) → C-USIM(512+1024). Only C-USIM uses the additional 1,024 calibration responses. Group measures use the representative grouping; response scales and test sizes follow Table 4 .
Figure 22: JP Anime with both plug-in baselines and C-USIM, using the representative grouping. (a,d) CEC-X over 50 seeds: boxes show quartiles and medians, whiskers extend to 1.5 IQR, all outliers are retained, and colored markers denote means. (b,e) Mean group coverage. (c,f) Mean within-seed absolute group error. Horizontal lines separate groups; group 9 has 15 test observations. Dashed lines mark 95% coverage. These summaries are not confidence intervals.
Figure 23: Allstate Claims Severity with both plug-in baselines and C-USIM, using the representative grouping. (a,d) CEC-X over 50 seeds: boxes show quartiles and medians, whiskers extend to 1.5 IQR, all outliers are retained, and colored markers denote means. (b,e) Mean group coverage. (c,f) Mean within-seed absolute group error. Horizontal lines separate all ten groups; dashed lines mark 95% coverage. These summaries are not confidence intervals.
Tabular Foundation Models (TFMs) are currently the best approach to tabular prediction problems. They are constructed as transformers that approximate the Bayesian posterior predictive distribution based on a pre-training prior. These univariate predictors can be converted into multivariate ones autoregressively by sampling one target and adding it to the features. However, the faithfulness of the resulting joint has not been investigated. Furthermore, TFMs cannot be evaluated against the posterior itself, at least not on real-world datasets, because the ground-truth distribution is unknown. We therefore propose asking a different question: could a model's predictions result from any joint distribution? To answer this question, we pose two requirements that any such model must satisfy. The first is marginalization consistency, which demands that marginalized conditionals are equal to directly predicted marginals. The second is factorization consistency, which demands that different factorization orders result in equal joint distributions. Every TFM that we evaluate violates both of these requirements for both classification and regression across all datasets.
Christian Klötergens, Vijaya Krishna Yalavarthi, Lars Schmidt-Thieme +1
Institute of Computer Science & VWFS DARC, University of Hildesheim, Hildesheim, Germany
Recent Tabular Foundation Models (TFMs) have demonstrated state-of-the-art predictive performance, often surpassing Gradient-Boosted Decision Trees (GBDTs). However, the trustworthiness of these models, particularly their uncertainty quantification, has been largely overlooked. We investigate this gap through an extensive study comparing TFMs, GBDTs, and classical baselines on the 112 datasets of the TALENT benchmark. Our results reveal a performance-uncertainty trade-off: although TFMs achieve the highest predictive performance, measured by AUC, they exhibit lower conditional coverage under conformal prediction, measured by SSCS, compared to GBDTs. Complementary experiments on synthetic datasets further characterize the regimes in which this effect intensifies. We conclude that while TFMs advance predictive frontiers, achieving well-calibrated uncertainty remains a major open challenge for their reliable adoption. Code is available at: https://github.com/jose-melo/high-performance-low-reliability
José Lucas De Melo Costa, Fabrice Popineau, Arpad Rimmel +1
CentraleSup´elec, Universit´e Paris-Saclay Gif-sur-Yvette - France
Modern tabular foundation models such as TabPFN and TabICL naturally produce full predictive distributions, while the benchmarks used to evaluate them (TabArena, TALENT, and others) still rely almost exclusively on point-estimate metrics (RMSE, R2). This mismatch implicitly rewards machine learning models or pipelines that elicit a good conditional mean while ignoring the quality of the predictive distribution. We make the case for using proper scoring rules for training, fine-tuning, and benchmarking (ranking) of tabular foundation models. Although all strictly proper scoring rules are theoretically equivalent at the population level, they may differ on finite data: We demonstrate analytically and empirically that different scoring rules can induce different inductive biases during finite-sample optimization, leading to different model performance. We validate this finding by running fine-tuning experiments with TabPFN and TabICL using different scoring rules for various data sets, revealing non-trivial interactions between training objectives and evaluation metrics. Our results show that practitioners can adapt tabular foundation models to task-specific scoring objectives, and that the choice of scoring rule can influence model behavior in practice.
Jonas Landsgesell, Pascal Knoll, Tizian Wenzel
University of Stuttgart (Stuttgart, Germany) · Ludwig Maximilian University of Munich (Munich, Germany) · Munich Center for Machine Learning (Munich, Germany)