Tabular foundation models (TFMs) provide predictive distributions for regression, but their prediction regions can exhibit undercoverage or overcoverage even when point predictions are accurate. We introduce C-USIM (Conditionally-Uniformized Score Integration Method), a lightweight application of highest predictive density split conformal prediction that accommodates multimodal predictions. Given calibration and test outputs, it requires no additional training or model inference. It provides finite-sample marginal validity under our assumptions. We bound conditional-marginal coverage gaps using distribution-estimation error and score discreteness, and examine coverage heterogeneity through percentile rank-score plots. Experiments with TabPFN and TabICL show improved marginal coverage accuracy and lower average conditional and group coverage errors. Under a fixed data budget, allocating more observations to calibration can reduce marginal coverage error despite less accurate point predictions.
Figures & tables
Figure 1: C-USIM overview with ncal calibration scores and target coverage 1−α . Using a pretrained TFM, the procedure predicts densities, computes density-rank scores, calibrates a threshold, and constructs prediction regions without fitting an additional model. See Section 4 for details.
Figure 2: Percentile rank–score plots over test covariates (plot (a)), obtained using TabPFN on the sine with shifted exponential noise function ( B.2.1 ). Plot (b) approximates Pt(Xi) at the nominal score threshold t=1−α (blue) and the ideal C-USIM threshold t=q1−α (red), which is estimated by the empirical (1−α) -quantile of the pooled Monte Carlo scores. The histogram summarizes interpolated estimates of Pt(X1),…,Pt(Xntest) , approximating the distribution of Pt(X) . See Appendices B.2 and B.3 for details.
Figure 3: Synthetic and Journal SJR overview. (a,b) Synthetic seed means, averaged equally over the selected mechanisms. (c,d) Journal SJR seed means of marginal coverage and CEC-X, respectively; (d) uses a representative covariate grouping. See Table 1 for experiment settings and Appendix B.6 for detailed results.
Figure 4: Split-ratio sensitivity on B.2.2 ; settings follow Table 1 . Rows correspond to TabPFN and TabICL. Columns show mean-function RMSE, absolute marginal coverage error, and CCAD; lower values are better. Standard boxplots summarize variation across seeds; connected diamonds mark the means.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Percentile rank–score diagnostics for B.2.1 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 6: Percentile rank–score diagnostics for B.2.1 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 7: Percentile rank–score diagnostics for B.2.1 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 8: Percentile rank–score diagnostics for B.2.2 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 9: Percentile rank–score diagnostics for B.2.2 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Figure 10: Percentile rank–score diagnostics for B.2.2 , seed 2026: C-USIM ( ncal→∞ ) and plug-in HPD.
Setting
Synthetic
Journal SJR
Split ratio
Dataset size
1,792 per run
27,803 cleaned rows
1,792 per run
Total label budget
1,536
1,536
1,536
Input dimensions
1, 5, 10, 20
7 (categorical)
1, 5, 10, 20
Run seeds
100–109
12100–12299
25600–25649
Number of seeds
10
200
50
Plug-in context observations
1,536
1,536
–
Appendix
Table 1: Experimental settings. Synthetic and split-ratio sample counts are per mechanism and run; conditional response draws are additional Monte Carlo samples. Split-ratio coverage and CCAD use the known conditional CDF rather than the stored response draws. A dash denotes an unused setting.
Table 2: Synthetic comparison under a fixed budget of 1,536 labels, averaged over ten seeds per model and case. Each arrow denotes Plug-in HPD(1536) → C-USIM(512+1024). Coverage is in percent and CCAD in percentage points (pp). Cases are defined in Appendix B.2 .
Figure 17: Synthetic coverage (a,b) and CCAD (c,d) across ten seeds for Plug-in HPD(1536) and C-USIM(512+1024); methods and label budgets follow Table 2 . Boxes show Q1–Q3 and medians, with 1.5-IQR whiskers and all outliers. Dashed lines mark 95% coverage. CCAD is in percentage points on model-specific scales. These seed distributions are not confidence intervals.
Figure 18: Journal SJR across 200 seeds for Plug-in HPD(1536) and C-USIM(512+1024); methods and label budgets follow Table 3 . (a,b) CEC-X, weighted by group size; (c,d) group coverage; (e,f) absolute group error. Errors are computed within each seed. Boxes show Q1–Q3 and medians, with 1.5-IQR whiskers and all outliers, not confidence intervals. Red bands mark higher mean absolute group error with C-USIM; dashed lines mark 95% coverage. Errors are in percentage points. CEC-X shares a scale across models; group plots use model-specific scales with visual padding below zero.
Measure
TabPFN
TabICL
Plug-in HPD(1536) → C-USIM
Marginal coverage (%)
91.50→94.83
92.93→94.87
Mean marginal gap (pp)
3.499→0.553
2.241→0.611
Mean group gap (pp)
2.979→1.760
2.458→1.856
CEC-X (pp)
3.698→1.662
2.782→1.707
Mean set length
1.149→1.453
1.255→1.540
Appendix
Table 3: Journal SJR under a fixed budget of 1,536 labels at 95% target coverage, averaged over 200 seeds. Each arrow denotes Plug-in HPD(1536) → C-USIM(512+1024). Absolute marginal and group gaps and CEC-X are computed within each seed before averaging, and are reported in percentage points (pp).
Figure 19: Additional real-world examples under a total budget of 1,536 labels. Points are means over 50 seeds, not confidence intervals. (a,b) JP Anime; (c,d) Allstate Claims Severity. Dashed lines mark 95% coverage. CEC-X uses the representative covariate grouping and is computed within each seed before averaging. Settings and detailed results are in Table 4 and Appendix B.6.3 .
Dataset
Rows
Features
Categorical
Test rows
Response scale
JP Anime
15,351
10
8
5,372
ln(Score)
Allstate
188,318
130
116
9,731
Claim loss
Appendix
Table 4: Additional real-world datasets.
Dataset
Measure
TabPFN
TabICL
Plug-in HPD (1536) → C-USIM
JP Anime
Marginal coverage (%)
88.54→94.96
91.40→94.96
Mean marginal gap (pp)
6.461→0.590
3.598→0.540
Mean group gap (pp)
6.224→2.770
4.235→2.837
CEC-X (pp)
6.598→2.429
4.211→2.521
Mean set length
0.447→0.565
0.494→0.539
Appendix
Table 5: Additional real-world datasets under a fixed total budget of 1,536 labels, averaged over 50 seeds. Each arrow denotes Plug-in HPD(1536) → C-USIM(512+1024). Marginal and group errors are computed within each seed before averaging; group measures use the representative grouping. Length uses the first 256 fixed test inputs.
TabPFN
TabICL
Case
Gap (pp)
CCAD (pp)
Gap (pp)
CCAD (pp)
1229:307 ( ≃8:2 ) → 512:1024 ( 1:2 )
B.2.1
1.033→0.554
2.110→2.152
0.911→0.440
2.830→3.381
B.2.1
0.964→0.497
1.998→2.093
1.013→0.561
2.370→2.541
B.2.1
0.857→0.546
1.231→1.379
1.098→0.552
2.486→3.104
B.2.2
1.026→0.583
2.271→2.218
1.119→0.650
2.569→2.634
Appendix
Table 6: Split-ratio sensitivity across all six selected mechanisms with 1,536 total labels. Entries are means over the same 50 seeds. Gap is the per-seed absolute deviation of marginal coverage from 95%; gap and CCAD use percentage points (pp).
Model
Example
Coverage (%)
CCAD (pp)
Mean length
Plug-in HPD(512) → C-USIM
TabPFN
B.2.1
92.95→95.28
2.947→2.159
0.618→0.717
B.2.1
90.24→94.74
5.074→2.186
0.805→1.290
B.2.1
92.11→94.91
3.268→1.708
0.516→0.580
B.2.2
88.49→95.33
6.569→2.201
0.841→2.100
B.2.2
88.42→94.81
6.714→1.642
2.281→3.313
Appendix
Table 7: Synthetic comparison using identical predictive densities from a 512-observation context. Each arrow denotes Plug-in HPD(512) → C-USIM(512+1024), averaged over ten seeds. C-USIM additionally uses 1,024 calibration responses.
Figure 20: Synthetic coverage, CCAD, and prediction-set length for all three configurations. Points are means over ten seeds, not confidence intervals. Horizontal lines separate mechanisms. The two 512-context methods share predictive densities. Dashed lines mark 95% coverage. All interval components contribute to length, excluding gaps.
Figure 21: Journal SJR with both plug-in baselines and C-USIM, using the representative grouping. (a,d) CEC-X over 200 seeds: boxes show quartiles and medians, whiskers extend to 1.5 IQR, all outliers are retained, and colored markers denote means. (b,e) Mean coverage in each group. (c,f) Mean within-seed absolute group error. Horizontal lines in the group plots separate groups. These summaries are not confidence intervals. Dashed lines mark 95% coverage; all ten groups are retained.
Measure
TabPFN
TabICL
Plug-in HPD(512) → C-USIM
Marginal coverage (%)
91.06→94.83
91.72→94.87
Mean marginal gap (pp)
3.964→0.553
3.431→0.611
Mean group gap (pp)
3.846→1.760
3.630→1.856
CEC-X (pp)
4.333→1.662
3.863→1.707
Mean set length
1.261→1.453
1.359→1.540
Appendix
Table 8: Journal SJR using identical predictive densities from a 512-observation context, averaged over 200 seeds. Each arrow denotes Plug-in HPD(512) → C-USIM(512+1024). Both configurations use the same joint calibration/test query table; only C-USIM uses the 1,024 calibration responses. Group errors use the representative grouping. Coverage uses all 9,731 test rows; length uses the first 256 inputs in the fixed test order.
Dataset
Measure
TabPFN
TabICL
Plug-in HPD (512) → C-USIM
JP Anime
Marginal coverage (%)
91.05→94.96
94.01→94.96
Mean marginal gap (pp)
3.974→0.590
1.853→0.540
Mean group gap (pp)
4.697→2.770
3.330→2.837
CEC-X (pp)
4.681→2.429
3.103→2.521
Mean set length
0.486→0.565
0.522→0.539
Appendix
Table 9: Additional real-world datasets using identical predictive densities from a 512-observation context, averaged over 50 seeds. Each arrow denotes Plug-in HPD(512) → C-USIM(512+1024). Only C-USIM uses the additional 1,024 calibration responses. Group measures use the representative grouping; response scales and test sizes follow Table 4 .
Figure 22: JP Anime with both plug-in baselines and C-USIM, using the representative grouping. (a,d) CEC-X over 50 seeds: boxes show quartiles and medians, whiskers extend to 1.5 IQR, all outliers are retained, and colored markers denote means. (b,e) Mean group coverage. (c,f) Mean within-seed absolute group error. Horizontal lines separate groups; group 9 has 15 test observations. Dashed lines mark 95% coverage. These summaries are not confidence intervals.
Figure 23: Allstate Claims Severity with both plug-in baselines and C-USIM, using the representative grouping. (a,d) CEC-X over 50 seeds: boxes show quartiles and medians, whiskers extend to 1.5 IQR, all outliers are retained, and colored markers denote means. (b,e) Mean group coverage. (c,f) Mean within-seed absolute group error. Horizontal lines separate all ten groups; dashed lines mark 95% coverage. These summaries are not confidence intervals.
University of Stuttgart (Stuttgart, Germany) · Ludwig Maximilian University of Munich (Munich, Germany) · Munich Center for Machine Learning (Munich, Germany)