SimplexUQ: An Evaluation Framework and Benchmark for Conformal Uncertainty on Simplex-Valued Predictions
Organizations: University of Pittsburgh · Duke University
Abstract
Conformal prediction guarantees marginal coverage, but a single calibration threshold can still spread that coverage unevenly, over-covering easy regions and under-covering hard ones. SimplexUQ is, to our knowledge, the first benchmark and reproducible protocol for measuring this allocation problem on simplex-valued predictions; it compares existing conformal wrappers rather than proposing a new one. Its task suite, SimplexTasks-12, combines six controlled synthetic regimes with six frozen-predictor real tasks spanning class probabilities, topic mixtures, spectral abundances, cell-type fractions, age distributions, and emotion mixtures. Each comparison fixes the predictor, score, and response-free stratification map, varies only the wrapper, and reports marginal coverage, worst-stratum coverage, max disparity, and within-task radius and compute. Global calibration can look valid while failing badly: on CIFAR-10 it attains 0.900 marginal coverage but only 0.542 in the worst entropy stratum, and Mondrian calibration raises that stratum to 0.886 while reducing max disparity from 0.358 to 0.022. No wrapper dominates, however. Under smooth synthetic heterogeneity, several repairs are competitive; fixed-map analyses show that rankings depend on the evaluation groups and protocol; and in a 12-task comparison, Mondrian has lower disparity on its single target partition for all 12 tasks, whereas BatchMVP has lower disparity over overlapping groups on five. These are empirical comparisons, not new coverage guarantees. A controlled predictor-bias sweep shows that removing predictor bias only partly reduces global-threshold disparity. We release task cards, result provenance, permitted derived arrays, and rebuild instructions, and treat wrapper selection as a diagnostic comparison rather than a universal ranking.
Figures & tables
| Wrapper | Threshold rule | Guarantee or role |
| Global | One split-conformal threshold from all calibration scores | Marginal coverage (Theorem 6 ) |
| Mondrian | One threshold per stratum of the fixed map | Coverage within each stratum (Theorem 9 ) |
| TwoStage | Local scale fitted on separate data; calibrate | Marginal coverage (Theorem 7 ); primary normalization wrapper |
| OneShot | kNN scale fitted on the same calibration residuals | Not guaranteed (Proposition 2 ); diagnostic |
| TrainRes | Scale fitted on training residuals, then held fixed | Marginal, but can misallocate (Proposition 3 ); diagnostic |
| FullCP | Full-conformal local-scale variant | Exact reference; highest compute |
| Task | Domain | Design hypothesis | ||
| CIFAR-10 softmax | Vision | 10 | 10000 | Classification-style grouped stress test |
| 20 Newsgroups topics | NLP | 10 | 7539 | Mild / smooth heterogeneity |
| SemEval AffectiveText | NLP/Affect | 6 | 1000 | Small-sample weak structure |
| Samson NMF unmixing | Remote sensing | 3 | 9025 | Aligned grouped heterogeneity |
| PBMC deconvolution | Genomics | 8 | 5000 | Semi-synthetic sensitivity case |
| UTKFace age LDL | Demo- graphics | 10 | 23687 | Strong grouped heterogeneity |
| Observed pattern | Diagnostic evidence | Prefer | Main caveat |
| Near-homogeneous regime | Low max disparity and acceptable worst-stratum coverage | Global split CP | Extra adaptation can add variance or compute without improving allocation |
| Coarse aligned heterogeneity | Hard/easy regions line up with entropy bins, dominant class, or other meaningful groups | Group-wise / Mondrian CP | Fragmented or weakly aligned groups can become unstable in small samples |
| Smooth heterogeneity | Coverage changes gradually across entropy or boundary proximity | Compare two-stage normalization with Mondrian; exact variants if affordable | Naive OneShot normalization is not guaranteed, and exactness can be expensive |
| Bias-type or predictor-driven failure | High disparity remains after both grouping and normalization | No wrapper recommendation; investigate the predictor or score | Wrapper choice alone may not repair structural misspecification |
| Small-sample weak-structure setting | Few calibration points per group and similar floors across valid methods | Conservative global or smooth normalization strategies | Mondrian can overfit the calibration split and hurt worst-stratum behavior |
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Source asset | Citation | Usage terms in this benchmark |
| CIFAR-10 | Official CIFAR-10 image benchmark, accessed through cached torchvision downloads | [ 21 ] | Public benchmark dataset used only to build frozen softmax predictions. The paper cites the official technical report requested by the source page and does not redistribute the image archive as part of the benchmark tables or figures. |
| Topics | 20 Newsgroups text collection, accessed through scikit-learn’s fetcher | [ 22 ] | Public educational/research benchmark used only to derive cached topic-mixture arrays. We treat the corpus as a source-cited benchmark asset rather than republishing the raw text collection. |
| Affective- Text | SemEval-2007 Task 14 headlines and gold emotion labels | [ 23 ] | Shared-task benchmark data used locally for evaluation. We do not repackage the raw headline collection in the paper outputs, and the conformal benchmark operates on derived cached predictions and normalized gold labels. |
| Samson | Public 95 95 Samson ROI plus endmember/abundance benchmark files used in hyperspectral unmixing | [ 24 ] | Academic benchmark files used unmodified with attribution. The public bundle we used does not ship an explicit license file, so we document it as a source-cited research asset and avoid presenting it as a newly redistributable paper asset. |
| PBMC | 10x Genomics PBMC3K single-cell reference, accessed through scanpy.datasets.pbmc3k | [ 25 ] | Public 10x tutorial data. Our benchmark uses derived pseudobulk mixtures and frozen deconvolution outputs rather than redistributing the raw count matrix. Source terms and any permitted redistribution are documented separately in the task card. |
| UTKFace | Official UTKFace aligned-and-cropped face images | [ 26 ] | Non-commercial research only according to the dataset homepage. Copyright remains with the original image owners, so the benchmark uses derived features and predictions and does not repackage the face-image archive in the paper outputs. |
| Item | Benchmark setting |
| CPU-only conformal evaluation on a personally owned Apple MacBook Air. | |
| Predictor regeneration | Not part of repeated benchmarking. Each task first constructs or caches a fixed simplex predictor , after which all conformal runs operate on frozen arrays. |
| Synthetic repetitions | 200 random calibration/test splits per regime. |
| Real repetitions | 50 random calibration/test splits for most tasks; AffectiveText and PBMC use 200. |
| Cheap methods | Global, Mondrian/partition, and most normalization-style methods are typically sub-second per repetition once is cached. |
| Expensive methods | FullCP: 2.30 sec/rep on D2, 3.35 on Topics, 3.71 on Samson, and 3.25 on full-scale CIFAR-10. |
| Task | Boundary | Entropy | Dominant | KMeans | Winner stability |
| CIFAR-10 | Mondrian (0.021, 0.901) | Mondrian (0.022, 0.902) | Mondrian (0.020, 0.902) | Mondrian (0.024, 0.902) | Stable: Mondrian |
| Topics | Mondrian (0.024, 0.902) | Mondrian (0.025, 0.902) | Mondrian (0.038, 0.901) | Mondrian (0.028, 0.902) | Stable: Mondrian |
| AffectiveText | Mondrian (0.070, 0.919) | Mondrian (0.071, 0.919) | Global (0.106, 0.906) | Mondrian (0.073, 0.918) | Mixed: Global, Mondrian |
| Samson | Mondrian (0.020, 0.902) | Mondrian (0.025, 0.903) | Mondrian (0.011, 0.901) | Mondrian (0.079, 0.903) | Stable: Mondrian |
| UTKFace | Mondrian (0.014, 0.900) | Mondrian (0.013, 0.901) | Mondrian (0.027, 0.901) | Mondrian (0.019, 0.902) | Stable: Mondrian |
| PBMC | Mondrian (0.016, 0.902) | Mondrian (0.029, 0.903) | Mondrian (0.031, 0.903) | Mondrian (0.031, 0.904) | Stable: Mondrian |
| Task | Method | Coverage | Worst | Disparity | Radius | Runtime (s) |
| CIFAR-10 | Global | 0.900 | 0.542 | 0.358 | 0.153 | 0.000 |
| CIFAR-10 | Mondrian | 0.902 | 0.886 | 0.022 | 0.188 | 0.001 |
| CIFAR-10 | TwoStage | 0.901 | 0.641 | 0.259 | 0.119 | 0.079 |
| CIFAR-10 | FullCP | 0.900 | 0.656 | 0.244 | 0.107 | 3.247 |
| CIFAR-10 | Jackknife+ | 0.900 | 0.657 | 0.243 | 0.108 | 0.151 |
| CIFAR-10 | OneShot | 0.882 | 0.609 | 0.291 | 0.091 | 0.150 |
| Task | Method | Coverage | Worst | Disparity | Radius | Runtime (s) |
| AffectiveText | Global | 0.904 | 0.892 | 0.100 | 29.073 | 0.000 |
| AffectiveText | Mondrian | 0.902 | 0.582 | 0.356 | 28.893 | 0.000 |
| AffectiveText | TwoStage | 0.905 | 0.893 | 0.100 | 86.693 | 0.003 |
| AffectiveText | FullCP | 0.903 | 0.889 | 0.101 | 38.979 | 0.308 |
| AffectiveText | Jackknife+ | 0.894 | 0.880 | 0.101 | 37.746 | 0.004 |
| AffectiveText | OneShot | 0.890 | 0.873 | 0.102 | 34.973 | 0.004 |
| Stratum | Global | Mondrian | TwoStage | FullCP | Jackknife+ | OneShot | TrainRes | Weighted |
| S0 | 0.999 | 0.902 | 0.999 | 0.998 | 0.998 | 0.998 | 0.999 | 0.650 |
| S1 | 0.998 | 0.901 | 0.998 | 0.996 | 0.996 | 0.995 | 0.996 | 0.000 |
| S2 | 0.992 | 0.905 | 0.983 | 0.978 | 0.978 | 0.970 | 0.977 | 0.000 |
| S3 | 0.972 | 0.902 | 0.882 | 0.872 | 0.872 | 0.838 | 0.871 | 0.000 |
| S4 | 0.542 | 0.902 | 0.641 | 0.656 | 0.657 | 0.609 | 0.663 | 0.000 |
| Stratum | Global | Mondrian | TwoStage | FullCP | Jackknife+ | OneShot | TrainRes | Weighted |
| S0 | 0.762 | 0.902 | 0.953 | 0.956 | 0.956 | 0.953 | 0.955 | 0.649 |
| S1 | 0.972 | 0.902 | 0.883 | 0.885 | 0.885 | 0.878 | 0.888 | 0.529 |
| S2 | 0.958 | 0.902 | 0.895 | 0.900 | 0.901 | 0.884 | 0.901 | 0.510 |
| S3 | 0.934 | 0.904 | 0.893 | 0.877 | 0.878 | 0.864 | 0.879 | 0.673 |
| S4 | 1.000 | 0.904 | 0.800 | 0.805 | 0.805 | 0.783 | 0.803 | 0.873 |
| Stratum | Global | Mondrian | TwoStage | FullCP | Jackknife+ | OneShot | TrainRes | Weighted |
| S0 | 0.973 | 0.729 | 0.976 | 0.976 | 0.976 | 0.976 | 0.977 | 0.887 |
| S1 | 0.825 | 0.913 | 0.867 | 0.866 | 0.866 | 0.856 | 0.871 | 0.274 |
| S2 | 0.908 | 0.903 | 0.893 | 0.881 | 0.881 | 0.869 | 0.885 | 0.212 |
| S3 | 0.897 | 0.899 | 0.922 | 0.912 | 0.913 | 0.901 | 0.912 | 0.165 |
| S4 | 0.901 | 0.900 | 0.878 | 0.895 | 0.895 | 0.883 | 0.896 | 0.200 |
| Stratum | Global | Mondrian | TwoStage | FullCP | Jackknife+ | OneShot | TrainRes | Weighted |
| S0 | 0.892 | 0.906 | 0.894 | 0.892 | 0.882 | 0.877 | 0.892 | 0.753 |
| S1 | 1.000 | 0.899 | 0.996 | 1.000 | 0.999 | 0.999 | 0.999 | 0.903 |
| S2 | 1.000 | 0.950 | 0.995 | 0.999 | 0.997 | 0.996 | 0.997 | 0.903 |
| S3 | 1.000 | 0.916 | 0.997 | 0.998 | 0.997 | 0.995 | 0.998 | 0.936 |
| S4 | 1.000 | 0.687 | 0.998 | 0.997 | 0.997 | 0.994 | 1.000 | 0.973 |
| Stratum | Global | Mondrian | TwoStage | FullCP † | Jackknife+ | OneShot | TrainRes | Weighted |
| S0 | 0.991 | 0.901 | 0.811 | 0.988 | 0.842 | 0.831 | 0.844 | 0.456 |
| S1 | 0.949 | 0.901 | 0.800 | 0.973 | 0.796 | 0.778 | 0.793 | 0.378 |
| S2 | 0.981 | 0.900 | 0.917 | 0.987 | 0.926 | 0.921 | 0.927 | 0.397 |
| S3 | 0.921 | 0.901 | 0.905 | 0.861 | 0.905 | 0.891 | 0.905 | 0.327 |
| S4 | 0.410 | 0.904 | 0.999 | 0.764 | 0.933 | 0.902 | 0.931 | 0.000 |
| Stratum | Global | Mondrian | TwoStage | FullCP |
| S1 | 0.875 | 0.900 | 0.935 | 0.937 |
| S2 | 1.000 | 0.903 | 0.765 | 0.750 |
| Stratum | Jackknife+ † | OneShot † | TrainRes † | Weighted † |
| S0 | 0.919 | 0.911 | 0.917 | 0.588 |
| S1 | 0.759 | 0.753 | 0.764 | 0.970 |
| S2 | 0.802 | 0.795 | 0.808 | 0.987 |
| S3 | 0.833 | 0.826 | 0.836 | 0.974 |
| S4 | 0.530 | 0.528 | 0.539 | 0.993 |
| Method | CIFAR-10 | Samson | Topics | AffectiveText | UTKFace | PBMC |
| Global | 0.0001 | 0.0001 | 0.0001 | 0.0001 | 0.0002 | 0.0000 |
| Mondrian | 0.0005 | 0.0006 | 0.0005 | 0.0002 | 0.0012 | 0.0001 |
| TwoStage | 0.0794 | 0.0204 | 0.0703 | 0.0027 | 0.0752 | 0.0181 |
| FullCP | 3.2473 | 3.7123 | 3.3472 | 0.3085 | 0.7031 | 1.0547 |
| Jackknife+ | 0.1507 | 0.0274 | 0.1325 | 0.0036 | 0.1135 | 0.0471 |
| OneShot | 0.1495 | 0.0270 | 0.1317 | 0.0036 | 0.1161 | 0.0462 |
| Task | Global | Mondrian | Ratio |
| CIFAR-10 | 0.153 | 0.188 | 1.225 |
| Samson | 27.146 | 28.284 | 1.042 |
| Topics | 6.025 | 6.042 | 1.003 |
| AffectiveText | 29.073 | 28.893 | 0.994 |
| UTKFace | 3.164 | 1.892 | 0.598 |
| PBMC | 42.265 | 39.982 | 0.946 |
| Task | A - B | Difference [95% interval] |
| D2 | Global - FullCP | |
| D2 | Global - TwoStage | |
| D2 | Global - OneShot | |
| D2 | Global - TrainRes | |
| D2 | TrainRes - FullCP | |
| D2 | TrainRes - Jackknife+ |
| Task | Method | Cov. | Worst | Radius | ||
| D1 | Global | 0.901 | 0.020 | 0.023 | 0.891 | 0.431 |
| D1 | Mondrian | 0.907 | 0.038 | 0.039 | 0.879 | 0.438 |
| D1 | BatchMVP | 0.891 | 0.041 | 0.048 | 0.863 | 0.426 |
| D2 | Global | 0.903 | 0.198 | 0.253 | 0.702 | 0.518 |
| D2 | Mondrian | 0.917 | 0.052 | 0.069 | 0.880 | 0.499 |
| D2 | BatchMVP | 0.887 | 0.057 | 0.069 | 0.848 | 0.456 |
| Task | Baseline | [95% interval] | [95% interval] |
| D1 | Global | ||
| D1 | Mondrian | ||
| D2 | Global | ||
| D2 | Mondrian | ||
| D3 | Global | ||
| D3 | Mondrian |
| Task | Map | Groups | M - G [95% interval] | ||
| CIFAR-10 | boundary | 5 | 0.259 | 0.020 | |
| CIFAR-10 | entropy | 5 | 0.358 | 0.020 | |
| CIFAR-10 | dominant | 5 | 0.049 | 0.021 | |
| CIFAR-10 | kmeans | 5 | 0.071 | 0.025 | |
| Topics | boundary | 5 | 0.045 | 0.025 | |
| Topics | entropy | 5 | 0.014 | 0.025 |
| Task | Req./real. | M - G [95% interval] | Min. cal. | ||
| CIFAR-10 | 2/2 | 0.102 | 0.010 | 1982.2 | |
| CIFAR-10 | 3/3 | 0.188 | 0.014 | 1310.0 | |
| CIFAR-10 | 5/5 | 0.358 | 0.020 | 774.7 | |
| CIFAR-10 | 10/10 | 0.683 | 0.037 | 377.6 | |
| Topics | 2/2 | 0.007 | 0.010 | 1493.9 | |
| Topics | 3/3 | 0.010 | 0.017 | 984.5 |
| Bias scale | Mean score | Global | Global worst | Mondrian |
| 0.45 | 0.552 | 0.151 | 0.750 | 0.051 |
| 0.3 | 0.383 | 0.141 | 0.761 | 0.051 |
| 0.15 | 0.233 | 0.127 | 0.775 | 0.051 |
| 0.0 | 0.163 | 0.128 | 0.773 | 0.050 |