Optimal Transport Dropout for Structured Predictive Uncertainty
Organizations: MOX–Department of Mathematics, Politecnico di Milano, Italy
Abstract
Deterministic neural networks and neural operators provide point predictions with no intrinsic measure of reliability. Yet, predictive uncertainty may stem from irreducible outcome variability, finite data, or limitations of the chosen model class. Monte Carlo dropout offers a computationally convenient way to construct a predictive distribution through stochastic feature masking, without training multiple independent networks or explicitly inferring a posterior over model parameters. However, its perturbation law is largely prescribed a priori and typically factorised across latent coordinates. We introduce Optimal Transport Dropout (OTD), which instead learns the predictive mapping and the law of its latent perturbations jointly. Starting from a simple independent reference distribution, OTD transports latent perturbations through a learnable flow and propagates them through the predictive neural network, thereby inducing a structured predictive law. Training uses the strictly proper Energy Score, while a kinetic-action term geometrically regularises the transport. Synthetic benchmarks show that OTD captures multimodal predictive distributions, generates meaningful dispersion when the model is misspecified, and exhibits contracting dispersion as more training data or greater model capacity are provided. For a field-valued partial differential equation surrogate, predictive dispersion strongly aligns with the spatial pattern of prediction errors. On this task, compared with Monte Carlo dropout, OTD yields more accurate predictions and better-calibrated, substantially narrower intervals. On real-world regression benchmarks, it further shows competitive accuracy and better probabilistic predictions compared to established baselines. OTD therefore offers a way to learn structured predictive uncertainty without explicit posterior inference or ensembles of independently trained predictors.
Figures & tables
| Predictive performance | PICP % (Width) | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Method | RMSE | MAE | MACE | 50% | 75% | 80% | 90% | 95% | |
| Deterministic | 0.163 | 0.070 | 0.070 | – | – | – | – | – | – |
| Dropout | 0.172 | 0.102 | 0.071 | 0.016 | 50.5 (0.068) | 74.0 (0.115) | 78.4 (0.128) | 87.5 (0.164) | 92.7 (0.193) |
| OTD | 0.128 | 0.044 | 0.031 | 0.009 | 50.9 (0.025) | 76.7 (0.044) | 81.5 (0.050) | 90.5 (0.066) | 94.9 (0.081) |
| Model | RMSE | MAE | MACE | |
|---|---|---|---|---|
| MC Dropout | ||||
| MC Dropout + ES | ||||
| OTD-0 | ||||
| OTD, | ||||
| OTD |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | Value |
|---|---|
| Latent mask | |
| Concrete parameter | |
| Concrete temperature | |
| Latent flow | |
| Velocity hidden widths | |
| Activation | GeLU |
| Hyperparameter | Value |
|---|---|
| Latent mask | |
| Concrete parameter | |
| Concrete temperature | |
| Latent flow | |
| Velocity hidden widths | |
| Activation | GeLU |
| Hyperparameter | Value |
|---|---|
| Latent mask | |
| Mask dimension | |
| Concrete parameter | |
| Concrete temperature | |
| Latent flow | |
| Velocity hidden widths | |
| Hyperparameter | Value |
|---|---|
| Hidden-layer depth | |
| Reference hidden width | |
| Deterministic hidden widths | |
| Vanilla-dropout hidden widths | |
| OT-dropout hidden widths | |
| Activation | GeLU |
| Sample | Mean error | Std. dev. | |||
|---|---|---|---|---|---|
| 00000 | 0.0242 | 0.0253 | 0.0116 | 0.0167 | 0.0268 |
| 00011 | 0.0196 | 0.0232 | 0.0080 | 0.0114 | 0.0243 |
| 00036 | 0.0192 | 0.0229 | 0.0077 | 0.0106 | 0.0228 |
| 00050 | 0.0257 | 0.0277 | 0.0118 | 0.0167 | 0.0274 |
| 00058 | 0.0208 | 0.0253 | 0.0082 | 0.0109 | 0.0265 |
| 00077 | 0.0257 | 0.0288 | 0.0113 | 0.0171 | 0.0286 |
| Mean error | Std. dev. | |||
|---|---|---|---|---|
| 1 | 0.9765 | – | 0.9735 | |
| 0.9765 | 1 | 0.9441 | – | |
| Mean error | – | 0.9441 | 1 | 0.9294 |
| Std. dev. | 0.9735 | – | 0.9294 | 1 |
| Sample | ||||
|---|---|---|---|---|
| 00000 | 0.6290 | 0.1138 | 0.5881 | 0.1264 |
| 00011 | 0.6909 | -0.0576 | 0.6559 | -0.0404 |
| 00036 | 0.7722 | 0.1702 | 0.7455 | 0.1886 |
| 00050 | 0.8393 | 0.3215 | 0.8049 | 0.3426 |
| 00058 | 0.8659 | 0.3631 | 0.8326 | 0.3810 |
| 00077 | 0.9168 | 0.4516 | 0.8773 | 0.4826 |
| Mean error | Std. dev. | |||
|---|---|---|---|---|
| 1 | 0.9263 | – | 0.9095 | |
| 0.9263 | 1 | 0.6085 | – | |
| Mean error | – | 0.6085 | 1 | 0.5926 |
| Std. dev. | 0.9095 | – | 0.5926 | 1 |
| RMSE | MAE | MACE | ||
|---|---|---|---|---|
| RMSE | MAE | MACE | ||
|---|---|---|---|---|
| Dataset | CADET | CART-h | CART-k | CKDE | LSCDE | NKDE | CDE-HT | LinCDE | MDN | NF | CPFN | OTD |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| energy | 3.55 | 3.09 | 3.06 | 2.47 | 3.38 | 3.00 | 2.93 | 2.93 | 2.78 | 2.86 | 2.73 | 1.08 |
| synchronous | -2.93 | -1.63 | -1.86 | -3.59 | -1.25 | -1.57 | -2.11 | -1.85 | -2.94 | -2.64 | -4.48 | -5.02 |
| localization | -0.23 | -0.55 | -0.01 | -0.26 | -0.61 | -0.28 | -0.66 | -0.95 | -0.68 | -0.43 | 4.19 | 1.03 |
| toxicity | 1.80 | 1.50 | 1.38 | 1.32 | 1.34 | 1.55 | 1.53 | 1.29 | 1.24 | 1.23 | 2.13 | 1.40 |
| concrete | 4.17 | 3.75 | 3.93 | 3.32 | 3.66 | 3.91 | 3.72 | 3.47 | 2.97 | 3.18 | 3.40 | 3.07 |
| slump | 3.42 | 3.55 | 3.43 | 2.35 | 2.91 | 3.08 | 3.34 | 2.98 | 2.23 | 2.39 | 4.48 | 2.40 |
| Dataset | RMSE | Sharpness | MACE | |
|---|---|---|---|---|
| energy | 0.857 | 0.464 | 1.231 | 0.323 |
| synchronous | 0.001 | 0.001 | 0.006 | 0.064 |
| localization | 0.829 | 0.395 | 1.866 | 0.142 |
| toxicity | 1.028 | 0.569 | 1.605 | 0.278 |
| concrete | 5.172 | 2.965 | 6.830 | 0.355 |
| slump | 2.345 | 1.613 | 0.721 | 0.690 |
| Method | RMSE | MAE | MACE | Sharpness | |||
|---|---|---|---|---|---|---|---|
| OTD | 1536 | ||||||
| Low-latent OTD | 512 | ||||||
| Low-latent OTD | 128 | ||||||
| Low-latent OTD | 32 | ||||||
| Conditional OT | 1536 | ||||||
| Conditional OT | 512 |