Organizations: UMR CNRS 7253, Heudiasyc, Universit´e de Technologie de Compi`egne, Compi`egne, France. · School of Knowledge Science, Japan Advanced Institute of Science and Technology, Nomi, Japan.
This paper tackles visible challenges in deep ensemble learning, where deep neural networks serve as ensemble members: training and storage burdens, and robustness of cautious (set-valued) predictions targeting multiple utilities, which may involve reward-sensitivity. To mitigate the training and storage burdens, we propose to employ compact ensembles, such as Bayesian Neural Networks and Convolutional Neural Networks with the Monte-Carlo dropout prediction option, to produce probabilistic predictions. For each query instance, these probabilistic predictions are then used to define a representative distribution optimizing some statistical distance. The representative distribution is then employed to define the Bayes-optimal prediction (BOP) of any utility. To address the potential unrobustness of singleton prediction making, we propose a family of set-utilities satisfying some desirable properties and whose set-valued BOPs can be found efficiently. Empirical evidence is then given to illustrate the potential (dis)advantages of the proposed ensemble learning framework.
Figures & tables
Figure 1 : Illustration of the conventional ensemble learning framework
Figure 2 : Illustration of compact ensembles.
#
Name
# of instances
image type
# of classes
1
CIFAR-10
60,000
32×32 color images
10
2
Fashion-MNIST
70,000
28×28 gray images
10
3
Plant Leaf Diseases
61,486
256×256 color images
39
Table 1 : Data sets.
Figure 3 : BNNs with clean training data. Results with Uα on clean (noisy) test data are plotted with dashed (thick) lines. Classifiers are color-coded: clf_sE , clf_L1 , and clf_KL .
Figure 4 : BNNs with noisy training data. Results with Uα on clean (noisy) test data are plotted with dashed (thick) lines. Classifiers are color-coded: clf_sE , clf_L1 , and clf_KL .
Figure 5 : The probabilistic prediction p(Y∣x) and set-valued BOP Y^U1.63 ( 4 ) of a few cat images from the CIFAR-10 data set. The chosen value of α is the median reward level.
Figure 6 : The probabilistic prediction p(Y∣x) and set-valued BOPs Y^U2.050/1 and Y^U1.63 ( 4 ) of a clean cat image from the CIFAR-10 data set. The ensembles are trained using clean training data D . The chosen values of α are the median reward levels.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol/Acronym
Meaning
X , x
Instance space, instance
Y , Y , y^,y^u
Output space, set prediction, BOP, BOP with utility u
Y^Uα0/1
Set-valued BOP produced by utility-discounted accuracy function Uα0/1
Xp , Yk
Feature, class variable
Yk
Output space of Yk
YRk\vbox..=Yk∪{Yk}
Set of predictions with rejection
Appendix
Table 2 : Notation and acronyms
Figure 7 : BNNs with clean training data. Results with Uα0/1 on clean (noisy) test data are plotted with dashed (thick) lines. Classifiers are color-coded: clf_sE , clf_L1 , and clf_KL .
Figure 8 : BNNs with noisy training data. Results with Uα0/1 on clean (noisy) test data are plotted with dashed (thick) lines. Classifiers are color-coded: clf_sE , clf_L1 , and clf_KL .
Figure 9 : CNNs with clean training data. Results with Uα on clean (noisy) test data are plotted with dashed (thick) lines. Classifiers are color-coded: clf_sE , clf_L1 , and clf_KL .
Figure 10 : CNNs with noisy training data. Results with Uα on clean (noisy) test data are plotted with dashed (thick) lines. Classifiers are color-coded: clf_sE , clf_L1 , and clf_KL .
Figure 11 : CNNs with clean training data. Results with Uα0/1 on clean (noisy) test data are plotted with dashed (thick) lines. Classifiers are color-coded: clf_sE , clf_L1 , and clf_KL .
Figure 12 : CNNs with noisy training data. Results with Uα0/1 on clean (noisy) test data are plotted with dashed (thick) lines. Classifiers are color-coded: clf_sE , clf_L1 , and clf_KL .
CIFAR-10
Fashion-MNIST
Plan Leaf Diseases
BNNs
CNNs
BNNs
CNNS
BNN
CNNs
c-n
ECEcl
7.99 ± 0.38
8.75 ± 0.86
2.82 ± 0.04
6.62 ± 0.06
2.45 ± 0.14
3.68 ± 0.13
MCEcl
71.11 ± 3.55
70.80 ± 5.51
59.30 ± 11.09
77.87 ± 1.42
97.02 ± 0.34
95.74 ± 0.62
ECEco
26.53 ± 5.06
31.46 ± 5.69
9.50 ± 0.51
5.33 ± 0.62
11.53 ± 2.69
33.20 ± 3.91
MCEco
43.17 ± 10.86
51.74 ± 8.57
50.80 ± 25.43
21.52 ± 6.76
18.59 ± 3.86
61.42 ± 2.99
c-c
ECEcl
0.90 ± 0.04
1.29 ± 0.03
0.40 ± 0.02
3.48 ± 0.08
0.50 ± 0.01
1.52 ± 0.03
Appendix
Table 3 : Calibration errors of the representative distributions ( 8 ) produced by dsE ( 9 ): The shorthand notations ECEcl , MCEcl , ECEco , and MCEco are used instead of ECEclasswise , MCEclasswise , ECEconfidence , and MCEconfidence . Results on (clean train data, clean test data), (clean train data, noisy test data), (noisy train data, clean test data), and (noisy train data, noisy test data) are marked as c-c, c-n, n-c, and n-n, respectively.
CIFAR-10
Fashion-MNIST
Plan Leaf Diseases
BNNs
CNNs
BNNs
CNNS
BNN
CNNs
c-n
ECEcl
9.35 ± 0.64
11.57 ± 0.88
2.59 ± 0.04
5.79 ± 0.13
2.49 ± 0.16
4.23 ± 0.08
MCEcl
81.42 ± 4.37
88.76 ± 2.33
68.29 ± 3.99
81.41 ± 0.89
95.52 ± 0.29
98.22 ± 1.14
ECEco
38.59 ± 5.24
51.72 ± 5.25
8.03 ± 0.54
5.06 ± 0.87
37.21 ± 2.55
62.66 ± 3.21
MCEco
61.86 ± 14.26
68.27 ± 4.53
42.55 ± 29.52
24.81 ± 16.83
50.94 ± 2.12
82.51 ± 3.14
c-c
ECEcl
0.78 ± 0.01
1.07 ± 0.07
0.46 ± 0.02
2.08 ± 0.09
0.07 ± 0.00
0.30 ± 0.00
Appendix
Table 4 : Calibration errors of the representative distributions ( 8 ) produced by dL1 ( 10 ): The shorthand notations ECEcl , MCEcl , ECEco , and MCEco are used instead of ECEclasswise , MCEclasswise , ECEconfidence , and MCEconfidence . Results on (clean train data, clean test data), (clean train data, noisy test data), (noisy train data, clean test data), and (noisy train data, noisy test data) are marked as c-c, c-n, n-c, and n-n, respectively.
CIFAR-10
Fashion-MNIST
Plan Leaf Diseases
BNNs
CNNs
BNNs
CNNS
BNN
CNNs
c-n
ECEcl
9.39 ± 0.62
11.67 ± 0.87
2.46 ± 0.05
5.97 ± 0.12
2.62 ± 0.17
4.25 ± 0.08
MCEcl
82.07 ± 5.07
90.11 ± 1.77
67.95 ± 4.12
82.45 ± 0.81
97.10 ± 2.05
99.27 ± 0.49
ECEco
39.58 ± 5.04
52.73 ± 5.14
7.35 ± 0.51
6.68 ± 0.69
41.77 ± 2.87
65.71 ± 3.71
MCEco
56.61 ± 6.75
68.20 ± 3.29
41.64 ± 30.06
17.77 ± 1.82
57.19 ± 1.94
83.65 ± 2.30
c-c
ECEcl
0.88 ± 0.02
1.15 ± 0.09
0.36 ± 0.02
2.16 ± 0.09
0.06 ± 0.00
0.26 ± 0.01
Appendix
Table 5 : Calibration errors of the representative distributions ( 8 ) produced by dKL ( 11 ): The shorthand notations ECEcl , MCEcl , ECEco , and MCEco are used instead of ECEclasswise , MCEclasswise , ECEconfidence , and MCEconfidence . Results on (clean train data, clean test data), (clean train data, noisy test data), (noisy train data, clean test data), and (noisy train data, noisy test data) are marked as c-c, c-n, n-c, and n-n, respectively.
We introduce an efficient Bayesian deep ensemble method for predictive regression designed to enhance interpretability while maintaining competitive predictive performance and computational efficiency. Our method combines the statistical rigor of Bayesian inference with the scalability of deep ensembles, providing calibrated uncertainty estimates that enable its use not only for standalone prediction but also as a component within broader learning systems. To achieve these goals, our work relies on three key design components: (i) low-dimensional ensemble representation: predictions are expressed as a combination of a small number of trained neural predictors, enabling scalable inference whose cost depends on ensemble size rather than dataset size; (ii) closed-form Bayesian aggregation: ensemble predictions are combined using Bayesian linear regression, yielding interpretable posterior weights and calibrated uncertainty without approximate inference; and (iii) Independent ensemble training: multiple neural networks are trained separately, producing diverse predictive representations that improve robustness and uncertainty calibration. Empirical results on standard regression benchmarks demonstrate that the proposed approach achieves competitive predictive performance while maintaining reliable uncertainty estimates across settings.
Sina Aghaee Dabaghan Fard, Marie Maros, Jaesung Lee
Wm Michael Barnes ’64 Department of Industrial and Systems Engineering Texas A&M University College Station, TX, USA
While deep ensembles are widely considered to be the default method for uncertainty quantification in deep learning, their effectiveness for graph-structured data is often simply assumed based on successes in domains like computer vision. We investigate standard deep ensembles specifically for message-passing graph neural networks. Benchmarking across seven datasets representing varied tasks and complexities, we reveal that ensembles provide surprisingly little improvement over a single model. Instead, the observed marginal gains stem primarily from stabilizing optimization noise in point predictions rather than yielding meaningfully better uncertainty estimates. Through an aleatoric-epistemic decomposition, we identify epistemic collapse: independently trained networks consistently converge to overly similar predictions. Because disagreement is the fundamental mechanism through which ensembles capture epistemic uncertainty, this lack of diversity neutralizes their key advantage. Analyzing this phenomenon further, we suggest this collapse is driven by functional rather than weight-space convexity, where distinct parameter solutions induce almost identical behavior. Our results suggest that deep ensemble success does not seamlessly transfer to graph machine learning.
Pedro C. Vieira, Pedro Ribeiro, Viacheslav Borovitskiy
University of Edinburgh · DCC/FCUP, University of Porto
Parallel ensemble methods were compared on 56 small-to-medium tabular classification tasks drawn from OpenML CC18. A set of ``best practice'' recommendations on the use of ensemble methods was derived from these observations. It was later validated on 28 additional tasks using TabArena's precomputed data, where the recommendation set significantly outperformed Single Best and matched or exceeded individual ensemble methods. Two key observations were made. First, Blending and Stacking are inconsistent, but their inconsistencies are independent and happen on different tasks. Second, while Hard Voting's probabilistic classification is rather weak, a consequence of using vote proportions as posterior estimates, Robust Soft Voting's probabilistic classification is particularly successful, especially in the multiclass case.