Semantic segmentation networks operate on a fixed set of classes and therefore fail when out-of-distribution (OOD) objects appear during deployment, a critical limitation for safety-critical applications such as autonomous driving. Reliably identifying OOD objects requires well-calibrated epistemic uncertainty, yet common softmax-based confidence scores remain overconfident, while Bayesian alternatives such as Monte Carlo dropout or deep ensembles require costly repeated forward passes. Evidential Deep Learning (EDL) offers an efficient alternative by modeling class probabilities as a Dirichlet distribution learned from a single deterministic forward pass. Existing EDL formulations rely on Euclidean objectives that push predictions towards the simplex vertices, encouraging overconfidence rather than preserving uncertainty for unfamiliar inputs. We instead employ Wasserstein-based objectives, which respect the geometry of the probability simplex, and study the influence of the Wasserstein order on segmentation accuracy and OOD detection within a unified evidential framework. We evaluate this framework on a convolutional (DeepLabV3+) and a transformer-based (SegFormer) architecture on the SegmentMeIfYouCan benchmark, including LostAndFound, RoadObstacle21, RoadAnomaly21, and Fishyscapes. Our results show the optimal Wasserstein order is architecture-dependent: second-order objectives dominate on the convolutional backbone, third-order objectives on the transformer backbone, and our framework surpasses comparable baselines on most metrics, with a single deterministic forward pass.
Figures & tables
Figure 1: Top: Semantic segmentation by a deep neural network. Bottom: Evidential uncertainty heatmap obtained by our method. Brighter colors correspond to higher uncertainty.
Figure 2: Illustration of probability mass transport for support extension.
Figure 3: Illustration of the MSE effect on the probability simplex.
Figure 4: Illustration of the Wasserstein effect on the probability simplex.
Figure 5: Left: Semantic segmentation prediction. Right: OOD heatmap obtained by our method using DeeplabV3+, W2 , 0.45 MSE and 0.75 Dice. The top images are from the LostAndFound dataset (a kid playing on the street next to a pile of rubble) and the bottom images from RoadAnomaly21 (a car with an attached caravan).
LostAndFound test-NoKnown
RoadObstacle21
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
Ensemble
2.9
82.0
6.7
7.6
2.7
1.1
77.2
8.6
4.7
1.3
MC Dropout
36.8
35.6
17.4
34.7
13.0
4.9
50.3
5.5
5.8
1.1
Maximum Softmax
30.1
33.2
14.2
62.2
10.3
15.7
16.6
19.7
15.9
6.3
Entropy
47.1
21.6
30.7
42.1
30.2
28.4
26.7
14.7
20.5
9.7
PGN
69.3
9.8
50.0
44.8
45.4
16.5
19.7
19.5
14.8
7.4
Table 1: OOD segmentation benchmark results for the LostAndFound and RoadObstacle21 dataset.
Fishyscapes LostAndFound
RoadAnomaly21
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
Ensemble
0.3
90.4
3.1
1.1
0.4
17.7
91.1
16.4
20.8
3.4
MC Dropout
14.4
47.8
4.8
18.1
4.3
28.9
69.5
20.5
17.3
4.3
Maximum Softmax
5.6
40.5
3.5
9.5
1.8
28.0
72.1
15.5
15.3
5.4
Entropy
15.8
36.1
7.5
16.3
8.6
30.0
73.0
17.8
15.6
5.1
PGN
26.9
36.6
14.8
29.6
16.5
42.8
56.4
25.8
21.8
9.7
Table 2: OOD segmentation benchmark results for the Fishyscapes LostAndFound and RoadAnomaly21 dataset.
DeepLabV3+
SegFormer
Baseline
80.09
84.00
W1
67.67−15.55+9.43
80.61−0.24+0.27
W2
56.14−5.01+6.53
80.60−0.44+0.38
W3
66.88−11.04+9.46
80.54−0.36+0.20
Table 3: In-distribution evaluation on the Cityscapes dataset using mIoU. The rows correspond to the three Wasserstein orders W1 , W2 , and W3 , and the two columns to the DeepLabV3+ and SegFormer backbones. +a and −b describe the size to the largest and smallest deviation respectively.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
LostAndFound test-NoKnown
RoadObstacle21
λW1
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
25.7
56.8
25.9
27.7
14.6
1.6
62.1
18.5
5.6
2.0
1.00
0.00
0.00
0.00
47.5
27.4
21.5
39.2
19.7
4.1
66.0
6.9
8.5
2.0
1.00
0.00
0.00
0.40
50.4
25.8
21.2
38.0
19.2
30.7
42.3
11.6
26.4
8.7
1.00
0.00
0.00
0.45
53.2
23.2
25.3
37.7
20.7
20.2
31.2
11.0
18.3
5.8
1.00
0.00
0.00
0.50
44.4
30.5
25.2
42.8
23.6
12.2
32.6
3.6
34.0
3.4
Appendix
Table A.1: OOD segmentation results for LostAndFound and RoadObstacle21 using DeepLabV3+ with Wasserstein order W1 .
LostAndFound test-NoKnown
RoadObstacle21
λW2
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
25.7
56.8
25.9
27.7
14.6
1.6
62.1
18.5
5.6
2.0
1.00
0.00
0.00
0.00
43.8
38.9
18.9
53.9
21.6
1.1
44.5
14.1
3.6
1.7
1.00
0.00
0.00
0.40
50.4
36.9
18.2
45.3
20.3
13.7
53.9
8.0
15.2
3.3
1.00
0.00
0.00
0.45
46.1
47.6
22.7
53.2
26.3
13.5
40.9
7.2
20.6
4.4
1.00
0.00
0.00
0.45
48.8
38.9
17.0
44.4
18.4
17.2
49.9
6.4
13.7
3.0
Appendix
Table A.2: OOD segmentation results for LostAndFound and RoadObstacle21using DeepLabV3+ with Wasserstein order W2 .
LostAndFound test-NoKnown
RoadObstacle21
λW3
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
25.7
56.8
25.9
27.7
14.6
1.6
62.1
18.5
5.6
2.0
1.00
0.00
0.00
0.00
34.0
88.1
13.8
40.1
17.0
12.0
74.1
8.4
18.6
5.2
1.00
0.00
0.00
0.40
10.5
93.2
1.3
12.0
1.2
13.7
50.6
4.7
19.3
3.8
1.00
0.00
0.00
0.45
12.6
89.5
8.4
23.5
8.0
2.8
63.1
5.0
6.4
1.9
1.00
0.00
0.00
0.50
11.7
84.9
2.1
10.6
1.6
5.4
73.7
2.9
6.0
1.4
Appendix
Table A.3: OOD segmentation results for LostAndFound and RoadObstacle21using DeepLabV3+ with Wasserstein order W3 .
Fishyscapes LostAndFound
RoadAnomaly21
λW1
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
25.7
56.8
25.9
27.7
14.6
1.6
62.1
18.5
5.6
2.0
1.00
0.00
0.00
0.00
47.5
27.4
21.5
39.2
19.7
4.1
66.0
6.9
8.5
2.0
1.00
0.00
0.00
0.40
50.4
25.8
21.2
38.0
19.2
30.7
42.3
11.6
26.4
8.7
1.00
0.00
0.00
0.45
53.2
23.2
25.3
37.7
20.7
20.2
31.2
11.0
18.3
5.8
1.00
0.00
0.00
0.50
44.4
30.5
25.2
42.8
23.6
12.2
32.6
3.6
34.0
3.4
Appendix
Table A.4: OOD segmentation results for Fishyscapes and RoadAnomaly21 using DeepLabV3+ with Wasserstein order W1 .
Fishyscapes LostAndFound
RoadAnomaly21
λW2
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
25.7
56.8
25.9
27.7
14.6
1.6
62.1
18.5
5.6
2.0
1.00
0.00
0.00
0.00
4.6
53.6
7.1
12.6
9.2
33.0
60.3
21.1
19.5
5.0
1.00
0.00
0.00
0.40
10.1
38.4
4.6
20.8
4.2
50.3
57.9
20.0
22.6
5.6
1.00
0.00
0.00
0.45
2.5
42.0
11.0
11.8
7.3
42.9
59.2
18.9
23.5
6.0
1.00
0.00
0.00
0.50
3.9
46.5
8.5
10.4
6.9
35.6
65.2
23.4
20.3
6.8
Appendix
Table A.5: OOD segmentation results for Fishyscapes and RoadAnomaly21 using DeepLabV3+ with Wasserstein order W2 .
Fishyscapes LostAndFound
RoadAnomaly21
λW3
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
25.7
56.8
25.9
27.7
14.6
1.6
62.1
18.5
5.6
2.0
1.00
0.00
0.00
0.00
7.7
53.3
6.2
23.6
20.7
36.0
85.0
16.5
23.8
6.5
1.00
0.00
0.00
0.40
2.7
68.3
1.6
13.3
2.4
40.6
67.2
17.0
26.8
6.3
1.00
0.00
0.00
0.45
2.2
54.6
8.9
9.7
7.8
42.3
76.8
20.7
25.5
7.4
1.00
0.00
0.00
0.50
2.3
52.5
5.8
10.8
4.7
40.2
76.3
15.2
28.3
6.0
Appendix
Table A.6: OOD segmentation results for Fishyscapes and RoadAnomaly21 using DeepLabV3+ with Wasserstein order W3 .
LostAndFound test-NoKnown
RoadObstacle21
λW1
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
22.4
63.1
7.0
13.1
26.9
21.2
64.7
4.2
55.7
3.9
1.00
0.00
0.00
0.00
16.1
57.7
4.5
11.6
1.3
8.9
66.1
3.0
13.6
0.8
1.00
0.00
0.00
0.40
41.9
43.9
23.3
25.5
13.1
29.1
48.6
13.0
20.9
5.2
1.00
0.00
0.00
0.45
34.1
39.8
22.4
19.2
9.7
31.5
35.8
9.8
28.6
5.8
1.00
0.00
0.00
0.50
34.1
49.8
16.7
23.8
9.2
24.3
39.0
11.0
13.8
3.0
Appendix
Table A.7: OOD segmentation results for LostAndFound and RoadObstacle21 using SegFormer with Wasserstein order W1 .
LostAndFound test-NoKnown
RoadObstacle21
λW2
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
22.4
63.1
7.0
13.1
26.9
21.2
64.7
4.2
55.7
3.9
1.00
0.00
0.00
0.00
23.2
65.7
10.3
12.8
3.7
15.1
46.5
3.3
24.3
1.5
1.00
0.00
0.00
0.40
21.0
62.9
7.3
14.6
2.9
18.9
67.9
7.7
11.6
2.2
1.00
0.00
0.00
0.45
8.8
79.9
3.2
8.1
0.9
9.4
81.0
3.0
40.9
2.7
1.00
0.00
0.00
0.50
34.8
62.9
17.8
20.5
9.1
27.7
54.7
10.0
24.6
4.7
Appendix
Table A.8: OOD segmentation results for LostAndFound and RoadObstacle21 using SegFormer with Wasserstein order W2 .
LostAndFound test-NoKnown
RoadObstacle21
λW3
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
22.4
63.1
7.0
13.1
26.9
21.2
64.7
4.2
55.7
3.9
1.00
0.00
0.00
0.00
33.0
63.0
13.2
21.6
6.9
20.5
54.1
6.1
27.9
3.4
1.00
0.00
0.00
0.40
38.8
59.7
20.0
30.7
14.1
42.1
57.8
23.3
26.7
11.8
1.00
0.00
0.00
0.45
21.1
64.2
10.1
14.5
4.1
20.9
64.6
10.0
14.8
3.1
1.00
0.00
0.00
0.50
35.5
47.5
18.7
19.3
8.9
29.5
34.6
14.5
24.1
6.8
Appendix
Table A.9: OOD segmentation results for LostAndFound and RoadObstacle21 using SegFormer with Wasserstein order W3 .
Fishyscapes LostAndFound
RoadAnomaly21
λW1
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
32.6
70.1
14.4
49.7
14.4
48.2
96.4
30.1
29.5
8.9
1.00
0.00
0.00
0.00
28.9
67.1
8.9
40.3
11.7
49.4
86.6
32.2
27.9
9.9
1.00
0.00
0.00
0.40
50.4
25.8
21.2
38.0
19.2
30.7
42.3
11.6
26.4
8.7
1.00
0.00
0.00
0.45
53.2
23.2
25.3
37.7
20.7
20.2
31.2
11.0
18.3
5.8
1.00
0.00
0.00
0.50
44.4
30.5
25.2
42.8
23.6
12.2
32.6
3.6
34.0
3.4
Appendix
Table A.10: OOD segmentation results for Fishyscapes and RoadAnomaly21 using SegFormer with Wasserstein order W1 .
Fishyscapes LostAndFound
RoadAnomaly21
λW2
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
32.6
70.1
14.4
49.7
14.4
48.2
96.4
30.1
29.5
8.9
1.00
0.00
0.00
0.00
28.2
72.2
6.8
32.0
8.1
48.9
85.1
33.5
24.4
8.1
1.00
0.00
0.00
0.40
35.6
55.0
6.9
63.3
11.5
50.3
87.3
31.0
26.9
8.6
1.00
0.00
0.00
0.45
16.2
73.0
2.7
31.0
3.5
44.6
75.0
35.3
26.0
10.6
1.00
0.00
0.00
0.50
36.5
70.5
8.2
53.7
11.9
59.6
79.4
37.8
28.4
12.0
Appendix
Table A.11: OOD segmentation results for Fishyscapes and RoadAnomaly21 using SegFormer with Wasserstein order W2 .
Fishyscapes LostAndFound
RoadAnomaly21
λW3
λDice
λKL
λMSE
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
AuPRC ↑
FPR 95 ↓
sIoU↑
PPV↑
F1↑
0.00
0.00
0.00
1.00
32.6
70.1
14.4
49.7
14.4
48.2
96.4
30.1
29.5
8.9
1.00
0.00
0.00
0.00
28.4
76.0
9.1
49.1
12.2
57.5
84.4
34.3
28.9
10.6
1.00
0.00
0.00
0.40
37.3
71.2
12.7
54.4
18.8
64.5
72.4
35.7
31.7
13.6
1.00
0.00
0.00
0.45
22.4
69.0
8.4
58.5
13.6
42.4
90.1
30.3
26.3
8.3
1.00
0.00
0.00
0.50
38.0
60.3
10.3
66.7
15.5
56.8
85.0
36.6
28.6
11.4
Appendix
Table A.12: OOD segmentation results for Fishyscapes and RoadAnomaly21 using SegFormer with Wasserstein order W3 .
Reliable semantic segmentation for mobile robots requires both accurate dense prediction and robust uncertainty estimation under distribution shift. Strong uncertainty baselines such as Monte Carlo Dropout often require repeated stochastic forward passes and are difficult to deploy on edge platforms. We propose Energy-Aware NECO, a single-pass pixel-wise out-of-distribution (OOD) detector for semantic segmentation. The method combines a centered NECO-style geometric ratio computed from decoder features with a logit-based Energy score. Both components are standardized using statistics fitted on a pure in-distribution validation split and fused through a convex combination. We evaluate the method on the miniMUAD subset using true pixel-level OOD labels. The proposed hybrid score achieves an AUROC of 0.8539, outperforming NECO-only (0.8280), Energy-only (0.8171), and an ensemble predictive-entropy baseline (0.8124). Additional qualitative and operating-point analyses show that the hybrid detector improves overall ranking performance while preserving the efficiency advantages of a single-pass design. Code is available at https://github.com/boyuan-zhangx/Energy-Aware_NECO
Boyuan Zhang, Huanshan Huang, Yifei Cao
Ecole Polytechnique, Institut Polytechnique de Paris, Palaiseau, France. · CIAD, UTBM, Universit´e Marie et Louis Pasteur, France. · U2IS, ENSTA, Institut Polytechnique de Paris, France.
While Deep Neural Networks (DNNs) achieve remarkable performance, their tendency to produce overconfident predictions. Evidential Deep Learning (EDL) mitigates this by formulating predictions as a Dirichlet distribution over class probabilities to explicitly quantify epistemic uncertainty. However, we found that the conventional EDL suffers from two fundamental limitations: a Kullback-Leibler (KL) penalty that only suppresses the evidence of negative classes, producing excessively high evidence therefore decreasing the model's ability to quantify uncertainty, and an absence in theoretical guarantee of setting Dirichlet parameter α=e+1. In this paper, we propose a mathematically principled framework, Variational Inference Evidential Deep Learning (VI-EDL). By reformulating evidential learning through the lens of variational inference, we derive an Evidence Lower Bound (ELBO), which prevents the evidence from growing excessively. Theoretically, we rigorously establish a generalization bound and reveal how the predicted uncertainty, feature and network complexity affect this bound, and why setting α=e+1 can minimize it. Extensive experiments on standard visual and medical datasets demonstrate that VI-EDL achieves state-of-the-art performance, showing excellent performance in out-of-distribution detection, noise detection and autonomous driving scenario. The code is available in https://github.com/seutjw/VI-EDL.
Jiawei Tang, Xinyan Du, Hui Liu +2
School of Computer Science and Engineering Southeast University Nanjing 210096, China · School of Computing Information Sciences Saint Francis University2026 Hong Kong SAR, China · CityDepartmentUniversityof Computerof Hong KongScience Hong Kong SAR, China
Evidential Deep Learning (EDL) has emerged as an efficient, sampling-free strategy for uncertainty estimation. A series of EDL variants have been proposed to address specific limitations of the original framework, achieving notable success. However, the underlying theoretical structure of EDL and the relationships among these variants have received limited systematic investigation. In this work, we establish a principled theoretical foundation for EDL by interpreting it within a generalized Bayesian framework that includes prior specification, posterior update, and training objective. We further characterize evidential uncertainty from a Bayesian distributional uncertainty viewpoint, established via asymptotic analysis. Building on this perspective, we further propose Generalized Evidential Deep Learning (GEDL), a unified and extensible framework that explicitly disentangles the roles of individual components and systematically relates GEDL to existing variants. Extensive experiments demonstrate that GEDL yields comparable results on classification, uncertainty estimation and OOD detections, with theoretical grounding.
Yuanye Liu, Yibo Gao, Yuanyang Chen +1
School of Data Science, Fudan University, Shanghai, China.