Localisation-Aware Uncertainty for Pretrained Object Detection
Authors: Charmaine Barker, Daniel Bethell, Simos Gerasimou
Organizations: Department of Computer Science, University of York, UK · Department of Elect. Eng., and Computer Science and Eng., Cyprus University of Technology, Cyprus
Reliable uncertainty estimation is essential for deploying object detectors when distribution/covariate shift and adversarial attacks may occur. Existing approaches often require detector retraining, architectural modification, or repeated inference, which may be infeasible or incur significant overheads. We introduce a lightweight post-hoc evidential meta-model that learns when object localisations should be considered uncertain while keeping the base detector frozen. Our approach automatically identifies localisation-relevant features and uses saliency-guided modification to construct an increasingly challenging curriculum. Detection-level targets combine localisation error, modification level, and prediction instability to guide an evidential meta-model to estimate uncertainty for each predicted bounding box. Our approach requires no changes to the detector and preserves its original localisation outputs. Across adversarial attacks and evaluated strengths, GRACE improves TP-FP AUROC by 22% relative to the strongest comparator in some cases while maintaining in-distribution detection performance.
Figures & tables
Figure 1: Overview of GRACE meta-model approach, showing the Saliency Calibration and Uncertainty Guided Training stages. GRACE extracts salient features and weight maps from a pretrained model in a fully post-hoc manner, then generates a noise-driven curriculum to teach the meta-model when to be uncertain. In the figure, g0,g1,g2 are evidential projection branches from salient layers, and u is the predicted uncertainty from the Dirichlet evidence.
Figure 2: Mean detection IoU across equal-count increasing uncertainty percentile bins within each run for all datasets. Stronger downward trends and more negative mean Spearman uncertainty–IoU correlations ( rs ) indicate better uncertainty–localisation alignment.
Method
TP-FP AUROC ↑
FP AUPRC ↑
Risk-coverage AUC ↓
IoU-coverage AUC ↑
Uncertainty-IoU correlation ↓
Added inference time (ms) ↓
CelebA
GRACE
0.83±0.02
0.72±0.05
0.04±0.01
0.69±0.01
−0.36±0.07
2.66±0.48
EMM
0.50±0.03
0.12±0.02
0.10±0.01
0.64±0.01
0.00±0.05
11.76±0.75
ModelNet
0.42±0.18
0.39±0.18
0.20±0.05
0.59±0.04
0.12±0.20
42.93±3.44
MetaDetect
0.74±0.00
0.47±0.00
0.05±0.00
0.67±0.00
−0.35±0.00
27.00±15.77
COCO
Table 1: Uncertainty performance across four datasets: TP-FP AUROC and FP AUPRC measure incorrect-detection identification; AURC and IoU-coverage AUC, selective prediction; uncertainty-IoU correlation, Spearman alignment with localisation quality; and added inference time, overhead relative to the frozen detector. Best results are bold and highlighted .
Figure 3: CelebA between uncertainty assigned to clean detections and their localisation error under five perturbations at severity t=8 . Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation.
Figure 4: CelebA disappearance-prediction AUROC using uncertainty assigned to clean detections under five perturbations at severity t=8 . Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation.
Figure 5: CelebA TP–FP AUROC under five adversarial attacks across increasing attack strength ϵ . Lines show mean performance across five runs for all methods; higher AUROC indicates better discrimination between TP and FP detections under adversarial perturbation.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
mAP 50:95 ↑
AP 50 ↑
AP 75 ↑
Precision 50 ↑
Recall 50 ↑
Median best IoU ↑
CelebA
0.40±0.00
0.99±0.00
0.04±0.00
0.91±0.00
0.99±0.00
0.70±0.00
COCO
0.29±0.00
0.41±0.00
0.31±0.00
0.74±0.00
0.44±0.00
0.84±0.00
TT100K
0.52±0.00
0.70±0.00
0.62±0.00
0.60±0.00
0.84±0.00
0.85±0.00
VisDrone
0.29±0.00
0.52±0.00
0.27±0.00
0.51±0.00
0.60±0.00
0.67±0.00
Appendix
Table 2: Frozen detector performance across the four evaluation datasets. mAP 50:95 reports average precision across IoU thresholds from 0.50 to 0.95, while AP 50 and AP 75 report performance at fixed IoU thresholds. Precision 50 and recall 50 are calculated at an IoU threshold of 0.50. Median best IoU describe the localisation overlap between predicted detections and their best-matching ground-truth boxes.
Method
FPR at 95% TPR ↓
Brier score ↓
SSC ↓
Pearson uncertainty-IoU correlation ↓
CelebA
GRACE
0.281±0.062
0.284±0.177
0.421±0.235
−0.773±0.031
EMM
0.907±0.026
0.797±0.035
0.844±0.021
−0.003±0.075
ModelNet
0.647±0.185
0.783±0.033
0.846±0.023
0.324±0.300
MetaDetect
0.483±0.000
0.084±0.000
0.028±0.000
−0.524±0.000
COCO
Appendix
Table 3: Detection uncertainty calibration and alignment across four datasets. FPR at 95% TPR measures the false-positive rate required to retain 95% of correct detections. Brier score and success-score calibration error (SSC) measure the calibration of predicted detection success probabilities. Pearson uncertainty-IoU correlation measures linear alignment between uncertainty and localisation quality, where a more negative correlation is better. Results are mean ± standard deviation over up to five runs. Best results are bold and highlighted .
Figure 6: Mean detection IoU across uncertainty percentile bins for CelebA. Detections are ordered from lowest to highest assigned uncertainty within each run and divided into equal-count bins. rs reports the mean Spearman uncertainty–IoU correlation, where a stronger downward trend and more negative rs indicate better uncertainty–localisation alignment.
Figure 7: Mean detection IoU across uncertainty percentile bins for Coco. Detections are ordered from lowest to highest assigned uncertainty within each run and divided into equal-count bins. rs reports the mean Spearman uncertainty–IoU correlation, where a stronger downward trend and more negative rs indicate better uncertainty–localisation alignment.
Figure 8: Mean detection IoU across uncertainty percentile bins for TT100K. Detections are ordered from lowest to highest assigned uncertainty within each run and divided into equal-count bins. rs reports the mean Spearman uncertainty–IoU correlation, where a stronger downward trend and more negative rs indicate better uncertainty–localisation alignment.
Figure 9: Mean detection IoU across uncertainty percentile bins for VisDrone. Detections are ordered from lowest to highest assigned uncertainty within each run and divided into equal-count bins. rs reports the mean Spearman uncertainty–IoU correlation, where a stronger downward trend and more negative rs indicate better uncertainty–localisation alignment.
Method
Added time (ms) ↓
Added time (%) ↓
Time per box (ms) ↓
End-to-end FPS ↑
CelebA
GRACE
2.66±0.48
11.66±0.67
2.616±0.479
41.84±0.65
EMM
11.76±0.75
25.19±1.26
11.555±0.749
25.69±1.46
ModelNet
42.93±3.44
118.02±0.64
42.071±3.381
14.81±1.01
MetaDetect
27.00±15.77
78.67±20.44
26.476±15.493
23.50±5.44
COCO
Appendix
Table 4: Computational efficiency across four datasets. Added time reports the latency introduced by each uncertainty method relative to the frozen detector, both in milliseconds and as a percentage. Time per box measures uncertainty-processing latency per detected object, while end-to-end FPS includes both detector and uncertainty-method inference. Best results are bold and highlighted .
Figure 10: Pearson correlations between detection box size and assigned uncertainty across four datasets. Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation. Values closer to zero indicate less systematic dependence on box size, while more negative correlations indicate greater uncertainty for smaller detections.
Figure 11: COCO TP–FP AUROC under five adversarial attacks across increasing attack strength ϵ . Lines show mean performance across five runs for GRACE, EMM, ModelNet, and MetaDetect under PGD ( L∞ ), PGD ( L2 ), TOG Vanishing, TOG Fabrication, and Square Attack ( L∞ ); higher AUROC indicates better discrimination between true- and false-positive detections under adversarial perturbation.
Figure 12: TT100K TP–FP AUROC under five adversarial attacks across increasing attack strength ϵ . Lines show mean performance across five runs for GRACE, EMM, ModelNet, and MetaDetect under PGD ( L∞ ), PGD ( L2 ), TOG Vanishing, TOG Fabrication, and Square Attack ( L∞ ); higher AUROC indicates better discrimination between true- and false-positive detections under adversarial perturbation.
Figure 13: VisDrone TP–FP AUROC under five adversarial attacks across increasing attack strength ϵ . Lines show mean performance across five runs for GRACE, EMM, ModelNet, and MetaDetect under PGD ( L∞ ), PGD ( L2 ), TOG Vanishing, TOG Fabrication, and Square Attack ( L∞ ); higher AUROC indicates better discrimination between true- and false-positive detections under adversarial perturbation.
Method
Adversarial TP-FP AUROC ↑
Adversarial FP AUPRC ↑
Adversarial risk-coverage AUC ↓
Adversarial IoU-coverage AUC ↑
CelebA
GRACE
0.825±0.020
0.850±0.025
0.555±0.009
0.400±0.005
EMM
0.519±0.028
0.661±0.009
0.641±0.010
0.317±0.009
ModelNet
0.501±0.241
0.715±0.090
0.692±0.050
0.306±0.067
MetaDetect
0.676±0.000
0.757±0.000
0.589±0.000
0.340±0.000
COCO
Appendix
Table 5: Uncertainty performance under adversarial attacks across four datasets. Results aggregate performance across PGD ( L∞ ), PGD ( L2 ), TOG Vanishing, TOG Fabrication, and Square Attack ( L∞ ) and their tested attack strengths within each run, and report mean ± standard deviation across runs. TP-FP AUROC and FP AUPRC evaluate identification of unreliable detections, while risk-coverage AUC and IoU-coverage AUC evaluate selective prediction and retained localisation quality under adversarial perturbation. Best results are bold and highlighted .
Figure 14: CelebA correlations between uncertainty assigned to clean detections and their localisation error under five perturbations at different severity levels. Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation; higher Pearson r is better.
Figure 15: Coco correlations between uncertainty assigned to clean detections and their localisation error under five perturbations at different severity levels. Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation; higher Pearson r is better.
Figure 16: TT100K correlations between uncertainty assigned to clean detections and their localisation error under five perturbations at different severity levels. Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation; higher Pearson r is better.
Figure 17: VisDrone correlations between uncertainty assigned to clean detections and their localisation error under five perturbations at different severity levels. Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation; higher Pearson r is better.
Figure 18: CelebA disappearance-prediction AUROC using uncertainty assigned to clean detections under five perturbations at different severity levels. Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation; a higher AUROC is better.
Figure 19: Coco disappearance-prediction AUROC using uncertainty assigned to clean detections under five perturbations at different severity levels. Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation; a higher AUROC is better.
Figure 20: TT100K disappearance-prediction AUROC using uncertainty assigned to clean detections under five perturbations at different severity levels. Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation; a higher AUROC is better.
Figure 21: VisDrone disappearance-prediction AUROC using uncertainty assigned to clean detections under five perturbations at different severity levels.. Violins show the distribution across five runs, dots individual runs, diamonds means, solid horizontal lines medians, and the dashed line zero correlation; a higher AUROC is better.
Figure 22: Ablation of curriculum length T and corruption rate γ on the COCO dataset. We report TP–FP AUROC, risk–coverage AUC, uncertainty–IoU correlation, and uncertainty–perturbation correlation for a combination of T and γ .
Figure 23: Noise-driven curriculum schedules for varying curriculum lengths T and corruption rate γ . Each curve shows the target corruption level ( st=1−e−γt ) across curriculum steps t .
Figure 24: Ablation of the saliency coverage threshold η on COCO. We report TP–FP AUROC, risk–coverage AUC, IoU–coverage AUC, and added inference time per detected box.
Figure 25: Example images from the datasets used in our experimental evaluation.
Reliable uncertainty estimation is essential for deploying object detectors in autonomous systems operating in uncertain environments. Evidential Deep Learning (EDL) provides a principled framework for uncertainty-aware classification by representing network outputs as evidence and interpreting predictions through subjective logic. However, existing evidential object detectors typically combine evidential classification with regression uncertainty models that do not share the same theoretical foundation. In this work, we propose an evidential version of YOLOv8 in which both classification and bounding-box regression are formulated within a common evidential framework. Our approach exploits YOLOv8's distribution-based bounding-box representation, allowing the evidential formulation to be applied not only to classification but also to localisation. As a result, both tasks produce belief, uncertainty, and probability estimates that can be interpreted within the Dempster--Shafer framework. Experiments on KITTI, MUSES, and nuScenes show that the resulting detector remains broadly competitive with standard YOLOv8 in terms of detection accuracy while providing a localisation uncertainty that effectively discriminates between correct and erroneous detections. Moreover, this uncertainty becomes increasingly discriminative under domain shift.
Simon Barbarit-Gaboriau, Hind Laghmara, Rémi Boutteau +1
LITIS - STI, INSA Rouen Normandie · INSA Rouen Normandie, Univ Rouen Normandie, Univ Le Havre Normandie, Normandie Univ, LITIS UR 4108, F-76000 Rouen, France. · LITIS - STI +1
Reliable uncertainty estimation for 3D object detection is critical for deploying safe autonomous systems, yet modern detectors remain poorly calibrated, especially under distribution shifts. Although post-hoc calibration methods address this issue and provide improved calibration for in-distribution tests, they fail to adapt in distribution-shifted scenarios. In this work, we address this issue and introduce a density-aware calibration method that couples post-hoc calibrators with the feature density of latent object queries from DETR-style 3D object detectors. These queries form a compact, location and class-aware feature, ideal for density estimation, allowing our approach to adjust model confidences in distribution-shift scenarios. By fitting a density estimator on these query features, our approach jointly recalibrates both classification and bounding box regression uncertainties. On both a multi-view camera and LiDAR-based detector, our approach consistently outperforms standard post-hoc methods in both in-distribution and distribution-shifted scenarios. Code available https://tillbeemelmanns.github.io/query2uncertainty/ .
Till Beemelmanns, Alexey Nekrasov, Stefan Vilceanu +4
Institute for Automotive Engineering, RWTH Aachen · Computer Vision Institute, RWTH Aachen
Object detection is a safety-critical component of autonomous driving. It is essential to quantify the uncertainty in bounding-box predictions for safety assurance. Post hoc uncertainty quantification without retraining aligns with real-world deployment requirements; therefore, we employ the Laplace approximation. Because instance-level uncertainty is needed, linearized inference methods that require multiple backpropagations are not time-efficient, and sampling-based methods are not fully post hoc. We propose Monte-Carlo generalized linearized model (MC-GLM), which provides instance-level and approximately post hoc uncertainty quantification. The number of samples required in the Monte Carlo step is constant and independent of the number of output instances, so it can be parallelized. Experiments on the nuScenes dataset with the CenterPoint detector validate the effectiveness of our method, and the resulting uncertainties exhibit good quality.
Chongzhe Zhang, Zifan Zeng, Qunli Zhang +2
RAMS Lab, Huawei Heisenberg Research Center, Munich, Germany · Technische Universität Berlin, Berlin, Germany · Technische Universität München, Munich, Germany