Mixed-Prior Decision Risk for Open-Set Recognition
Authors: L. A. Erlygin, P. D. Proskura, A. A. Zaytsev
Organizations: Skolkovo Institute of Science and Technology (Skoltech), Moscow, Russia · Risk Management, Sber, Moscow, Russia · The Institute for Information Transmission Problems (IITP RAS), Moscow, Russia
In open-set recognition (OSR), a probe must either be identified as one of the known gallery classes or rejected as unknown, so three error types coexist: false acceptance, false rejection, and misidentification. An uncertainty score for selective recognition should rank probes by the risk of the decision the system has made. Bayesian gallery-aware models such as Holistic Uncertainty Estimation (HolUE) summarize the posterior over known and unknown classes by Kullback--Leibler (KL) divergence components and map them to an uncertainty score with a supervised nonlinear calibrator. We show that the KL summary is not generally monotone in decision risk: linear fusion of the KL components tuned on validation data yields negative filtering quality on several benchmarks. We propose MPRisk, a mixed-prior posterior decision-risk score that keeps the same Bayesian posterior but directly scores the error events associated with the selected decision: false-acceptance, misidentification, and false-rejection risks, plus a non-specificity penalty for rejections, enabled by modeling unknown identities as a continuous component. Four nonnegative weights tuned on a validation set suffice for ranking; no nonlinear supervised model is required. Across nine image, audio, and text benchmarks, MPRisk achieves the best or tied-best Prediction Rejection Ratio at every operating point on the image and audio benchmarks and on most text operating points, with bootstrap-confirmed gains over HolUE on five benchmarks (up to +0.19 PRR) at comparable or lower runtime.
Figures & tables
Рис. 1: MPRisk overview. (a) Embedding space of a fixed OSR system with K gallery classes (prototypes ∙ , samples ■ ) under the mixed prior: unknown identities form a continuous component c∈(K,K+1] (tinted reject region); dashed lines are decision boundaries; blue shading shows the decision-risk field 1−maxap(a∣x) . Markers show the three error types: false acceptance ( ★ ), misidentification ( ⧫ ), and false rejection ( × ) — a corrupted known probe whose diffuse embedding distribution (green density, low κx ) yields a non-specific unknown posterior, i.e., high N0(x) . (b) Top: the KL summary measures information gain and can invert the risk ordering of two posteriors (A vs. B), which is why KL features require the nonlinear calibration of HolUE. Bottom: MPRisk scores the selected decision with four components rFA , rID , rFR , and rNS=P0N0 ; stacked bars show the components at the three marked probes — for the confidently rejected corrupted probe only rNS flags the error. The final score uλ combines the components with four nonnegative validation-tuned weights.
Dataset
Spearman
Inv. rate
KL AUROC
KL AUPRC
IJB-C
-0.03
0.55
0.45
0.04
Yahoo Answers
-0.47
0.81
0.18
0.15
Таблица 1: Disagreement between the raw (uncalibrated) KL summary and the error-risk ordering. Spearman: rank correlation between the KL summary and the per-probe error indicator. Inv. rate: fraction of (erroneous, correct) probe pairs ordered incorrectly by the KL summary. KL AUROC/AUPRC: any-error detection quality of the raw KL summary used directly as an uncertainty score.
Method
IJB-C
IJB-B
Whale
VB-Eval-L-5
0.05
0.1
0.2
0.05
0.1
0.2
0.05
0.1
0.2
0.01
0.05
0.1
SCF
0.40
0.31
0.23
0.29
0.25
0.22
0.16
0.02
-0.06
0.55
0.26
0.12
AccScr
0.73
0.72
0.66
0.65
0.68
0.62
0.77
0.75
0.66
0.66
0.76
0.72
MSP
0.74
0.75
0.70
0.66
0.70
0.65
0.77
0.77
0.70
0.38
0.88
0.86
Margin
0.74
0.75
0.70
0.66
0.70
0.66
0.77
0.77
0.70
0.68
0.88
0.86
GalUE
0.74
0.74
0.67
0.66
0.69
0.60
0.78
0.76
0.70
0.69
0.89
0.87
Таблица 2: Prediction Rejection Ratios (PRR, ↑ ) for F1 filtering on image, Whale, and audio open-set recognition benchmarks. HolUE uses its validation-trained calibration; MPRisk uses validation-tuned component weights; both receive identical validation data. Best results in bold, second-best underlined.
IJB-C (FPIR 0.05 )
Whale (FPIR 0.1 )
Method
Any
FA
FR
ID
Any
FA
FR
ID
SCF
0.70
0.61
0.77
0.91
0.61
0.55
0.84
0.86
AccScr
0.87
0.91
0.80
0.85
0.88
0.87
0.86
0.81
MSP
0.87
0.91
0.80
0.88
0.88
0.88
0.85
0.88
GalUE
0.88
0.92
0.80
0.90
0.89
0.88
0.86
0.90
HolUE
0.87
0.96
0.70
0.97
0.93
0.93
0.85
0.90
Таблица 10: Error-type detection quality on IJB-C at FPIR 0.05 and Whale at FPIR 0.1 . AUROC is reported for detecting any OSR error (Any), false acceptance (FA), false rejection (FR), and misidentification (ID).
Yahoo Answers (FPIR 0.1 )
PAN-20-AV (FPIR 0.1 )
Method
Any
FA
FR
ID
Any
FA
FR
ID
SCF
0.45
0.66
0.41
–
0.50
0.37
0.56
0.36
AccScr
0.59
0.77
0.54
–
0.61
0.89
0.48
0.83
MSP
0.54
0.76
0.48
–
0.64
0.82
0.53
0.86
GalUE
0.58
0.77
0.53
–
0.61
0.88
0.48
0.90
HolUE
0.89
0.58
0.92
–
0.59
0.75
0.52
0.70
Таблица 11: Error-type detection quality on Yahoo Answers and PAN-20-AV, both at FPIR 0.1 . AUROC is reported for detecting any OSR error (Any), false acceptance (FA), false rejection (FR), and misidentification (ID); misidentification is absent in the Yahoo Answers protocol at this operating point.
Open-set recognition systems face a neglected failure mode: high-confidence near-known unknowns, which lie outside the known label set but are close enough to known classes that a closed-set classifier accepts them with high confidence. We show that this failure is widespread across scalar-threshold methods, including recent post-hoc detectors, and that stronger encoders can amplify rather than remove the risk. We propose EGUR-A, which changes the decision from is this sample's score high enough?'' to does this predicted known class have sufficient evidence to accept this sample?'' EGUR-A combines class-conditional local acceptance evidence with global residual evidence, and selects their relative weight from known-sample statistics without unknown validation data. Across CUB, FGVC-Aircraft, and ImageNet-hard, EGUR-A substantially reduces high-confidence false known acceptance at matched known-rejection operating points. The result is not a stronger threshold; it is a different question: whether a known class is entitled to accept a sample.
Xi Chen, Yingjun Xiao, Gang Fang
School of Computer Science and Cyber Engineering, Guangzhou University, Guangzhou, China · School of Computer Science and Cyber Engineering, Guangzhou University, Guangzhou 510006, China · School of Artificial Intelligence, Guangzhou University, Guangzhou, China +3
Open-set recognition (OSR) requires a classifier to reject inputs from unseen classes which is essential in safety-critical settings such as medical imaging. Simplex based methods, which fix class prototypes at the vertices of a regular simplex and then reject via a distance-ratio score, perform well empirically but lack theoretical justification, and existing analysis applies only when the embedding dimension d is at least C-1, which is the regime in which a regular simplex exists. We give a theoretical account of simplex-ratio OSR that holds in every embedding dimension, including d < C-1. Our analysis centers on balanced equal-norm codes: prototype configurations with equal lengths and zero sum, which exist for all d >= 2 and include the regular simplex as a special case. For these codes we show that an auxiliary squared ratio score has sublevel sets that are exact unions of Euclidean balls, which in turn bracket the acceptance region of the operational score; and we prove a sharp dichotomy: the prototypes attain one-distance symmetry, behaving like a regular simplex, if and only if d >= C-1, with controlled degradation governed by an explicit defect parameter below that threshold. We further show the false-acceptance rate decays exponentially in d under natural isotropy assumptions, and that the operational score is globally Lipschitz with compact acceptance regions. Empirically, we study balanced prototype geometry as both an analytic tool and a representation-learning prior, rather than as a stand-alone state-of-the-art detector. Across CIFAR and MedMNIST open-set splits, the geometry provides useful structure, but OSR performance remains strongly dependent on the scoring rule: raw ratio scores typically underperform nearest-neighbor and logit-based alternatives.
Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-optimal decision strategies for these tasks do not coincide with the scores commonly used to quantify epistemic uncertainty. Building on the epistemic reject-option framework, we evaluate epistemic uncertainty using its ability to identify regret, the reducible error. Formulating selective prediction as a constrained optimization over coverage, expected risk, and regret, we prove the optimal selector is a thresholded convex combination of the ground-truth aleatoric and epistemic uncertainties. This theoretical unification exposes a weakness in recent uncertainty disentanglement literature: we demonstrate that standard correlation metrics between learned components do not necessarily predict their actual operational utility. We instead propose to evaluate the achievable risk, regret, coverage surface of the decomposition as a diagnostic for joint disentanglement and utility. Benchmarking standard methods on datasets with dense human annotations reveals that decision-theoretic rankings can disagree substantially with proxy-task rankings, including pairwise rank inversions between methods that are top-ranked on one criterion and bottom-ranked on other.
Jakub Paplhám, Willem Waegeman, Eyke Hüllermeier +1
Czech Technical University in Prague · Ghent University · LMU Munich, MCML, DFKI