Uncertainty-aware explanation methods often produce several alternatives for the same prediction. Selecting among them requires a policy for balancing prediction confidence, uncertainty, and application constraints. This paper presents a framework for applying such policies to a fixed set of generated explanations. Candidates are characterised by uncertainty change, prediction direction, and, when available, interval position relative to a decision boundary. The framework combines these properties with eligibility rules, optional bidirectional Pareto screening, and policy-aware ranking. A fictitious prostate-cancer example illustrates how different explanatory purposes lead to different selections from the same candidate set. We instantiate the framework with Calibrated Explanations for classification, thresholded regression, and plain regression. Across 41 benchmark datasets, mean candidate counts range from 11.57 to 21.75 for single-feature explanations and from 29.48 to 69.53 when conjunctions are included. Equal-weight and confidence-only policies yield an average selection-disagreement rate of 28.7% while favouring the same confidence direction. A supporting δ-CLUE experiment demonstrates use with a second generator. By making the selection policy explicit, the framework allows applications to compare and prioritise explanations according to their intended use.
Figures & tables
Approach
Primary selection stage
Relation to the present framework
Explanation Sets [ 5 ]
Representation and restriction
Unifies counterfactual and semifactual explanations and supports restrictions. Not organised around candidate-level uncertainty trade-offs.
CLUE family [ 1 , 9 , 10 ]
During generation
Searches for uncertainty-aware alternatives. Multiple returned candidates can still require an external policy.
Stępka et al. [ 21 ]
After generation
Closest selection precedent. Applies multi-criteria filtering and ideal-point selection to counterfactuals from an ensemble of explainers.
Proposed framework
After generation
Separates uncertainty change, threshold ambiguity, confidence direction, eligibility, and policy across compatible alternative-explanation records.
Table 1 : Positioning against closely related approaches.
Figure 1 : The framework operates after generation. It preserves candidate records and makes application-specific selection choices explicit.
Property
Condition
Interpretation
Ensured
Uj<U0 . Optionally Uj≤U0−ϵU with ϵU>0
Strict reduction in the specified uncertainty signal.
Uncertainty-non-increasing
Uj≤U0
Includes equal uncertainty, used by the CE diagnostic reported later.
Uncertainty-increasing
Uj>U0
Candidate carries greater uncertainty than the factual reference.
Potential
ℓj≤τ≤hj
Candidate interval spans or touches the supplied boundary.
Counter-oriented
(Cj−τ)(C0−τ)<0
Point prediction crosses to the opposite side of the boundary.
Semi-oriented
(Cj−τ)(C0−τ)>0 and ∣Cj−τ∣<∣C0−τ∣
Remains on the factual side but moves toward the boundary.
Table 2 : Operational candidate properties. “Ensured” concerns the selected uncertainty signal; it is not a guarantee of correctness, calibration after selection, feasibility, or actionability.
Hypothetical alternative
Cj
Interval
Uj
Properties and possible eligibility
a1
Refined prostate-volume measurement yields PSA density <0.10
0.25
[0.20,0.30]
0.10
Ensured and counter-oriented: interval wholly below τ .
a2
Patient’s age is 65 years
0.30
[0.15,0.45]
0.30
Ensured, potential, and counter-oriented: excluded if age changes are prohibited.
a3
MRI re-evaluation yields PI-RADS 5
0.85
[0.80,0.90]
0.10
Ensured and super-oriented: interval wholly above τ .
a4
PI-RADS 5 together with PSA density >0.20
0.90
[0.80,1.00]
0.20
Ensured and super-oriented: higher confidence but greater uncertainty than a3 .
Table 3 : Fictitious prostate-cancer alternatives. The factual record is C0=0.40 , interval [0.20,0.60] , U0=0.40 , and τ=0.35 . All predictive quantities and alternatives are illustrative.
Figure 2 : The worked example separates direction, uncertainty, and eligibility. Candidates a1 and a3 have equal uncertainty but opposite confidence directions; a3 and a4 move in the same direction but trade confidence against uncertainty.
Figure 3 : Exact score paths for the four-candidate example when higher confidence is preferred. The upper envelope selects a3 for 0<w<2/3 and a4 for w>2/3 ; a1 and a3 tie at w=0 .
Provider
Candidate coordinates and qualification
Calibrated Explanations
Task-specific calibrated estimate and associated interval width: main evaluated instantiation.
CLUE / δ -CLUE
Decision-relevant prediction and the generator’s uncertainty signal, such as predictive entropy or variance: δ -CLUE is evaluated in a supporting transfer check.
QUCE
Prediction and a candidate-comparable summary of path uncertainty: path semantics must remain explicit.
Multi-objective or robust recourse
Compatible when records expose a decision-relevant prediction and an explicit uncertainty or sensitivity signal: ordinary validity and proximity objectives alone do not establish the mapping.
Table 4 : Candidate-provider mappings and qualifications. Proposed mappings are interface arguments, not additional transfer experiments.
Task
Prediction coordinate
Uncertainty and directional interpretation
Binary classification
Calibrated support for the decision-relevant class
Width of the associated Venn–Abers output; boundary-relative types when a decision threshold is defined.
Multiclass classification
Calibrated support for a specified target class
Width for that class; decreasing target support does not select a particular competing class.
Thresholded regression
Calibrated estimate for the supplied threshold event
Width of the event-specific bounds; threshold is part of the task definition.
Plain regression
Numeric prediction
Width of a conformal prediction interval; direction is relative to the factual value or an application reference.
Table 5 : Task-specific interpretation of the CE mapping.
Task and mode
Total
Non-increasing
Share (%)
Binary, single
18.12
6.48
35.8
Binary, conjunctive
69.26
22.77
32.9
Multiclass, single
21.75
9.97
45.8
Multiclass, conjunctive
69.53
31.59
45.4
Regression 25th threshold, single
11.82
5.59
47.3
Regression 25th threshold, conjunctive
29.90
13.66
45.7
Table 6 : Candidate multiplicity. Total is the mean number generated. Non-increasing candidates satisfy Uj≤U0 . Conjunctive mode includes two-feature conjunctions.
Policy comparison
Disagree
Mean med. ∣ΔC∣
Mean med. ∣ΔU∣
Equal weight vs confidence only ( 0.5 vs 1 )
28.7%
0.13
0.46
Confidence decrease vs increase ( −0.5 vs 0.5 )
85.5%
0.65
0.18
Equal weight vs uncertainty only ( 0.5 vs 0 )
51.9%
0.53
0.18
Uncertainty only vs confidence only ( 0 vs 1 )
66.4%
0.50
0.47
Table 7 : Policy disagreement averaged over dataset-setting-mode blocks. Coordinate differences are means of blockwise medians over changed selections, using the coordinates employed for ranking.
Figure 4 : Policy sensitivity for a fixed thresholded California Housing candidate set. The panels show the generated set and the top ten records under three policies.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
DS
t
n
f
DS
t
n
f
DS
t
n
f
diabetes
b
768
8
german
b
955
27
kc1
b
1192
21
pc4
b
1343
37
ttt
b
958
27
cars
m
1728
6
cmc
m
1473
9
cool
m
768
8
heat
m
768
8
image
m
2310
19
steel
m
1941
27
vehicle
m
846
18
vowel
m
990
11
wave
m
5000
40
wineR
m
1599
11
wineW
m
4898
11
yeast
m
1484
8
abalone
r
4177
8
Appendix
Table 1 : Datasets used in the diagnostic evaluation. Columns are repeated three times to fit the table: DS = dataset, t = task code (b = binary, m = multiclass, r = regression), n = instances, f = features.
Figure 5 : Feasible confidence–uncertainty regions for probability intervals. Solid lines delimit feasible records; dashed lines separate intervals wholly on one side of τ=1/2 from potential intervals containing it. The feasible region above the dashed boundaries is potential. These are representation boundaries, not Pareto frontiers. The curved guides in Figure 4 use the second geometry.