How Much Imprecision is Enough Imprecision in my Classifier? A Practical Elicitation Procedure
Authors: Victor F. Lopes de Souza, Sébastien Destercke, Abdelhak Imoussaten
Organizations: EuroMov Digital Health in Motion, Univ. Montpellier, IMT Mines Al`es, Montpellier, France · LIRMM, Universit´e de Montpellier, CNRS, Montpellier, France · SyCoIA, IMT Mines Al`es, Al`es, France · Heudiasyc, Universit´e de Technologie de Compi`egne, CNRS, Compi`egne, France
Set-valued classifiers, whether derived from precise probabilities and an adapted cost function, from convex sets with a robust inference mechanism, or from conformal methods, are routine options to obtain more robust, trustworthy predictions. However, there is a lack of operational tools to measure how robust or imprecise a given user is ready to be when receiving predictions, that is how much precision he/she is ready to let go in exchange of more accuracy. This is why we propose, in this paper, practical and operational elicitation procedures to measure the user proneness to set-valued predictions. The effectiveness of the iterative elicitation procedure in converging to the target parameter value is demonstrated on both tabular and image datasets drawn from standard machine learning benchmarks. The results show that the procedure also presents the user with a small number of instances, highlighting the practicality of the approach for real-world applications aimed at identifying the decision maker's optimal behavior when faced with imprecision.
Figure 2: Elicitation Examples for the Fashion MNIST dataset for α∗=1.4 . GT stands for ground-truth of the image. M1 is the prediction of the precise classifier, M2 of the (possibly) imprecise classifier
Figure 3: Evolution of the interval size for parameter α∗ across iterations for six datasets using Algorithm 1 ). The y-axis shows the interval size [β,γ] in logarithmic scale, demonstrating exponential convergence towards the true parameter value. The ideal curve (orange) corresponds to halving the interval size at each iteration, while the actual curve (blue) deviates slightly from this ideal behavior due to the tolerance parameter ϵ in Algorithm 1 .
Iteration
#1
#2
#3
#4
#5
wine quality
8
8
7
7
11
glass identification
8
8
7
6
11
heart disease
6
9
9
8
10
students success
5
8
8
7
10
fashion mnist
7
9
12
4
6
mnist
4
5
8
10
5
Table 1: Number of samples retrieved by Algorithm 1 ) to reach the target value α∗ .
Figure 4: Elicitation Examples for the Fashion MNIST dataset for α∗=2.2 .
School of Agriculture and Science, University of KwaZulu-Natal, Pietermaritzburg, KwaZulu-Natal, South Africa · 2AIMS Research and Innovation Centre, African Institute for Mathematical Sciences, Kigali, Rwanda · 3African Institute for Mathematical Sciences, Kigali, Rwanda