As artificial intelligence is increasingly deployed, algorithmic unfairness has raised growing concerns and intensified demands for transparent fairness auditing. In practice, the tolerable degree of algorithmic unfairness depends on the specific legal, ethical, or application context. Given a prespecified tolerance threshold, an important statistical question is how to determine whether a group disparity exceeds the allowable tolerance across different auditing objectives. To address this problem, we develop a unified tolerance-based fairness auditing framework for two complementary auditing objectives: violation certification, which prioritizes control of false violation declarations, and sensitivity screening, which prioritizes reducing missed violations. For the first objective, we develop a constrained empirical likelihood test for formal settings that uses least-favorable-point calibration and can be combined with false flagging rate control for simultaneous subgroup auditing. For the second objective, we develop split empirical likelihood and adjusted split empirical likelihood tests using an adaptive boundary-proxy principle for early-warning settings. Numerical experiments show the distinct error-control--sensitivity trade-offs of these procedures. A COMPAS analysis illustrates the framework in predictive fairness auditing.
Figures & tables
Figure 1 : The theoretical limiting rejection probabilities of the EL, CEL, SEL, and ASEL statistics, together with the CEL2, SEL2, and ASEL2 variants. The threshold is set to 0.1 , with α=α1=0.05 and α2=0.08 .
Figure 2 : Rejection frequency plots under the symmetric normal model: (a)–(d) correspond to n=30,100,300,1000 with N(ϵ,1) data. The horizontal gray line is the nominal level α=0.05 and the vertical dashed line marks ϵ=0.25 .
Figure 3 : Rejection frequency plots under the heavy-tailed Student t model: (a)–(d) correspond to n=30,100,300,1000 with t(3)+ϵ data. The horizontal gray line is the nominal level α=0.05 and the vertical dashed line marks ϵ=0.25 .
Figure 4 : Rejection frequency plots under the right-skewed exponential design Exp(ϵ1) : (a)–(d) correspond to n=30,100,300,1000 . The horizontal gray line is the nominal level α=0.05 and the vertical dashed line marks ϵ=0.25 .
N(ϵ,1)
Exp(1/ϵ)
ϵ
CEL-BH
SEL-BH
ASEL-BH
t -BH
ϵ
CEL-BH
SEL-BH
ASEL-BH
t -BH
0.00
0.00
0.23
0.30
0.00
0.00
-
-
-
-
0.05
0.00
0.18
0.26
0.00
0.05
0.00
0.32
0.32
0.00
0.10
0.00
0.23
0.38
0.00
0.10
0.00
0.25
0.25
0.00
0.15
0.01
0.28
0.56
0.01
0.15
0.00
0.26
0.26
0.00
0.20
0.04
0.18
0.70
0.03
0.20
0.00
0.40
0.42
0.00
Table 1 : Empirical FFR under the global null π1=0 . Bold entries correspond to the tolerance boundary.
N(ϵ,1)
Exp(1/ϵ)
δ
CEL-BH
SEL-BH
ASEL-BH
t -BH
δ
CEL-BH
SEL-BH
ASEL-BH
t -BH
0.05
0.0092
0.0319
0.1562
0.0054
0.05
0.1042
0.1408
0.3692
0.0058
0.20
0.1165
0.1785
0.4638
0.0908
0.10
0.5988
0.5635
0.6931
0.2827
0.25
0.2854
0.3238
0.5858
0.2485
0.15
0.8735
0.8415
0.8896
0.7496
0.35
0.6235
0.6123
0.7915
0.5912
0.20
0.9619
0.9446
0.9596
0.9127
0.50
0.9112
0.8881
0.9496
0.8973
0.25
0.9915
0.9858
0.9896
0.9796
Table 2 : Empirical power for selected configurations with π1=1 .
Figure 5 : Empirical FFR and power under the normal (top) and exponential (bottom) designs. The dashed lines mark the nominal FFR level 0.05 .
Nominal level
EL
CEL
SEL
ASEL
90%
0.0212
0.0251
0.0209
0.0235
95%
0.0180
0.0212
0.0175
0.0235
Table 3 : Lower bounds for the African-American group at nominal levels 90% and 95%.
Figure 6 : Pointwise inversion regions for the COMPAS PPV disparity of each African-American subgroup relative to the entire Caucasian reference group. Top row: 90% nominal level and bottom row: 95% nominal level.
Figure 7 : Pointwise inversion regions for the COMPAS PPV disparity of each subgroup formed by intersections of age and sex relative to the entire Caucasian reference group. Top row: 90% nominal level and bottom row: 95% nominal level.
Figure 8 : Pointwise inversion regions for the COMPAS PPV disparity of each female subgroup defined by the intersection of race and age relative to the overall PPV among defendants receiving a positive COMPAS prediction. Top row: 90% nominal level; bottom row: 95% nominal level. Six race–age intersection subgroups are omitted because the corresponding cell sizes fall below the minimum of n=8 : <25 , Hispanic ; <25 , Other ; 25 – 45 , Hispanic ; 25 – 45 , Native American ; >45 , Hispanic ; and >45 , Native American .
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure A.1 : Rejection-frequency behavior for lower-tail tolerance auditing under the symmetric normal model and the heavy-tailed Student t model: (a)–(b) correspond to n=100,1000 under the normal model, and (c)–(d) correspond to n=100,1000 under the heavy-tailed Student t model. The horizontal gray line is the nominal level α=0.05 . The vertical dashed line marks the null boundary.
Fairness audits are increasingly mandated in high-stakes applications such as hiring, lending, and automated decision-making. Recent work has established fundamental impossibility results for black-box fairness auditing, showing that sufficiently expressive models can evade any auditing strategy. We complement these results by quantifying the extent of unavoidable post-audit manipulation under finite audit resources. We formulate fairness auditing as a min-max optimization between a computationally unbounded company and a budget-constrained auditor. We study two auditing regimes: (i) a budgeted auditor that certifies fairness using a fixed-size audit set, and (ii) a budgeted α-tolerant auditor that additionally requires the audit set to estimate the fairness of the certified model within an α approximation. For both settings, we derive explicit lower bounds on the worst-case post-audit demographic parity deviation as functions of the audit budget, group imbalance, and fairness tolerance. Finally, we empirically illustrate these theoretical limits using simple audit-set construction heuristics with linear and neural network classifiers. Our results demonstrate that increasing audit resources reduces, but does not eliminate, the scope for post-audit manipulation, highlighting fundamental limitations of finite-budget fairness certification.
Rachit Verma, Padala Manisha, Sujit Gujar
Indian Institute of Technology Gandhinagar · International Institute of Information Technology Hyderabad
External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often have limited access to deployed models and must rely on query-based interactions. Most existing fairness evaluation methods assume static datasets and fixed-sample statistical tests, making them poorly suited to real-world auditing scenarios in which evidence must be collected sequentially under query constraints. In this work, we formulate fairness auditing as a tolerance-aware sequential hypothesis-testing problem under limited model output access. We develop a sequential generalized likelihood-ratio framework that allows auditors to accumulate evidence from a finite audit pool and stop once sufficient support for compliance or violation has been obtained. The framework is instantiated for decision-based Statistical Parity and Equal Opportunity audits, and extended to score- and logit-based proxy audits when richer observables are available. Our results show that both the fairness metric and the level of model access significantly affect audit efficiency, and that the benefits of richer output information are not uniform across auditing settings. In particular, richer outputs can substantially reduce the number of queries required for some fairness metrics and operating regimes, while offering limited gains in near-threshold cases. This work provides a practical statistical framework for sequential fairness auditing under realistic deployment constraints.
Ioannis Pitsiorlas, Martha V. Sourla, Marios Kountouris
EURECOM, France · DaSCI, University of Granada, Spain
Fairness audits are a key component of responsible machine-learning deployment. Yet, audit-recommendation reliability under incomplete protected-label access is still poorly understood. In this work, we focused on protected-label missingness in fairness mitigation audits. We introduced a seed-calibrated stress test to separate missingness effects from seed-to-seed movement already present under complete labels. Across ACS/Folktables tasks, missingness settings that retain some protected labels usually do not move selected mitigation methods beyond a complete-label seed-to-seed baseline. At 0 protected-label access, candidates collapse to an empirical-risk-minimization baseline and deterministic tie-breaking rather than revealing a broad missingness effect. We also found that threshold optimization can turn fairness gains on a single protected axis into intersectional harm above a seed baseline, and this threshold-optimizer finding persists under random-forest validation. Overall, our results highlight that protected-label missingness should be reported with seed-null calibration, candidate-set context, and intersectional consequences before it is treated as evidence of audit fragility.