Fairness Theatre: Evaluating Post-Hoc Fairness Interventions in Vendor-Controlled Early Warning Systems
Organizations: University of Toronto, Toronto, Ontario, Canada · Georgia Institute of Technology, Atlanta, Georgia, USA
Abstract
Public institutions increasingly procure AI systems whose design they cannot inspect or change. In higher education, proprietary Early Warning Systems (EWS) leave colleges with few options beyond adjusting model outputs to address inequity. This raises the question of how fairness work is coordinated among vendors, institutions, advisors, and students with unequal power to change these systems? Using student records from a public college in Ontario, Canada, we evaluate six post-hoc fairness interventions on a research EWS under simulated procurement constraints. We compare fairness, accuracy, and demographic disparities, introducing error-type profiling to trace how interventions redistribute false positives and false negatives. Interventions redistributed disparities without consistently reducing them. Two implementations favored already-advantaged groups because they used group size to define disadvantage; small, marginalized groups remained poorly served. These findings show how procurement constraints and implementation choices shape the possibilities for fairness work. We call the resulting condition fairness theatre; dashboard metrics converge while groups' error burdens persist or worsen.
Figures & tables
| Method | Approach | Primary Optimization Target | Source | Key Limitations |
|---|---|---|---|---|
| GetFair | Learns a single global decision threshold via reinforcement learning-inspired optimization | Configurable: SP, EOp, or EO (via reward function) | Sikdar et al. (2022) ( Sikdar et al., 2022 ) | Adjusts one attribute at a time; extreme classifications for small groups; substantial accuracy loss at high fairness weights |
| Decoupled classifiers | Trains group-specific logistic regression models on base model probability scores (post-hoc adaptation) | Configurable (joint accuracy–fairness loss) | Dwork et al. (2018) ( Dwork et al., 2018 ) | Requires large per-group samples; reverts toward pooled baseline for small groups; cannot handle intersections |
| Reject option fairness | Relabels uncertain predictions near the decision boundary based on group membership | Statistical parity | Kamiran et al. (2018) ( Kamiran et al., 2018 ) | Effectiveness depends on density of predictions near boundary; minimal adjustment for groups with few borderline cases |
| MBS | Learns bias scores from group-conditional models; selectively flips predictions exceeding bias threshold | Equalized odds (with accuracy floor) | Chen et al. (2024) ( Chen et al., 2024 ) | Requires group labels and observed outcomes during configuration; results depend on the selected fairness constraint and accuracy floor |
| Bias mitigation post-processing | Audits predictions for group-level disparities and selectively corrects largest imbalances | SP, EOp, and EO simultaneously | Lohia et al. (2018) ( Lohia et al., 2018 ) | Our adaptation requires institutional input features and observed outcomes to train surrogate models; surrogate predictions may not reproduce the inaccessible model’s counterfactual behaviour |
| GBCP | Adapts group-balanced conformal prediction to set group-specific probability thresholds that equalize positive prediction rates to a shared target | Statistical parity (adapted from coverage equalization in original) | Angelopoulos & Bates (2022) ( Angelopoulos and Bates, 2022 ) | Original method equalizes coverage (prediction set reliability) across groups; our adaptation repurposes the quantile-threshold mechanism to target statistical parity, which produces uniform positive rates regardless of base rates; we use 0.50 target rate |
| Cohort | Total | Training | Calibration | Test |
|---|---|---|---|---|
| Domestic | 97,599 | 48,799 | 24,400 | 24,400 |
| International | 70,951 | 35,475 | 17,738 | 17,738 |
| Total | 168,550 | 84,274 | 42,138 | 42,138 |
| Under 25 | 25 to 35 | Over 35 | |||||||
| Method | TPR | FPR | Pos.R | TPR | FPR | Pos.R | TPR | FPR | Pos.R |
| Base | 0.883 | 0.339 | 0.721 | 0.953 | 0.543 | 0.880 | 0.945 | 0.619 | 0.886 |
| Calibrated | 0.938 | 0.430 | 0.787 | 0.979 | 0.621 | 0.915 | 0.966 | 0.668 | 0.912 |
| GetFair | 0.958 | 0.490 | 0.819 | 0.988 | 0.673 | 0.932 | 0.979 | 0.710 | 0.930 |
| Decoupled | 0.938 | 0.429 | 0.786 | 0.979 | 0.621 | 0.915 | 0.968 | 0.681 | 0.916 |
| Reject Option | 0.899 | 0.360 | 0.739 | 0.963 | 0.559 | 0.892 | 0.950 | 0.630 | 0.892 |
| Female | Male | Unknown Gender | |||||||
| Method | TPR | FPR | Pos.R | TPR | FPR | Pos.R | TPR | FPR | Pos.R |
| Base | 0.931 | 0.476 | 0.832 | 0.876 | 0.321 | 0.706 | 0.667 | 0.125 | 0.357 |
| Calibrated | 0.966 | 0.561 | 0.878 | 0.931 | 0.406 | 0.770 | 0.667 | 0.292 | 0.452 |
| GetFair | 0.977 | 0.609 | 0.897 | 0.949 | 0.450 | 0.796 | 0.722 | 0.417 | 0.548 |
| Decoupled | 0.966 | 0.561 | 0.878 | 0.931 | 0.406 | 0.770 | 0.667 | 0.292 | 0.452 |
| Reject Option | 0.941 | 0.496 | 0.845 | 0.894 | 0.341 | 0.724 | 0.722 | 0.417 | 0.548 |
| Domestic | International | |||||
| Method | TPR | FPR | Pos. Rate | TPR | FPR | Pos. Rate |
| Base | 0.875 | 0.374 | 0.704 | 0.937 | 0.413 | 0.856 |
| Calibrated | 0.921 | 0.438 | 0.756 | 0.981 | 0.563 | 0.916 |
| GetFair | 0.948 | 0.499 | 0.794 | 0.981 | 0.564 | 0.916 |
| Decoupled | 0.921 | 0.438 | 0.756 | 0.981 | 0.563 | 0.916 |
| Reject Option | 0.883 | 0.384 | 0.713 | 0.983 | 0.582 | 0.921 |
| Method | Fairness Gains | Trade-offs / Limitations |
|---|---|---|
| Calibration | Domestic TPR +0.046; gender TPR gap narrowed (0.055 0.035); pos. rate convergence improved. | Residual FPR gaps (gender: 0.406 vs. 0.561); modest small-group improvements. |
| GetFair ( ) | Domestic TPR 0.921 0.948; under-25 TPR 0.938 0.958; age pos. rate range narrowed by 0.02. | Single threshold cannot equalize heterogeneous base rates; gender DP immovable (0.628–0.634); converged to baseline. |
| Decoupled Classifiers | High TPR for large groups (age 25–35: 0.979; female: 0.966). | Unknown-gender TPR=0.667; defaults to negative class for small groups; reproduced calibrated baseline. |
| Reject Option | Domestic FPR reduced (0.438 0.384); modest age improvements. | Misdirected residency correction: intl. pos. rate 0.916 0.921, domestic 0.756 0.713 (gap widened); group size unreliable proxy for disadvantage. |
| MBS | Gender EOp 0.05 met (F: 0.998, M: 0.995, N: 1.000); residency 0.10 met. | Gender FPRs surged (F: 0.561 0.837, M: 0.406 0.734); 0.05 infeasible for residency; strict enforcement produced all-zero predictions. |
| Bias Mitigation | Unknown age TPR 0.937 0.975; modest smallest-group corrections. | Misdirected residency correction: intl. FPR 0.563 0.773, pos. rate 0.916 0.956 (gap widened). Gender corrections near baseline. |
| Method | Age | Gender | Residency |
|---|---|---|---|
| GetFair ( ) | 1,189 | 942 | 942 |
| Decoupled Classifiers | 23 | 0 | 0 |
| Reject Option | 1,693 | 1,688 | 1,141 |
| MBS | 0 | 4,842 | 76 |
| Bias Mitigation | 86 | 0 | 706 |
| GBCP | 14,326 | 14,329 | 15,485 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Attribute | Group | TPR (range) | FPR (range) | Pos. Rate (range) | |
| Age | Under 25 | 0.938 (0) | 0.429 (0) | 0.786 (0) | 29,114 |
| 25 to 35 | 0.979 (0) | 0.621 (0) | 0.915 (0) | 8,078 | |
| Over 35 | 0.966 (0) | 0.668 (0) | 0.912 (0) | 4,168 | |
| Unknown | 0.975 (0) | 0.666 (0.005) | 0.887 (0.002) | 778 | |
| Gender | Female | 0.966 (0) | 0.561 (0) | 0.878 (0) | 20,830 |
| Male | 0.931 (0) | 0.406 (0) | 0.770 (0) | 21,266 |
| Method | TPR | FPR | Pos. Rate |
|---|---|---|---|
| Base | 0.667 [0.44, 0.88] | 0.125 [0.00, 0.27] | 0.357 [0.21, 0.50] |
| Calibrated | 0.667 [0.43, 0.88] | 0.292 [0.12, 0.48] | 0.452 [0.31, 0.60] |
| GetFair | 0.722 [0.50, 0.92] | 0.417 [0.22, 0.62] | 0.548 [0.40, 0.69] |
| Decoupled | 0.667 [0.44, 0.88] | 0.292 [0.12, 0.48] | 0.452 [0.31, 0.60] |
| Reject Option | 0.722 [0.50, 0.93] | 0.417 [0.23, 0.62] | 0.548 [0.40, 0.69] |
| MBS | 1.000 [1.00, 1.00] | 0.667 [0.47, 0.85] | 0.810 [0.69, 0.93] |
| A: smallest = deprived | B: lowest pos. rate = deprived | ||||||||
| Method | Group | TPR | FPR | Pos.R | chg. | TPR | FPR | Pos.R | chg. |
| Reject Option | domestic | 0.883 | 0.384 | 0.712 | 1,141 | 0.948 | 0.499 | 0.794 | 1,579 |
| international | 0.983 | 0.582 | 0.922 | 0.957 | 0.462 | 0.881 | |||
| Bias Mitigation | domestic | 0.921 | 0.438 | 0.756 | 706 | no change (0 predictions altered) | |||
| international | 0.990 | 0.773 | 0.956 | ||||||