Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction
Organizations: Virginia Tech Blacksburg, Virginia, USA · University of Texas at Arlington Arlington, Texas, USA · University of Melbourne Melbourne, Victoria, Australia · University of Toronto Toronto, Ontario, Canada · University of Virginia Charlottesville, Virginia, USA · Bangladesh University of Engineering and Technology Dhaka, Bangladesh
Abstract
Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness interventions. As a complementary case study, we introduce a lightweight adaptation strategy that jointly balances ethnicity--gender--insurance representation without conditioning on mortality outcomes, allowing demographic representation balancing to be examined separately from outcome-conditioned or direct error-rate interventions. We evaluate its behavior at both marginal and corresponding three-way intersectional subgroup levels, while accounting for the statistical support of finer-grained estimates. The results show that interventions can receive substantially different assessments across accuracy/AUROC, sensitivity, and false-positive rate, and that marginal demographic summaries can conceal heterogeneous error profiles within their constituent intersections, including among larger subgroups. These findings highlight the importance of evaluating fairness interventions at both complementary metric and subgroup resolutions, while accounting for the intervention target and the reliability of subgroup estimates.
Figures & tables
| Work | Data/task | Intervention target | Main distinction |
|---|---|---|---|
| Mao et al. ( Mao et al., 2023 ) | Vision benchmarks | Sensitive group label; optional fairness regularization | Last-layer adaptation outside clinical EHRs |
| Lett et al. ( Lett et al., 2025 ) | MIMIC-IV ED admission | Ethnoracial identity gender performance objectives | Direct intersectional performance optimization; different endpoint |
| FairPlay ( Theodorou et al., 2025 ) | MIMIC-IV/eICU mortality | Demographic–outcome-conditioned augmentation | Synthetic records with explicit outcome conditioning |
| Yoon and Kwak ( Yoon and Kwak, 2026 ) | MIMIC-IV mortality | Race or gender targeted separately | Single-axis interventions with cross-axis/intersectional evaluation |
| This study | MIMIC-IV ICU mortality | Ethnicity gender insurance representation | Joint three-way balancing of observed records without outcome conditioning |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| Ethnicity | Gender | Insurance | Prev. | TPR | FPR | Spec. | Acc. | AUROC | ||
|---|---|---|---|---|---|---|---|---|---|---|
| Asian | M | Other | 48 | 6 | 0.125 | 0.833 | 0.381 | 0.619 | 0.646 | 0.877 |
| Black | F | Medicare | 113 | 18 | 0.159 | 0.944 | 0.116 | 0.884 | 0.894 | 0.946 |
| Black | F | Other | 106 | 5 | 0.047 | 0.400 | 0.188 | 0.812 | 0.792 | 0.774 |
| Black | M | Medicare | 81 | 11 | 0.136 | 0.727 | 0.200 | 0.800 | 0.790 | 0.800 |
| Hispanic/Latino | F | Other | 39 | 5 | 0.128 | 1.000 | 0.147 | 0.853 | 0.872 | 0.965 |
| Hispanic/Latino | M | Medicare | 27 | 7 | 0.259 | 0.571 | 0.350 | 0.650 | 0.630 | 0.750 |