Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction
Authors: Abdullah Al Noman, Fahmid Al Rifat, Tahrima Hashem, Syed Muhammad Ibne Zulfiker, Rishov Paul, Tanzima HAshem
Organizations: Virginia Tech Blacksburg, Virginia, USA · University of Texas at Arlington Arlington, Texas, USA · University of Melbourne Melbourne, Victoria, Australia · University of Toronto Toronto, Ontario, Canada · University of Virginia Charlottesville, Virginia, USA · Bangladesh University of Engineering and Technology Dhaka, Bangladesh
Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness interventions. As a complementary case study, we introduce a lightweight adaptation strategy that jointly balances ethnicity--gender--insurance representation without conditioning on mortality outcomes, allowing demographic representation balancing to be examined separately from outcome-conditioned or direct error-rate interventions. We evaluate its behavior at both marginal and corresponding three-way intersectional subgroup levels, while accounting for the statistical support of finer-grained estimates. The results show that interventions can receive substantially different assessments across accuracy/AUROC, sensitivity, and false-positive rate, and that marginal demographic summaries can conceal heterogeneous error profiles within their constituent intersections, including among larger subgroups. These findings highlight the importance of evaluating fairness interventions at both complementary metric and subgroup resolutions, while accounting for the intervention target and the reliability of subgroup estimates.
Figures & tables
Work
Data/task
Intervention target
Main distinction
Mao et al. ( Mao et al., 2023 )
Vision benchmarks
Sensitive group × label; optional fairness regularization
Direct intersectional performance optimization; different endpoint
FairPlay ( Theodorou et al., 2025 )
MIMIC-IV/eICU mortality
Demographic–outcome-conditioned augmentation
Synthetic records with explicit outcome conditioning
Yoon and Kwak ( Yoon and Kwak, 2026 )
MIMIC-IV mortality
Race or gender targeted separately
Single-axis interventions with cross-axis/intersectional evaluation
This study
MIMIC-IV ICU mortality
Ethnicity × gender × insurance representation
Joint three-way balancing of observed records without outcome conditioning
Table 1: Closest methodological contrasts in intervention target and demographic resolution.
Figure 1 : Study design. IBHFT performs equal-count adaptation over observed ethnicity–gender–insurance intersections and updates only the replacement classifier head. The resulting model is evaluated along two dimensions: predictive-utility versus subgroup-error metrics, and marginal versus intersectional demographic resolution.
Figure 2 : Group-wise true-positive rate (top) and false-positive rate (bottom) across repeated evaluations. The series labeled “Ours” corresponds to IBHFT. Reweighting attains very low FPR but substantially lower TPR, whereas IBHFT preserves a sensitivity profile closer to the baseline.
Figure 3 : Predictive utility and sensitivity support different assessments of reweighting. High accuracy/AUROC (top) coexists with low sensitivity (bottom).
Figure 4 : Ethnicity-marginal sensitivity (dashed line) and sensitivity estimates for larger nested ethnicity–gender–insurance intersections. Main-text points include intersections with N≥100 and at least five mortality-positive cases; horizontal bars show 95% Wilson intervals. Dashed lines denote marginal sensitivity from the IBHFT predictions used for the intersectional analysis. Appendix B reports intersections with N+≥5 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Ethnicity
Gender
Insurance
N
N+
Prev.
TPR
FPR
Spec.
Acc.
AUROC
Asian
M
Other
48
6
0.125
0.833
0.381
0.619
0.646
0.877
Black
F
Medicare
113
18
0.159
0.944
0.116
0.884
0.894
0.946
Black
F
Other
106
5
0.047
0.400
0.188
0.812
0.792
0.774
Black
M
Medicare
81
11
0.136
0.727
0.200
0.800
0.790
0.800
Hispanic/Latino
F
Other
39
5
0.128
1.000
0.147
0.853
0.872
0.965
Hispanic/Latino
M
Medicare
27
7
0.259
0.571
0.350
0.650
0.630
0.750
Appendix
Table 2 : Three-way intersectional audit for subgroups with at least five mortality-positive cases ( N+≥5 ). Point estimates should be interpreted together with subgroup support; Figure 6 reports Wilson intervals for sensitivity.
Figure 5 : TPR–FPR profiles for ethnicity–gender–insurance intersections with at least five mortality-positive cases ( N+≥5 ). Marker size scales with subgroup size.
Figure 6 : Sensitivity estimates with 95% Wilson intervals for ethnicity–gender–insurance intersections with at least five mortality-positive cases ( N+≥5 ).
Accurately predicting mortality risk in intensive care unit (ICU) patients is critical for clinical decision-making. Large language models (LLMs) are increasingly explored for clinical prediction using structured medical data, but their outputs may exhibit demographic disparities. Mitigating such disparities without degrading predictive performance remains challenging. We systematically investigate demographic bias in LLM-based ICU mortality prediction and propose Case-Based Prompting (CAP), a training-free framework designed to improve the empirical balance between predictive performance and subgroup fairness. CAP retrieves clinically similar historical misprediction and demographic-sensitive cases with known outcomes and incorporates them as case-level contextual evidence. We further evaluate model discrimination, subgroup fairness, and feature-dependence consistency. On MIMIC-IV, CAP improved AUROC from 0.806 to 0.873 and AUPRC from 0.497 to 0.694 compared with the Baseline Prompt. CAP also reduced several subgroup disparities, particularly for sex and White-Black comparisons, although age-related disparities remained non-negligible. Controlled demographic-information ablations showed that removing demographic information did not uniformly improve both predictive performance and subgroup fairness, indicating distinct performance-fairness patterns rather than a uniform benefit from demographic suppression. Feature-dependence analysis showed generally consistent expressed clinical factor patterns across demographic subgroups. These findings suggest that CAP is a promising inference-time strategy for LLM-based clinical prediction under the evaluated retrospective setting. Further validation across additional LLMs, external datasets, and prospective clinical scenarios is required before deployment.
Gangxiong Zhang, Yongchao Long, Yuxi Zhou +2
School of Computer Science and Engineering, Tianjin University of Technology, Tianjin, China · Institute for Artificial Intelligence, Peking University, Beijing, China · DCST, BNRist, RIIT, Institute of Internet Industry, Tsinghua University, Beijing, China +1
Objective: To propose and retrospectively validate an integrated framework addressing three barriers to clinical translation of readmission prediction: lack of explainability, absence of deployment reliability infrastructure, and inadequate demographic fairness evaluation. Materials and Methods: We constructed a cohort of 415231 adult admissions from the MIMIC-IV database (30-day readmission prevalence 18.0%), split 70/15/15. Logistic regression, XGBoost, and LightGBM models were trained on 26 features. SHAP provided per-patient explanations. Fairness was evaluated across 16 subgroups using AUC-ROC, false negative rate (FNR), and positive predictive value (PPV). Calibration was assessed using Brier scores and calibration curves. Results: XGBoost achieved AUC-ROC 0.696 (95% CI 0.691-0.701), outperforming or matching the LACE baseline (AUC 0.60-0.68). LightGBM achieved best calibration (Brier 0.146). Prior admissions were the dominant predictor. All subgroups met equity thresholds (delta AUC <= 0.05, delta FNR <= 0.10). Conclusion: This framework delivers competitive performance, clinically actionable explanations, and strong demographic equity. Code is publicly available at https://github.com/Tomisin92/readmission-prediction.
Isaac Tosin Adisa
Department of Statistics, Florida State University, Tallahassee, FL 32306, USA
Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes. This limits practical guidance for selecting fairness strategies, where disparities may arise across intersectional subgroups and across multiple stages of the modeling lifecycle. This work presents FairSelect, a toolkit for systematically evaluating fairness mitigation strategies applied individually and in combination across preprocessing, inprocessing, and postprocessing stages. FairSelect supports multiple model architectures, intersectional subgroup evaluation, and comparison of fairness utility tradeoffs across baseline, single method, and multi level configurations. The framework was validated using synthetic clinical datasets designed to represent specific bias mechanisms and a real-world replication of two-year stroke risk prediction among patients with atrial fibrillation. Synthetic experiments showed that targeted fairness methods generally reduced intended subgroup disparities, while combined strategies produced larger average fairness improvements with modest utility tradeoffs. In the clinical prediction task, mitigation effects were highly variable, with some combinations improving both fairness and predictive performance while others were ineffective or counterproductive. These findings demonstrate that fairness interventions interact in nonadditive and context dependent ways. FairSelect provides a practical framework for systematically identifying fairness strategies that improve subgroup equity while preserving model performance in clinical machine learning.
Nick Souligne, Isabella Mixton-Garcia, Vignesh Subbian
College of Engineering, University of Arizona, Tucson, AZ