cs.LGOct 1, 2026

Beyond Demographic Balance: Multi-Metric and Intersectional Evaluation of Fairness in MIMIC-IV Mortality Prediction

Authors: Abdullah Al Noman, Fahmid Al Rifat, Tahrima Hashem, Syed Muhammad Ibne Zulfiker, Rishov Paul, Tanzima HAshem

Organizations: Virginia Tech Blacksburg, Virginia, USA · University of Texas at Arlington Arlington, Texas, USA · University of Melbourne Melbourne, Victoria, Australia · University of Toronto Toronto, Ontario, Canada · University of Virginia Charlottesville, Virginia, USA · Bangladesh University of Engineering and Technology Dhaka, Bangladesh

Abstract

Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness interventions. As a complementary case study, we introduce a lightweight adaptation strategy that jointly balances ethnicity--gender--insurance representation without conditioning on mortality outcomes, allowing demographic representation balancing to be examined separately from outcome-conditioned or direct error-rate interventions. We evaluate its behavior at both marginal and corresponding three-way intersectional subgroup levels, while accounting for the statistical support of finer-grained estimates. The results show that interventions can receive substantially different assessments across accuracy/AUROC, sensitivity, and false-positive rate, and that marginal demographic summaries can conceal heterogeneous error profiles within their constituent intersections, including among larger subgroups. These findings highlight the importance of evaluating fairness interventions at both complementary metric and subgroup resolutions, while accounting for the intervention target and the reliability of subgroup estimates.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Improving Fairness of Large Language Model-Based ICU Mortality Prediction via Case-Based Prompting

    Dec 17, 2025Gangxiong Zhang, Yongchao Long, Yuxi Zhou +2Intensive Care Unit Mortality PredictionAlgorithmic Fairness

  2. An Integrated Framework for Explainable, Fair, and Observable Hospital Readmission Prediction: Development and Validation on MIMIC-IV

    Apr 24, 2026Isaac Tosin AdisaReadmission PredictionRoc-Auc

  3. FairSelect: A Systematic Evaluation of Multi-Level and Intersectional Algorithmic Fairness

    Jul 9, 2026Nick Souligne, Isabella Mixton-Garcia, Vignesh SubbianAlgorithmic FairnessSubgroups