Organizations: School of Computer Science and Engineering, Tianjin University of Technology, Tianjin, China · Institute for Artificial Intelligence, Peking University, Beijing, China · DCST, BNRist, RIIT, Institute of Internet Industry, Tsinghua University, Beijing, China · National Institute of Health Data Science, Peking University, Beijing, China
Accurately predicting mortality risk in intensive care unit (ICU) patients is critical for clinical decision-making. Large language models (LLMs) are increasingly explored for clinical prediction using structured medical data, but their outputs may exhibit demographic disparities. Mitigating such disparities without degrading predictive performance remains challenging. We systematically investigate demographic bias in LLM-based ICU mortality prediction and propose Case-Based Prompting (CAP), a training-free framework designed to improve the empirical balance between predictive performance and subgroup fairness. CAP retrieves clinically similar historical misprediction and demographic-sensitive cases with known outcomes and incorporates them as case-level contextual evidence. We further evaluate model discrimination, subgroup fairness, and feature-dependence consistency. On MIMIC-IV, CAP improved AUROC from 0.806 to 0.873 and AUPRC from 0.497 to 0.694 compared with the Baseline Prompt. CAP also reduced several subgroup disparities, particularly for sex and White-Black comparisons, although age-related disparities remained non-negligible. Controlled demographic-information ablations showed that removing demographic information did not uniformly improve both predictive performance and subgroup fairness, indicating distinct performance-fairness patterns rather than a uniform benefit from demographic suppression. Feature-dependence analysis showed generally consistent expressed clinical factor patterns across demographic subgroups. These findings suggest that CAP is a promising inference-time strategy for LLM-based clinical prediction under the evaluated retrospective setting. Further validation across additional LLMs, external datasets, and prospective clinical scenarios is required before deployment.
Figures & tables
Figure 1 : Overview of the Case-Based Prompting (CAP) framework for ICU mortality prediction. For a target patient, CAP retrieves clinically similar historical cases that were misclassified by the baseline prompting strategy and sensitive to counterfactual demographic perturbations from a curated case repository, and incorporates their structured features and known outcomes into the prompt. The framework is designed to provide case-level contextual evidence for LLM-based mortality prediction without model retraining.
Category
Subgroup / Type
Number (%)
Prediction error
FP
278 (69.5%)
FN
122 (30.5%)
Sex
Male
227 (56.8%)
Female
173 (43.2%)
Age group
18–59
126 (31.5%)
60+
274 (68.5%)
Table 1 : Composition of the case repository used by CAP. FP, false positive; FN, false negative.
Method
Core Design
Key Characteristics
Purpose
Baseline Prompt
Task instruction
Basic ICU mortality prediction without fairness instruction, structured reasoning, or retrieved cases
Serves as the reference prompting strategy.
Fairness-Aware Prompt
Fairness instruction
Adds an instruction to avoid unsupported reliance on demographic attributes
Tests whether abstract fairness guidance can reduce subgroup disparities.
System 2 Prompting
Structured reasoning
Guides step-by-step analysis of physiological indicators, laboratory values, severity scores, and interventions
Encourages evidence-based reasoning and reduces shortcut reasoning.
Case-Based Prompting (CAP)
Retrieved case evidence
Incorporates clinically similar historical misprediction and demographic-sensitive cases with known outcomes
Provides case-level contextual evidence for clinically grounded prediction.
Table 2 : Comparison of prompting strategies for ICU mortality prediction.
Parameter
Value
Backbone LLM
Qwen3-32B
Model update
Frozen; no fine-tuning
Temperature
0.1
Top-p
0.9
Maximum output tokens
1024
Random seed
42
Table 3 : Main LLM inference settings.
Feature
Overall (n=45,153)
Train (n=31,607)
Test (n=13,546)
p-value
Demographics
Age (years), mean ± SD
66.1 ± 16.2
66.1 ± 16.2
66.0 ± 16.2
0.55
Female, n (%)
20,138 (44.6)
14,072 (44.5)
6,066 (44.8)
0.62
Race, n (%)
0.45
White
24,404 (54.1)
17,107 (54.1)
7,297 (53.9)
Black
7,043 (15.6)
4,970 (15.7)
2,073 (15.3)
Table 4 : Baseline characteristics of the study cohort. SD, standard deviation. P-values were calculated using the t-test for continuous variables and the chi-square test for categorical variables to compare the training and test sets.
Model
Method
AUROC
AUPRC
F1
Prec
Sen
Spec
Brier
Qwen3 32B
Baseline
0.806
0.497
0.438
0.313
0.725
0.731
0.110
(0.792–0.820)
(0.470–0.525)
(0.412–0.465)
(0.285–0.342)
(0.695–0.755)
(0.715–0.747)
(0.104–0.115)
Fairness-Aware
0.791
0.479
0.402
0.281
0.705
0.694
0.111
(0.777–0.805)
(0.452–0.506)
(0.378–0.428)
(0.255–0.308)
(0.675–0.735)
(0.678–0.710)
(0.106–0.116)
System 2
0.785
0.475
0.426
0.322
0.632
0.774
0.110
(0.771–0.799)
(0.448–0.502)
(0.402–0.452)
(0.294–0.350)
(0.600–0.664)
(0.760–0.788)
(0.105–0.116)
Table 5: Performance comparison of LLM prompting strategies and the XGBoost tabular reference baseline. Metrics are reported with 95% confidence intervals estimated using 1,000 patient-level bootstrap resamples. Statistical significance was evaluated against the Baseline Prompt among LLM-based methods using paired bootstrap tests (* p<0.05 , ** p<0.01 ). Bold values indicate the best value for each metric. Lower Brier scores indicate better calibration.
Figure 2 : Subgroup-specific AUROC performance and disparity analysis across prompting strategies. Left panel: AUROC values across sex, age, and race/ethnicity subgroups. The horizontal dashed lines indicate the overall mean AUROC for the Baseline Prompt and CAP. Right panel: maximum absolute subgroup AUROC difference within each demographic dimension for each prompting strategy. The dashed line at 0.05 is shown as a practical visual reference for subgroup AUROC differences and should not be interpreted as a statistical significance threshold.
Dimension
Comparison
Baseline
Fairness-Aware
System 2
CAP
Sex
Male vs Female
0.083
0.070
0.035
0.003
FNR
0.306/0.223
0.336/0.267
0.242/0.207
0.259/0.256
95% CI
(0.041–0.126)
(0.030–0.112)
(0.006–0.073)
(0.000–0.036)
p -value
0.018
0.031
0.032
0.912
Age
18–59 vs 60+
0.180
0.164
0.152
0.126
FNR
0.143/0.323
0.191/0.355
0.119/0.272
0.170/0.296
Table 6: False-negative-rate (FNR) gaps across demographic subgroups. FNR gap is computed as the absolute difference in FNR between the target subgroup and the reference subgroup, corresponding to Equal Opportunity Difference (EOD) in this setting. Following prior clinical bias-mitigation work, 0.05 was used as a practical reference threshold for potentially meaningful subgroup disparity. Values are interpreted together with 95% bootstrap confidence intervals and permutation-test p -values. The permutation-test p -values assess whether the observed subgroup FNR gap within each method is larger than expected under random demographic group-label assignment, rather than directly testing the gap reduction between methods. Lower values indicate smaller subgroup disparity.
Strategy
AUROC
AUPRC
Sen
Mean FNR gap
Sex gap
Age gap
Race gap
Baseline Prompt
0.806
0.497
0.725
0.101
0.083
0.180
0.071
Clinical-only Retrieval
0.846
0.604
0.805
0.069
0.031
0.143
0.051
Demographic-only Retrieval
0.765
0.431
0.690
0.118
0.078
0.207
0.094
Demographic-aware Retrieval
0.839
0.628
0.782
0.061
0.024
0.151
0.034
Full CAP
0.873
0.694
0.850
0.039
0.003
0.126
0.014
Table 7: Retrieval-strategy ablation for Case-Based Prompting. Clinical-only Retrieval uses clinical variables only for case selection, whereas Demographic-only Retrieval uses age group, sex, and race/ethnicity only. Demographic-aware Retrieval applies demographic candidate filtering before clinical similarity ranking. All retrieval-based strategies provide the selected case’s clinical features and known outcome as auxiliary context, while Full CAP additionally includes historical error type and demographic-sensitivity annotation. Mean FNR gap is calculated across the sex, age, White–Black, and White–Other comparisons; race gap is the average of the White–Black and White–Other FNR gaps. Lower gap values indicate smaller subgroup disparity. Sen, sensitivity.
Setting
AUROC
AUPRC
F1
Sen
Brier
Mean FNR gap
Sex gap
Age gap
Race gap
Full CAP
0.873
0.694
0.525
0.850
0.100
0.039
0.003
0.126
0.014
Demographic-masked Prompt
0.861
0.671
0.511
0.826
0.103
0.036
0.006
0.112
0.012
No-demographic Filtering
0.865
0.681
0.518
0.838
0.102
0.047
0.010
0.133
0.023
Table 8: Controlled demographic-information ablation and fairness–performance trade-off. Mean FNR gap is calculated across the sex, age, White–Black, and White–Other comparisons; race gap is the average of the White–Black and White–Other FNR gaps. Lower Brier scores and FNR gaps indicate better calibration and smaller subgroup disparity. Bold values indicate the numerically best value within each column and do not imply statistical significance; paired-bootstrap confidence intervals for between-method differences are reported in the text. Sen, sensitivity.
Dimension
Comparison
Method
Top-3 Jac.
All Jac.
Cosine
JS Div.
Sex
Male vs Female
Baseline
0.935
0.903
0.996
0.003
Fairness-Aware
0.955
0.928
0.998
0.002
System 2
0.929
0.897
0.996
0.003
CAP
0.965
0.910
0.997
0.004
Age
18–59 vs 60+
Baseline
0.816
0.736
0.966
0.033
Fairness-Aware
0.904
0.855
0.994
0.007
Table 9: Feature-dependence consistency across demographic subgroups. Higher Top-3 Jaccard, all-feature Jaccard, and cosine similarity, together with lower JS divergence, indicate more similar expressed feature-dependence patterns across subgroups.
Figure 3 : Representative case studies illustrating reasoning shifts under Case-Based Prompting (CAP). Each panel compares the Baseline Prompt and CAP in terms of predicted mortality probability, retrieved historical case information, and generated reasoning.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Method
AUROC
AUPRC
F1
Sen
Brier
Mean FNR gap
Qwen3-8B
Baseline Prompt
0.772
0.428
0.401
0.701
0.119
0.112
Fairness-Aware Prompt
0.760
0.416
0.386
0.677
0.121
0.092
System 2 Prompting
0.781
0.445
0.419
0.654
0.116
0.103
CAP
0.828
0.568
0.476
0.789
0.108
0.065
BioMistral-7B
Baseline Prompt
0.748
0.405
0.372
0.641
0.125
0.105
Fairness-Aware Prompt
0.752
0.397
0.356
0.610
0.124
0.088
Appendix
Table 10: Supplementary predictive performance and subgroup fairness evaluation on additional lightweight LLMs. Mean FNR gap is calculated across the main sex, age, White–Black, and White–Other comparisons. Lower Brier scores and lower Mean FNR gaps indicate better calibration and smaller subgroup disparity, respectively.
Model
Method
Sex gap
Age gap
Race gap
Mean FNR gap
Qwen3-8B
Baseline Prompt
0.074
0.205
0.085
0.112
Fairness-Aware Prompt
0.048
0.189
0.066
0.092
System 2 Prompting
0.040
0.220
0.076
0.103
CAP
0.018
0.176
0.033
0.065
BioMistral-7B
Baseline Prompt
0.060
0.190
0.085
0.105
Fairness-Aware Prompt
0.034
0.182
0.068
0.088
Appendix
Table 11: Supplementary subgroup FNR gap evaluation on additional lightweight LLMs. Race gap is calculated as the average FNR gap across the White–Black and White–Other comparisons. White–Asian was not included in the aggregated race gap because of the small Asian subgroup size.
Model
Initial valid JSON rate
Valid after one re-query
Rule-based parsed outputs
Qwen3-8B
94.6%
98.7%
1.3%
BioMistral-7B
78.9%
91.4%
8.6%
Appendix
Table 12: Structured-output validity of additional lightweight LLMs. Invalid outputs were re-queried once and then parsed using the same rule-based procedure as in the main experiments.
Dimension
Comparison
Method
FNR gap
FPR gap
EOdds gap
Sex
Male vs Female
Baseline Prompt
0.083
0.031
0.083
CAP
0.003
0.026
0.026
Age
18–59 vs 60+
Baseline Prompt
0.180
0.058
0.180
CAP
0.126
0.064
0.126
Race/Ethnicity
White vs Black
Baseline Prompt
0.058
0.044
0.058
CAP
0.005
0.031
0.031
Appendix
Table 13: Supplementary error-rate fairness metrics comparing the Baseline Prompt and CAP. FNR gap and FPR gap are computed as absolute differences between the target subgroup and the reference subgroup. Equalized odds gap is defined as the maximum of the FNR gap and FPR gap. Lower values indicate smaller subgroup disparity.
Dimension
Baseline Prompt
Fairness-Aware
System 2
CAP
Sex
0.052
0.044
0.045
0.016
Age
0.070
0.067
0.075
0.029
Race/Ethnicity
0.055
0.057
0.046
0.012
Appendix
Table 14: Maximum absolute subgroup AUROC differences within each demographic dimension across prompting strategies. For sex and age, the value corresponds to the absolute AUROC difference between the two subgroups. For race/ethnicity, the value corresponds to the maximum AUROC difference among the evaluated racial/ethnic subgroups. Lower values indicate smaller subgroup variation in discrimination.
Dimension
Comparison
Method
AUPRC gap
Brier gap
Sex
Male vs Female
Baseline Prompt
0.045
0.008
CAP
0.028
0.006
Age
18–59 vs 60+
Baseline Prompt
0.118
0.014
CAP
0.082
0.012
Race/Ethnicity
White vs Black
Baseline Prompt
0.071
0.010
CAP
0.043
0.007
Appendix
Table 15: Supplementary subgroup AUPRC and calibration gaps comparing the Baseline Prompt and CAP. AUPRC gap and Brier gap are computed as absolute differences between the target subgroup and the reference subgroup. Lower values indicate smaller subgroup disparity.
Similarity threshold
Retrieval coverage
AUROC
AUPRC
Sen
Mean FNR gap
0.7
86.4%
0.860
0.666
0.858
0.052
0.8
70.2%
0.873
0.694
0.850
0.039
0.9
48.7%
0.861
0.655
0.812
0.037
Appendix
Table 16: Sensitivity analysis of the retrieval similarity threshold. Retrieval coverage denotes the proportion of test patients for whom at least one eligible case was retrieved. Mean FNR gap is calculated across the main sex, age, White–Black, and White–Other comparisons.
Number of retrieved cases
AUROC
AUPRC
Sen
Mean FNR gap
k=1
0.873
0.694
0.850
0.039
k=3
0.868
0.701
0.842
0.044
k=5
0.861
0.681
0.857
0.051
Appendix
Table 17: Sensitivity analysis of the number of retrieved cases. The default CAP setting used one retrieved case. Mean FNR gap is calculated across the main sex, age, White–Black, and White–Other comparisons.
Retrieved case format
AUROC
AUPRC
Sen
Mean FNR gap
Clinical features and known outcome only
0.852
0.641
0.806
0.058
Without error-type annotation
0.859
0.670
0.855
0.056
Without demographic-sensitivity annotation
0.868
0.697
0.838
0.048
Full CAP
0.873
0.694
0.850
0.039
Appendix
Table 18: Sensitivity analysis of retrieved case formatting. Full CAP includes clinical similarity, known outcome, historical error type, and demographic-sensitivity annotation. Mean FNR gap is calculated across the main sex, age, White–Black, and White–Other comparisons.
Virginia Tech Blacksburg, Virginia, USA · University of Texas at Arlington Arlington, Texas, USA · University of Melbourne Melbourne, Victoria, Australia +3
Department of Population Health Science, Weill Cornell Medicine, New York, 10065, US · Department of Medicine, Weill Cornell Medicine, New York, 10065, US