Organizations: School of Computer Science and Engineering, Tianjin University of Technology, Tianjin, China · Institute for Artificial Intelligence, Peking University, Beijing, China · DCST, BNRist, RIIT, Institute of Internet Industry, Tsinghua University, Beijing, China · National Institute of Health Data Science, Peking University, Beijing, China
Accurately predicting mortality risk in intensive care unit (ICU) patients is critical for clinical decision-making. Large language models (LLMs) are increasingly explored for clinical prediction using structured medical data, but their outputs may exhibit demographic disparities. Mitigating such disparities without degrading predictive performance remains challenging. We systematically investigate demographic bias in LLM-based ICU mortality prediction and propose Case-Based Prompting (CAP), a training-free framework designed to improve the empirical balance between predictive performance and subgroup fairness. CAP retrieves clinically similar historical misprediction and demographic-sensitive cases with known outcomes and incorporates them as case-level contextual evidence. We further evaluate model discrimination, subgroup fairness, and feature-dependence consistency. On MIMIC-IV, CAP improved AUROC from 0.806 to 0.873 and AUPRC from 0.497 to 0.694 compared with the Baseline Prompt. CAP also reduced several subgroup disparities, particularly for sex and White-Black comparisons, although age-related disparities remained non-negligible. Controlled demographic-information ablations showed that removing demographic information did not uniformly improve both predictive performance and subgroup fairness, indicating distinct performance-fairness patterns rather than a uniform benefit from demographic suppression. Feature-dependence analysis showed generally consistent expressed clinical factor patterns across demographic subgroups. These findings suggest that CAP is a promising inference-time strategy for LLM-based clinical prediction under the evaluated retrospective setting. Further validation across additional LLMs, external datasets, and prospective clinical scenarios is required before deployment.
Figures & tables
Figure 1 : Overview of the Case-Based Prompting (CAP) framework for ICU mortality prediction. For a target patient, CAP retrieves clinically similar historical cases that were misclassified by the baseline prompting strategy and sensitive to counterfactual demographic perturbations from a curated case repository, and incorporates their structured features and known outcomes into the prompt. The framework is designed to provide case-level contextual evidence for LLM-based mortality prediction without model retraining.
Category
Subgroup / Type
Number (%)
Prediction error
FP
278 (69.5%)
FN
122 (30.5%)
Sex
Male
227 (56.8%)
Female
173 (43.2%)
Age group
18–59
126 (31.5%)
60+
274 (68.5%)
Table 1 : Composition of the case repository used by CAP. FP, false positive; FN, false negative.
Method
Core Design
Key Characteristics
Purpose
Baseline Prompt
Task instruction
Basic ICU mortality prediction without fairness instruction, structured reasoning, or retrieved cases
Serves as the reference prompting strategy.
Fairness-Aware Prompt
Fairness instruction
Adds an instruction to avoid unsupported reliance on demographic attributes
Tests whether abstract fairness guidance can reduce subgroup disparities.
System 2 Prompting
Structured reasoning
Guides step-by-step analysis of physiological indicators, laboratory values, severity scores, and interventions
Encourages evidence-based reasoning and reduces shortcut reasoning.
Case-Based Prompting (CAP)
Retrieved case evidence
Incorporates clinically similar historical misprediction and demographic-sensitive cases with known outcomes
Provides case-level contextual evidence for clinically grounded prediction.
Table 2 : Comparison of prompting strategies for ICU mortality prediction.
Parameter
Value
Backbone LLM
Qwen3-32B
Model update
Frozen; no fine-tuning
Temperature
0.1
Top-p
0.9
Maximum output tokens
1024
Random seed
42
Table 3 : Main LLM inference settings.
Feature
Overall (n=45,153)
Train (n=31,607)
Test (n=13,546)
p-value
Demographics
Age (years), mean ± SD
66.1 ± 16.2
66.1 ± 16.2
66.0 ± 16.2
0.55
Female, n (%)
20,138 (44.6)
14,072 (44.5)
6,066 (44.8)
0.62
Race, n (%)
0.45
White
24,404 (54.1)
17,107 (54.1)
7,297 (53.9)
Black
7,043 (15.6)
4,970 (15.7)
2,073 (15.3)
Table 4 : Baseline characteristics of the study cohort. SD, standard deviation. P-values were calculated using the t-test for continuous variables and the chi-square test for categorical variables to compare the training and test sets.
Model
Method
AUROC
AUPRC
F1
Prec
Sen
Spec
Brier
Qwen3 32B
Baseline
0.806
0.497
0.438
0.313
0.725
0.731
0.110
(0.792–0.820)
(0.470–0.525)
(0.412–0.465)
(0.285–0.342)
(0.695–0.755)
(0.715–0.747)
(0.104–0.115)
Fairness-Aware
0.791
0.479
0.402
0.281
0.705
0.694
0.111
(0.777–0.805)
(0.452–0.506)
(0.378–0.428)
(0.255–0.308)
(0.675–0.735)
(0.678–0.710)
(0.106–0.116)
System 2
0.785
0.475
0.426
0.322
0.632
0.774
0.110
(0.771–0.799)
(0.448–0.502)
(0.402–0.452)
(0.294–0.350)
(0.600–0.664)
(0.760–0.788)
(0.105–0.116)
Table 5: Performance comparison of LLM prompting strategies and the XGBoost tabular reference baseline. Metrics are reported with 95% confidence intervals estimated using 1,000 patient-level bootstrap resamples. Statistical significance was evaluated against the Baseline Prompt among LLM-based methods using paired bootstrap tests (* p<0.05 , ** p<0.01 ). Bold values indicate the best value for each metric. Lower Brier scores indicate better calibration.
Figure 2 : Subgroup-specific AUROC performance and disparity analysis across prompting strategies. Left panel: AUROC values across sex, age, and race/ethnicity subgroups. The horizontal dashed lines indicate the overall mean AUROC for the Baseline Prompt and CAP. Right panel: maximum absolute subgroup AUROC difference within each demographic dimension for each prompting strategy. The dashed line at 0.05 is shown as a practical visual reference for subgroup AUROC differences and should not be interpreted as a statistical significance threshold.
Dimension
Comparison
Baseline
Fairness-Aware
System 2
CAP
Sex
Male vs Female
0.083
0.070
0.035
0.003
FNR
0.306/0.223
0.336/0.267
0.242/0.207
0.259/0.256
95% CI
(0.041–0.126)
(0.030–0.112)
(0.006–0.073)
(0.000–0.036)
p -value
0.018
0.031
0.032
0.912
Age
18–59 vs 60+
0.180
0.164
0.152
0.126
FNR
0.143/0.323
0.191/0.355
0.119/0.272
0.170/0.296
Table 6: False-negative-rate (FNR) gaps across demographic subgroups. FNR gap is computed as the absolute difference in FNR between the target subgroup and the reference subgroup, corresponding to Equal Opportunity Difference (EOD) in this setting. Following prior clinical bias-mitigation work, 0.05 was used as a practical reference threshold for potentially meaningful subgroup disparity. Values are interpreted together with 95% bootstrap confidence intervals and permutation-test p -values. The permutation-test p -values assess whether the observed subgroup FNR gap within each method is larger than expected under random demographic group-label assignment, rather than directly testing the gap reduction between methods. Lower values indicate smaller subgroup disparity.
Strategy
AUROC
AUPRC
Sen
Mean FNR gap
Sex gap
Age gap
Race gap
Baseline Prompt
0.806
0.497
0.725
0.101
0.083
0.180
0.071
Clinical-only Retrieval
0.846
0.604
0.805
0.069
0.031
0.143
0.051
Demographic-only Retrieval
0.765
0.431
0.690
0.118
0.078
0.207
0.094
Demographic-aware Retrieval
0.839
0.628
0.782
0.061
0.024
0.151
0.034
Full CAP
0.873
0.694
0.850
0.039
0.003
0.126
0.014
Table 7: Retrieval-strategy ablation for Case-Based Prompting. Clinical-only Retrieval uses clinical variables only for case selection, whereas Demographic-only Retrieval uses age group, sex, and race/ethnicity only. Demographic-aware Retrieval applies demographic candidate filtering before clinical similarity ranking. All retrieval-based strategies provide the selected case’s clinical features and known outcome as auxiliary context, while Full CAP additionally includes historical error type and demographic-sensitivity annotation. Mean FNR gap is calculated across the sex, age, White–Black, and White–Other comparisons; race gap is the average of the White–Black and White–Other FNR gaps. Lower gap values indicate smaller subgroup disparity. Sen, sensitivity.
Setting
AUROC
AUPRC
F1
Sen
Brier
Mean FNR gap
Sex gap
Age gap
Race gap
Full CAP
0.873
0.694
0.525
0.850
0.100
0.039
0.003
0.126
0.014
Demographic-masked Prompt
0.861
0.671
0.511
0.826
0.103
0.036
0.006
0.112
0.012
No-demographic Filtering
0.865
0.681
0.518
0.838
0.102
0.047
0.010
0.133
0.023
Table 8: Controlled demographic-information ablation and fairness–performance trade-off. Mean FNR gap is calculated across the sex, age, White–Black, and White–Other comparisons; race gap is the average of the White–Black and White–Other FNR gaps. Lower Brier scores and FNR gaps indicate better calibration and smaller subgroup disparity. Bold values indicate the numerically best value within each column and do not imply statistical significance; paired-bootstrap confidence intervals for between-method differences are reported in the text. Sen, sensitivity.
Dimension
Comparison
Method
Top-3 Jac.
All Jac.
Cosine
JS Div.
Sex
Male vs Female
Baseline
0.935
0.903
0.996
0.003
Fairness-Aware
0.955
0.928
0.998
0.002
System 2
0.929
0.897
0.996
0.003
CAP
0.965
0.910
0.997
0.004
Age
18–59 vs 60+
Baseline
0.816
0.736
0.966
0.033
Fairness-Aware
0.904
0.855
0.994
0.007
Table 9: Feature-dependence consistency across demographic subgroups. Higher Top-3 Jaccard, all-feature Jaccard, and cosine similarity, together with lower JS divergence, indicate more similar expressed feature-dependence patterns across subgroups.
Figure 3 : Representative case studies illustrating reasoning shifts under Case-Based Prompting (CAP). Each panel compares the Baseline Prompt and CAP in terms of predicted mortality probability, retrieved historical case information, and generated reasoning.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Method
AUROC
AUPRC
F1
Sen
Brier
Mean FNR gap
Qwen3-8B
Baseline Prompt
0.772
0.428
0.401
0.701
0.119
0.112
Fairness-Aware Prompt
0.760
0.416
0.386
0.677
0.121
0.092
System 2 Prompting
0.781
0.445
0.419
0.654
0.116
0.103
CAP
0.828
0.568
0.476
0.789
0.108
0.065
BioMistral-7B
Baseline Prompt
0.748
0.405
0.372
0.641
0.125
0.105
Fairness-Aware Prompt
0.752
0.397
0.356
0.610
0.124
0.088
Appendix
Table 10: Supplementary predictive performance and subgroup fairness evaluation on additional lightweight LLMs. Mean FNR gap is calculated across the main sex, age, White–Black, and White–Other comparisons. Lower Brier scores and lower Mean FNR gaps indicate better calibration and smaller subgroup disparity, respectively.
Model
Method
Sex gap
Age gap
Race gap
Mean FNR gap
Qwen3-8B
Baseline Prompt
0.074
0.205
0.085
0.112
Fairness-Aware Prompt
0.048
0.189
0.066
0.092
System 2 Prompting
0.040
0.220
0.076
0.103
CAP
0.018
0.176
0.033
0.065
BioMistral-7B
Baseline Prompt
0.060
0.190
0.085
0.105
Fairness-Aware Prompt
0.034
0.182
0.068
0.088
Appendix
Table 11: Supplementary subgroup FNR gap evaluation on additional lightweight LLMs. Race gap is calculated as the average FNR gap across the White–Black and White–Other comparisons. White–Asian was not included in the aggregated race gap because of the small Asian subgroup size.
Model
Initial valid JSON rate
Valid after one re-query
Rule-based parsed outputs
Qwen3-8B
94.6%
98.7%
1.3%
BioMistral-7B
78.9%
91.4%
8.6%
Appendix
Table 12: Structured-output validity of additional lightweight LLMs. Invalid outputs were re-queried once and then parsed using the same rule-based procedure as in the main experiments.
Dimension
Comparison
Method
FNR gap
FPR gap
EOdds gap
Sex
Male vs Female
Baseline Prompt
0.083
0.031
0.083
CAP
0.003
0.026
0.026
Age
18–59 vs 60+
Baseline Prompt
0.180
0.058
0.180
CAP
0.126
0.064
0.126
Race/Ethnicity
White vs Black
Baseline Prompt
0.058
0.044
0.058
CAP
0.005
0.031
0.031
Appendix
Table 13: Supplementary error-rate fairness metrics comparing the Baseline Prompt and CAP. FNR gap and FPR gap are computed as absolute differences between the target subgroup and the reference subgroup. Equalized odds gap is defined as the maximum of the FNR gap and FPR gap. Lower values indicate smaller subgroup disparity.
Dimension
Baseline Prompt
Fairness-Aware
System 2
CAP
Sex
0.052
0.044
0.045
0.016
Age
0.070
0.067
0.075
0.029
Race/Ethnicity
0.055
0.057
0.046
0.012
Appendix
Table 14: Maximum absolute subgroup AUROC differences within each demographic dimension across prompting strategies. For sex and age, the value corresponds to the absolute AUROC difference between the two subgroups. For race/ethnicity, the value corresponds to the maximum AUROC difference among the evaluated racial/ethnic subgroups. Lower values indicate smaller subgroup variation in discrimination.
Dimension
Comparison
Method
AUPRC gap
Brier gap
Sex
Male vs Female
Baseline Prompt
0.045
0.008
CAP
0.028
0.006
Age
18–59 vs 60+
Baseline Prompt
0.118
0.014
CAP
0.082
0.012
Race/Ethnicity
White vs Black
Baseline Prompt
0.071
0.010
CAP
0.043
0.007
Appendix
Table 15: Supplementary subgroup AUPRC and calibration gaps comparing the Baseline Prompt and CAP. AUPRC gap and Brier gap are computed as absolute differences between the target subgroup and the reference subgroup. Lower values indicate smaller subgroup disparity.
Similarity threshold
Retrieval coverage
AUROC
AUPRC
Sen
Mean FNR gap
0.7
86.4%
0.860
0.666
0.858
0.052
0.8
70.2%
0.873
0.694
0.850
0.039
0.9
48.7%
0.861
0.655
0.812
0.037
Appendix
Table 16: Sensitivity analysis of the retrieval similarity threshold. Retrieval coverage denotes the proportion of test patients for whom at least one eligible case was retrieved. Mean FNR gap is calculated across the main sex, age, White–Black, and White–Other comparisons.
Number of retrieved cases
AUROC
AUPRC
Sen
Mean FNR gap
k=1
0.873
0.694
0.850
0.039
k=3
0.868
0.701
0.842
0.044
k=5
0.861
0.681
0.857
0.051
Appendix
Table 17: Sensitivity analysis of the number of retrieved cases. The default CAP setting used one retrieved case. Mean FNR gap is calculated across the main sex, age, White–Black, and White–Other comparisons.
Retrieved case format
AUROC
AUPRC
Sen
Mean FNR gap
Clinical features and known outcome only
0.852
0.641
0.806
0.058
Without error-type annotation
0.859
0.670
0.855
0.056
Without demographic-sensitivity annotation
0.868
0.697
0.838
0.048
Full CAP
0.873
0.694
0.850
0.039
Appendix
Table 18: Sensitivity analysis of retrieved case formatting. Full CAP includes clinical similarity, known outcome, historical error type, and demographic-sensitivity annotation. Mean FNR gap is calculated across the main sex, age, White–Black, and White–Other comparisons.
Fairness conclusions in clinical prediction can depend strongly on both the metrics reported and the demographic resolution at which performance is evaluated. We revisit these evaluation choices for ICU mortality prediction on MIMIC-IV, comparing predictive-utility and subgroup-error metrics across several fairness interventions. As a complementary case study, we introduce a lightweight adaptation strategy that jointly balances ethnicity--gender--insurance representation without conditioning on mortality outcomes, allowing demographic representation balancing to be examined separately from outcome-conditioned or direct error-rate interventions. We evaluate its behavior at both marginal and corresponding three-way intersectional subgroup levels, while accounting for the statistical support of finer-grained estimates. The results show that interventions can receive substantially different assessments across accuracy/AUROC, sensitivity, and false-positive rate, and that marginal demographic summaries can conceal heterogeneous error profiles within their constituent intersections, including among larger subgroups. These findings highlight the importance of evaluating fairness interventions at both complementary metric and subgroup resolutions, while accounting for the intervention target and the reliability of subgroup estimates.
Abdullah Al Noman, Fahmid Al Rifat, Tahrima Hashem +3
Virginia Tech Blacksburg, Virginia, USA · University of Texas at Arlington Arlington, Texas, USA · University of Melbourne Melbourne, Victoria, Australia +3
Clinical language models (LMs) are increasingly applied to support clinical risk prediction from free-text notes, yet their uncertainty estimates often remain poorly calibrated and clinically unreliable. In this work, we propose Clinical Uncertainty Risk Alignment (CURA), a framework that aligns clinical LM-based risk estimates and uncertainty with both individual error likelihoods and cohort-level ambiguities. CURA first fine-tunes domain-specific clinical LMs to obtain task-adapted patient embeddings, and then performs uncertainty fine-tuning of a multi-head classifier using a bi-level uncertainty objective. Specifically, an individual-level calibration term aligns predictive uncertainty with each patient's likelihood of error, while a cohort-aware regularizer pulls risk estimates toward event rates in their local neighborhoods in the embedding space and places extra weight on ambiguous cohorts near the decision boundary. We further show that this cohort-aware term can be interpreted as a cross-entropy loss with neighborhood-informed soft labels, providing a label-smoothing view of our method. Extensive experiments on MIMIC-IV clinical risk prediction tasks across various clinical LMs show that CURA consistently improves calibration metrics without substantially compromising discrimination. Further analysis illustrates that CURA reduces overconfident false reassurance and yields more trustworthy uncertainty estimates for downstream clinical decision support.
Objective: Survival analysis is central to medical prediction, yet large language models (LLMs) are rarely used as end-to-end survival models because censoring prevents straightforward supervised fine-tuning. Here we present LLMSurvival, a framework that enables censoring-aware survival analysis with unmodified LLMs operating directly on tabular clinical data. Materials and Methods: LLMSurvival reformulates time-to-event prediction as pairwise ranking among comparable subjects, and derives test-time risk by aggregating comparisons against anchor individuals from the training cohort. Results: Across two clinical tasks (ICU mortality prediction in MIMIC-IV and fragility fracture prediction in a NewYork-Presbyterian/Weill Cornell Medicine cohort), LLMSurvival improves overall concordance over Cox proportional hazards modeling by 3.1% for ICU mortality and 0.5% for fracture risk, 2.1% on average for ICU mortality and 2.8% for fracture risk over three established deep learning survival models. Discussion: The results show that survival modeling with censoring can be made compatible with LLM fine-tuning through comparison-based reformulation. The framework demonstrates high portability and superior performance over expert curated scores like SAPS-II and FRAX scores across diverse clinical context. Furthermore, the framework supports local deployment, as compact, publicly available base models provide sufficient performance. Conclusion: The LLMSurvival framework serves as a proof of concept for an integrated, censoring-conscious approach to survival analysis via LLMs.
Yishu Wei, Hexin Dong, Yi Lin +3
Department of Population Health Science, Weill Cornell Medicine, New York, 10065, US · Department of Medicine, Weill Cornell Medicine, New York, 10065, US