Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models
Authors: Star S. D. Liu, Xiyu Ding, Robert B. Barrett, Alberto Santamaria-Pang, Nic Dobbins, Harold P. Lehmann
Organizations: Biomedical Informatics and Data Science, Johns Hopkins University School of Medicine, Baltimore, MD, USA · Microsoft, Redmond, WA, USA · National Institutes of Health, Bethesda, MD, USA
How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations relate to clinical predictions, and whether decisions shift as predicted by the specified cost direction and magnitude. Using a public diabetes dataset, we varied 11 false-negative (FN) to false-positive (FP) cost ratios across three phrasings and examined representations and behavioral outputs. Patient risk was linearly recoverable on par with conventional classifiers (AUC ≈0.83), and cost direction was recoverable in every model. However, representational shifts in cost direction tracked output changes only in the two larger models, and responses to cost magnitude were predominantly direction-agnostic. Only 2 of 12 model-phrasings showed both opposing responses to increasing FN versus FP costs and cost-correct ordering. Representationally, a direction fitted on one cost side did not invert when transferred to the other, as expected under mirror-symmetric encoding. These findings suggest that LLMs encode risk and cost information but do not reliably integrate them into cost-correct decisions. Clinical evaluations should therefore include tradeoff tests, phrasing sensitivity, and default operating points alongside predictive performance.
tradeoff-free : No cost tradeoff cost_only : Cost tradeoff only (“FN is X times worse than a FP”). cost_plus_hint : Cost tradeoff + hint (“when uncertain, err towards high/low”)
Tradeoff phrasing (paraphrases of the cost_only sentence)
Baseline : “The cost of a false negative (missing a patient who will develop diabetes) is [ X ] times the cost of a false positive (flagging a patient who will not develop diabetes).” Variant A : “Missing a patient who will develop diabetes is [ X ] times as costly as flagging a patient who will not develop diabetes.” Variant B : “A missed diagnosis is [ X ]-fold more serious than an unnecessary alert.”
Figure 2 : ROC curve with each point from a single tradeoff prompt over the entire dataset for all models. The blue star pins the performance from tradeoff-free prompts. The blue circles are cost_only prompt results. The yellow squares are cost_plus_hint results.
Figure 3 : Layer-wise cost-direction, ground truth risk, and model’s own HIGH/LOW probes’ AUCs.
Figure 4 : Layer-wise per-patient activation projection versus behavioral delta per model, for a 20-fold contrast between FN and FP (20:1 versus 1:20). The vertical dotted lines mark the cliff region identified in the linear-probe trajectory. A high correlation indicates that patients whose representations shift more strongly along the cost direction also show larger logit-gap shifts under cost framing.
Appendix A Table 1 : Example of serializing tabular data into structured text for input.
(a) Prompt without the cost tradeoff framing sentence for a truly high-risk diabetic patient
(b) Prompt with symmetric 1:1 cost of false negative to false positive for a truly high-risk diabetic patient
Layer 0
Layer 0
Layer 29
Layer 29
Layer 79
Layer 79
Appendix
Appendix D Figure 1: Salience maps over two set of representative prompts at 3 different layers for a high-risk patient. (a) shows examples of tradeoff-free prompt. (b) shows example prompts with symmetric 1:1 tradeoff between a false negative and a false positive for a truly high-risk diabetic patient. RED = odds in favor of HIGH, BLUE = odds in favor of LOW. Chat-template boilerplate shared across prompts was omitted from the visualization. Additional implementation details are provided in Appendix B .
Model
AUC (layer)
Qwen-2.5-7B
0.830 (L18)
Llama-3.1-8B
0.813 (L11)
Qwen-2.5-32B
0.832 (L37)
Llama-3.1-70B
0.824 (L33)
Appendix
Appendix D Table 1 : Ground truth risk probe area under the curve (AUC) ceiling for all four models.
Qwen-2.5-7B
Llama-3.1-8B
Qwen-2.5-32B
Llama-3.1-70B
Appendix
Appendix D Figure 2 : Layer-wise probe performance under paraphrase variants of the original cost tradeoff prompt for all 4 models. Variant A was paraphrased as “Missing a patient who will develop diabetes is [X] times as costly as flagging a patient who will not develop diabetes.” Variant B was paraphrased as “A missed diagnosis is [X]-fold more serious than an unnecessary alert.”
Qwen-2.5-7B
Qwen-2.5-32B
Llama-3.1-8B
Llama-3.1-70B
Appendix
Appendix D Figure 3 : Layer-wise probe performance under cost_plus_hint prompt condition for all 4 models. The cost_plus_hint sentence reads “when uncertain, err towards high/low”, depending on the cost of false negative to cost of false positive ratio.
5-fold contrast between FN and FP (5:1 versus 1:5)
10-fold contrast between FN and FP (10:1 versus 1:10)
Appendix
Appendix D Figure 4 : Layer-wise per-patient activation projection versus behavioral delta per model for a 2-fold, 3-fold, 5-fold, and 10-fold contrast between FN and FP. The vertical dotted lines mark the cliff region identified in the previous linear probe trajectory. Pearson and Spearman correlations between s(P)=Δ(P)⋅d−P population-direction projection score, and b(P)=(logitHIGH−logitLOW)FN−(logitHIGH−logitLOW)FP , per-patient behavioral delta across all 768 patients. A high correlation indicates that patients whose representations shift more strongly along the cost-direction also show larger logit-gap shifts under cost-framing, that is representation predicts behavior at the individual level.
model
phrasing
FN-side (want + )
Proportion (%)
FP-side (want − )
Proportion (%)
Differential / Common-mode
Dominance
Llama-3.1-70B
original
0.713
767/768 (99.9%)
− 0.244
604/768 (78.6%)
+0.479 / +0.235
2.04
variant A
0.182
479/768 (62.4%)
0.036
374/768 (48.7%)
+0.073 / +0.109
0.67
variant B
− 0.142
152/768 (19.8%)
− 0.217
745/768 (97.0%)
+0.038 / − 0.180
0.21
Qwen-2.5-32B
original
0.459
725/768 (94.4%)
− 0.672
659/768 (85.8%)
+0.566 / − 0.107
5.31
variant A
0.151
569/768 (74.1%)
0.256
170/768 (22.1%)
− 0.052 / +0.203
0.26
variant B
− 1.450
2/768 (0.3%)
0.083
287/768 (37.4%)
− 0.767 / − 0.684
1.12
Appendix
Appendix D Table 2 : Behavioral mirror test across 4 models and 3 phrasings contrasting 20 versus 2 cost tradeoff magnitude. Cost-correct here requires positive logit delta on the FN-side and negative logit delta on the FP-side.
model
phrasing
FN-side (want + )
Proportion (%)
FP-side (want − )
Proportion (%)
Differential / Common-mode
Dominance
Llama-3.1-70B
original
0.611
768/768 (100.0%)
− 0.206
651/768 (84.8%)
+0.408 / +0.203
2.01
variant A
0.283
687/768 (89.5%)
− 0.029
390/768 (50.8%)
+0.156 / +0.127
1.22
variant B
0.087
532/768 (69.3%)
− 0.548
768/768 (100.0%)
+0.317 / − 0.231
1.38
Qwen-2.5-32B
original
0.634
753/768 (98.0%)
− 0.802
725/768 (94.4%)
+0.718 / − 0.085
8.49
variant A
0.224
609/768 (79.3%)
0.329
136/768 (17.7%)
− 0.053 / +0.277
0.19
variant B
− 0.850
11/768 (1.4%)
− 0.454
693/768 (90.2%)
− 0.198 / − 0.652
0.30
Appendix
Appendix D Table 3 : Behavioral mirror test across 4 models and 3 phrasings contrasting 10 versus 2 cost tradeoff magnitude. Cost-correct here requires positive logit delta on the FN-side and negative logit delta on the FP-side.
model
phrasing
FN-side (want + )
Proportion (%)
FP-side (want − )
Proportion (%)
Differential / Common-mode
Dominance
Llama-3.1-70B
original
0.370
768/768 (100.0%)
− 0.100
570/768 (74.2%)
+0.235 / +0.135
1.74
variant A
0.107
561/768 (73.0%)
− 0.007
432/768 (56.3%)
+0.057 / +0.050
1.15
variant B
− 0.030
300/768 (39.1%)
− 0.046
508/768 (66.1%)
+0.008 / − 0.038
0.20
Qwen-2.5-32B
original
0.422
714/768 (93.0%)
− 0.454
674/768 (87.8%)
+0.438 / − 0.016
27.6
variant A
0.261
676/768 (88.0%)
0.287
137/768 (17.8%)
− 0.013 / +0.274
0.05
variant B
− 0.897
4/768 (0.5%)
1.800
28/768 (3.6%)
− 1.368 / − 0.452
2.99
Appendix
Appendix D Table 4 : Behavioral mirror test across 4 models and 3 phrasings contrasting 20 versus 5 cost tradeoff magnitude. Cost-correct here requires positive logit delta on the FN-side and negative logit delta on the FP-side.
model
phrasing
FN side median τ
FP side median τ
quadrant
Percent with positive FN and FP τ
Llama-3.1-70B
original
1.00
0.80
both positive
77.7%
variant A
0.46
− 0.20
driven to HIGH
17.1%
variant B
0.00
0.60
driven to LOW
37.8%
Qwen-2.5-32B
original
0.60
0.74
both positive
84.8%
variant A
0.32
− 0.40
driven to HIGH
9.8%
variant B
− 1.00
0.00
inversion
0.1%
Appendix
Appendix D Table 5 : Per-patient coherence test across 4 models and 3 phrasings. Given the mirror tradeoff ratio experimental grid, the full-grid evaluation will be 0, and τ must be separately positive. Positive median FP τ and FN τ indicate cost-correct on both sides. Positive median FP τ and negative median FN τ indicate more patients driven towards LOW. Positive median FN τ and negative median FP τ indicate more patients driven towards HIGH. Negative median FP τ and FN τ indicate sign inversion.
Contrast
Llama-3.1-70B
20 versus 2
10 versus 2
20 versus 5
Appendix
Appendix D Figure 5 : Representational mirror test, Llama-3.1-70B for cost ratio contrasts under cost_only prompt. Top: cosine between the FN side and the FP side difference-of-means vectors, d_FN and d_FP, at each layer. The dashed line is the split-half reliability ceiling and the grey band is ± 1 chance sd ( 1/d ). Bottom: held-out patient AUC for the FN (green) and FP (orange) specific differential scores (m). Dashed and dotted lines are FN and FP transfer of each difference vector applied to the numeric contrast.
Contrast
Llama-3.1-8B
20 versus 2
10 versus 2
20 versus 5
Appendix
Appendix D Figure 6 : Representational mirror test, Llama-3.1-8B for cost ratio contrasts under cost_only prompt. Top: cosine between the FN side and the FP side difference-of-means vectors, d_FN and d_FP, at each layer. The dashed line is the split-half reliability ceiling and the grey band is ± 1 chance sd ( 1/d ). Bottom: held-out patient AUC for the FN (green) and FP (orange) specific differential scores (m). Dashed and dotted lines are FN and FP transfer of each difference vector applied to the numeric contrast.
Contrast
Qwen-2.5-32B
20 versus 2
10 versus 2
20 versus 5
Appendix
Appendix D Figure 7 : Representational mirror test, Qwen-2.5-32B for cost ratio contrasts under cost_only prompt. Top: cosine between the FN side and the FP side difference-of-means vectors, d_FN and d_FP, at each layer. The dashed line is the split-half reliability ceiling and the grey band is ± 1 chance sd ( 1/d ). Bottom: held-out patient AUC for the FN (green) and FP (orange) specific differential scores (m). Dashed and dotted lines are FN and FP transfer of each difference vector applied to the numeric contrast.
Contrast
Qwen-2.5-7B
20 versus 2
10 versus 2
20 versus 5
Appendix
Appendix D Figure 8 : Representational mirror test, Qwen-2.5-7B for cost ratio contrasts under cost_only prompt. Top: cosine between the FN side and the FP side difference-of-means vectors, d_FN and d_FP, at each layer. The dashed line is the split-half reliability ceiling and the grey band is ± 1 chance sd ( 1/d ). Bottom: held-out patient AUC for the FN (green) and FP (orange) specific differential scores (m). Dashed and dotted lines are FN and FP transfer of each difference vector applied to the numeric contrast.
Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time series (ISMTS), must deliver both calibrated risk scores for patient triage and interpretable rationales that clinicians can verify. Large Language Models (LLMs) have been explored for this task, yet they collapse graded clinical risk into overconfident binary predictions. This risk polarization undermines both calibration and cross-patient comparability. To address this, we propose TRIAGE, a framework that trains an LLM to generate dialectical reasoning over competing clinical outcomes by eliciting outcome-specific rationales. This dialectical formulation mitigates risk polarization, enabling a single LLM to yield continuous risk scores grounded in explicit clinical reasoning. Evaluated on three ISMTS benchmarks, TRIAGE achieves an average AUPRC improvement of 3.3% and reduces calibration error by 81% compared to the competitive baselines. An LLM-as-a-judge assessment further shows that our rationales surpass post-hoc explanations from the baseline by 20% in clinical reasoning quality. The source code is available at https://github.com/HyeongWon-Jang/TRIAGE .
Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text. We ask whether output format changes the model's \emph{clinical representation} or only the mapping from a preserved representation to an answer. Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B, we find the same medical features fire on the shared clinical narrative under both formats but go {silent} at the multiple-choice decision token in all the cases at every model. Three independent methods (natural-language autoencoder verbalization, decision-token logit attribution, and top-feature characterization) agree that scaffold and format features, but not medical features, drive the decision logits. Behaviorally, the multiple-choice penalty inverts under both structured and natural-language input, option-order shuffle rules out positional bias, and the gap is dominated by off-by-one decision (the model picks an adjacent acuity letter to the gold answer) rather than knowledge failure. Thus, the failure originates in the output format and not in the clinical representation.
David Fraile Navarro, Berardino Como, Jialei Sheng +2
Large language models (LLMs) offer promising clinical decision support but remain vulnerable to hallucinated facts, unsupported recommendations, and citation errors. We present DIASENTINEL, a fully on-premise multi-agent system for one-year type 2 diabetes mellitus (T2DM) risk screening and guideline-grounded report generation from electronic health records (EHRs). The system integrates calibrated risk prediction, deterministic clinical signal extraction, Reciprocal Rank Fusion over American Diabetes Association (ADA) guidelines, and a hybrid verification layer combining rule-based checks with LLM entailment. The demonstration provides a real-time batch-screening dashboard and an interactive patient report interface with cited recommendations, verification results, and raw EHR comparison. DIASENTINEL demonstrates a practical framework for reliable, auditable, and privacy-preserving LLM-based clinical decision support.