LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage
Authors: Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, +1 more
Organizations: University of New South Wales, Sydney · Independent Researcher · St Vincent’s Hospital, Sydney · Royal Prince Alfred Hospital, Sydney · Campbelltown Hospital, Sydney
Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425 LLM-generated (SIMULATED) and 50 physician-authored (CLINICIAN) conversations, both labelled under the Emergency Severity Index (ESI). Every model, measured by quadratic weighted kappa (QWK), degrades from moderate-to-substantial agreement on completed records to fair-to-moderate agreement at every sequential checkpoint. Controlled perturbations show that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lift this plateau. Models extract clinically relevant content from later turns, yet the surprisal of the true label rises across the checkpoints. So the model fails to integrate the evidence. Three expert clinicians on the same conversations reach a QWK of 0.887-0.929, while the best model reaches 0.295. Predictions concentrate at ESI-2 and ESI-3, and models agree with each other more than with the ground truth, so ensembling worsens the failure. Deploying LLMs for ED triage based on offline benchmarks alone misses this sequential failure.
Figures & tables
Figure 1: Input format for the two information settings: offline and sequential . The offline input serialises the structured EHR record into text. The sequential input consists of conversation prefixes at five checkpoints indexed by k , with utterances accumulating at subsequent checkpoints ( ⋃ ). We evaluate the model at every checkpoint.
Figure 2: Performance comparison of all models across the five sequential checkpoints on the simulated corpus. Dashed line per panel represents the offline skyline.
Figure 3: Performance comparison of all models across the five sequential checkpoints on the simulated corpus under four prompting strategies.
Figure 4: Performance of all models on the clinician corpus and per-model Δ QWK from k =4 to +100% on both corpora.
Figure 5: gemma-l prediction shape and inter-model agreement across all checkpoints under sequential .
Figure 6: Change in QWK per perturbation, model, and checkpoint on the simulated corpus. Blue = perturbation degrades QWK relative to vanilla ; red = perturbation improves QWK.
Figure 7: Proportion of conversations committed at each checkpoint. Commit rates on the clinician corpus are in Appendix D.4 .
Figure 8: Commit rates per expert annotator on 50 conversations from simulated .
Figure 9: Extraction of red flags in both corpora and surprisal of the true ESI label per checkpoint in simulated .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: The ESI triage algorithm; accessed from the ESI Handbook .
Figure 11: Performance comparison of all models across the five sequential checkpoints on the clinician corpus under four prompting strategies. Companion of Figure 3 .
Figure 12: gemma-s confusion matrices at each checkpoint on simulated (top row, n=425) and clinician (bottom row, n=40).
Figure 13: qwen-s confusion matrices at each checkpoint on simulated (top row, n=425) and clinician (bottom row, n=40).
Figure 14: qwen-l confusion matrices at each checkpoint on simulated (top row, n=425) and clinician (bottom row, n=40).
Figure 15: claude confusion matrices at each checkpoint on simulated (top row, n=425) and clinician (bottom row, n=40).
Figure 16: gpt confusion matrices at each checkpoint on simulated (top row, n=425) and clinician (bottom row, n=40).
Figure 17: Change in QWK per perturbation, model, and checkpoint on the clinician corpus. Blue = perturbation degrades QWK relative to VANILLA; red = perturbation improves QWK. Companion of Figure 6 .
Figure 18: Proportion of conversations committed at each checkpoint, per model on clinician .
Figure 19: Screenshot of an example prefix provided to an annotator.
As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting. We introduce a benchmark of over 6,000 clinical scenarios, 7,000 physician annotations, and 225,000 model responses. Using this benchmark, we make two key observations. First, LLMs are more likely than physicians to recommend unnecessary care at baseline, and this tendency increases under perturbed inputs. Further, we find that LLM recommendations are more sensitive to gender and tone perturbations than human recommendations. Together, these results demonstrate that LLMs can vary under clinically irrelevant textual changes, highlighting the need for deployment-oriented evaluations grounded in expert physician behavior.
Massachusetts Institute of Technology, Department of Electrical Engineering and Computer Science · Worcester Polytechnic Institute, Department of Computer Science
At Noora Health, our nurses answer more than 50,000 medical queries per month on our WhatsApp-based service that provides caregivers with on-demand support. Their most time-critical task is emergency triage: deciding which queries need immediate in-person attention. To support them, we built a system that uses a large language model (LLM) to classify whether a message is an emergency and provide a rationale for interpretability. But the system was opaque: analyzing mistakes meant reading reasoning chains for each message, which is infeasible at our scale. Prompt changes meant re-running a full evaluation to prevent regressions, which was both costly and operationally challenging. Clinicians follow a decision tree to make this call, but it was never documented or passed to the model, which relied on a flat list of danger signs. To address these issues, we decomposed triage into two steps: an LLM extracts canonical symptoms and patient context from the query using a clinician-authored vocabulary, and a deterministic rule engine captures the scenarios that indicate an emergency. We show that the new system raised recall from 0.565 to 0.810 and F1 from 0.606 to 0.702, with structured rules driving most of the accuracy gains while the decomposition provides auditability: clinical experts can inspect each stage of the new system to see whether the query was mistranslated, symptoms were incorrectly extracted, patient context was wrongly inferred, or the necessary rules were missing. They can add new rules independently without causing regressions and avoid running costly evaluations. Since deployment, the new system has triaged 152,421 patient queries and flagged 28,535 (18.7%) as emergencies. The over-escalation rate has been 17.8%, without any increase in missed emergencies. Clinicians have also added 48 new rules since deployment, evidence of the faster correction loop we set out to build.
Shobhit Jagga, Aman Dalmia, Niharika Priyadarshini +7
Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the user should take next. We introduce CARE-Bench, a source-grounded benchmark that evaluates sequential patient-facing triage as a four-label per-turn current-action task. CARE-Bench contains 500 cases and 1,059 evaluated patient-disclosure prefixes reconstructed from medical dialogue, consultation, and follow-up-question sources. We evaluate 11 models on 269 held-out rounds under unprompted and minimally prompted open-ended protocols, using a fixed GPT-5.5 mapper to code each response into the four-label action space. Unprompted macro-F1 remains low, ranging from 31.2 to 50.4. Prompting improves 10 of 11 models, with prompted macro-F1 ranging from 46.9 to 63.4, but substantial threshold errors remain. Prompted models often recommend care before needed clarification is obtained; when the correct action was to ask for more information, only 33.5% of prompted outputs preserved the step. The persistence of these errors after prompting suggests that patient-facing triage is not a simple prompting problem and supports explicit evaluation of action timing before deployment.
Yining Hua, Hongbin Na, Cyrus Ayubcha
1Harvard University · University of Technology Sydney