cs.CLSep 19, 2026

LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage

Authors: Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, +1 more

Organizations: University of New South Wales, Sydney · Independent Researcher · St Vincent’s Hospital, Sydney · Royal Prince Alfred Hospital, Sydney · Campbelltown Hospital, Sydney

Abstract

Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425 LLM-generated (SIMULATED) and 50 physician-authored (CLINICIAN) conversations, both labelled under the Emergency Severity Index (ESI). Every model, measured by quadratic weighted kappa (QWK), degrades from moderate-to-substantial agreement on completed records to fair-to-moderate agreement at every sequential checkpoint. Controlled perturbations show that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lift this plateau. Models extract clinically relevant content from later turns, yet the surprisal of the true label rises across the checkpoints. So the model fails to integrate the evidence. Three expert clinicians on the same conversations reach a QWK of 0.887-0.929, while the best model reaches 0.295. Predictions concentrate at ESI-2 and ESI-3, and models agree with each other more than with the ground truth, so ensembling worsens the failure. Deploying LLMs for ED triage based on offline benchmarks alone misses this sequential failure.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

    Sep 29, 2026Abinitha Gourabathina, Haoran Zhang, Yuexing Hao +2Triage

  2. Auditable Emergency Triage for Maternal and Newborn Care in India

    Sep 10, 2026Shobhit Jagga, Aman Dalmia, Niharika Priyadarshini +7TriageInterpretability

  3. CARE-Bench: Benchmarking Patient-Facing LLM Triage

    Aug 4, 2026Yining Hua, Hongbin Na, Cyrus AyubchaTriageDisclosures