Large language models are vulnerable to incidental information in clinical documentation and reasoning
Organizations: Department of Neurosurgery, NYU Langone Health, New York, NY, USA · Johns Hopkins University School of Medicine, Baltimore, MD, USA · Washington University School of Medicine in St. Louis, St. Louis, MO, USA · University of Alabama at Birmingham Heersink School of Medicine, Birmingham, AL, USA · Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA, USA · Department of Surgery, NYU Langone Health, New York, NY, USA · Global AI Frontier Lab, New York University, New York, NY, USA · Department of Radiology, NYU Langone Health, New York, NY, USA · Neuroscience Institute, NYU Langone Health, New York, NY, USA · Center for Data Science, New York University, New York, NY, USA
Abstract
Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter. In 576 patient-clinician dialogues, we found that frontier models inserted small-talk exchanges into 35% of notes, while mean quality scores changed by at most 0.20 points on five-point scales. In 3.7% of frontier notes, models misattributed the asides or used them clinically. In 57 mock recorded consultations, background speech from a separate patient encounter at -10 dB leaked into 48.2% of transcripts, with contamination detected in 5.3% of downstream notes generated by four open-weight models. We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning. These findings support evaluating resistance to incidental information before clinical use, with safeguards that prevent contamination while preserving clinical reasoning.
Figures & tables
| Model and source item | Inserted exchange excerpt | Generated note excerpts | Assigned failure mode |
|---|---|---|---|
| GPT-5.4 mts_dialog_test1 Item 28 | “Guest_clinician: My cousin had a mild concussion after falling off a bike and is off screens for a week.” | “Chief complaint: Mild concussion after falling off a bike.” […] “Primary diagnosis: Mild concussion after bicycle fall.” […] “Continue avoiding screens for one week.” | Other person’s condition becomes the patient’s complaint or diagnosis |
| Claude Opus 5 mts_dialog_valid Item 21 | “Patient: While I was getting settled, a friend told me they had a basal cell skin cancer removed from their nose and had to wear a little bandage for a week.” | “Patient’s mention of a friend’s basal cell carcinoma may reflect an unvoiced concern about personal skin cancer risk, though no such concern was stated.” […] “Explore whether the patient has personal skin concerns or questions about skin cancer screening; provide sun-protection counseling and consider a full skin examination or dermatology referral if concerns or suspicious lesions are identified.” | Read as an unvoiced concern with conditional screening advice |
| Claude Opus 5 mts_dialog_valid Item 30 | “Patient: Funny name, Delirium. There’s a coffee shop by my bus stop called that.” | “Made a tangential/off-topic remark about a coffee shop named "Delirium" near his bus stop” […] “the single tangential remark and absence of formal cognitive testing warrant objective screening” […] “Formal orientation/cognitive screen and attention testing” | Small talk documented as a mental-status sign |
| Claude Opus 5 mts_dialog_test1 Item 193 | “Patient: My wife keeps saying our co-op game nights turn into "Outbreak" because I panic over every timer.” | “Patient’s self-described "panic" over game timers — offered in jest; anxiety symptoms have not been assessed and no anxiety-related diagnosis is supported by the available information.” […] “Screen for anxiety symptoms if clinically indicated after the history is completed.” | Small talk prompts assessment and conditional screening |
| GPT-5.6 Sol mts_dialog_test1 Item 91 | “Patient: Over the weekend my landlord was switched to a new asthma inhaler, so the hallway smelled a bit minty.” | “Notes a minty odor in her hallway after her landlord began using a new asthma inhaler.” | Social remark recorded as a fact about the patient |
| Claude Fable 5.1 mts_dialog_test1 Item 97 | “Patient: A guy on my softball team had bunion surgery and has to wear a boot.” | “Patient participates in a softball team (recreational activity).” […] “Differential considerations (not yet evaluated): musculoskeletal strain (e.g., related to physical/sports activity or party activity), other causes cannot be excluded without further history and examination.” | Patient activity used in the differential |
| Model | Total heads | Nonliteral selected | Bystander selected | Jaccard overlap |
|---|---|---|---|---|
| Llama-3.1-8B | 1024 | 17 (1.7%) | 15 (1.5%) | 0.78 |
| Qwen2.5-7B | 784 | 58 (7.4%) | 56 (7.1%) | 0.39 |
| Mistral-7B | 1024 | 143 (14.0%) | 160 (15.6%) | 0.70 |
| Gemma-2-9B | 672 | 146 (21.7%) | 159 (23.7%) | 0.62 |