cs.CLOct 6, 2026

Large language models are vulnerable to incidental information in clinical documentation and reasoning

Authors: Krithik Vishwanath, Brandon Ye, Anton Alyakin, John E. Markert, Aaron Hsieh, Michał Mańkowski, Eric K. Oermann

Organizations: Department of Neurosurgery, NYU Langone Health, New York, NY, USA · Johns Hopkins University School of Medicine, Baltimore, MD, USA · Washington University School of Medicine in St. Louis, St. Louis, MO, USA · University of Alabama at Birmingham Heersink School of Medicine, Birmingham, AL, USA · Perelman School of Medicine at the University of Pennsylvania, Philadelphia, PA, USA · Department of Surgery, NYU Langone Health, New York, NY, USA · Global AI Frontier Lab, New York University, New York, NY, USA · Department of Radiology, NYU Langone Health, New York, NY, USA · Neuroscience Institute, NYU Langone Health, New York, NY, USA · Center for Data Science, New York University, New York, NY, USA

Abstract

Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by assessing their sensitivity to information incidental to the patient encounter. In 576 patient-clinician dialogues, we found that frontier models inserted small-talk exchanges into 35% of notes, while mean quality scores changed by at most 0.20 points on five-point scales. In 3.7% of frontier notes, models misattributed the asides or used them clinically. In 57 mock recorded consultations, background speech from a separate patient encounter at -10 dB leaked into 48.2% of transcripts, with contamination detected in 5.3% of downstream notes generated by four open-weight models. We propose a dual encoding hypothesis of clinical reasoning and distraction in LLMs, with preliminary evidence that LLM components associated with disruption by incidental information also support clinical reasoning. These findings support evaluating resistance to incidental information before clinical use, with safeguards that prevent contamination while preserving clinical reasoning.

Figures & tables

Explore similar work

CardsList
  1. Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale

    May 18, 2026Jinghui Liu, Sarvesh Soni, Anthony NguyenInternational Classification Of Disease CodingNotes

  2. Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text

    Jun 16, 2026Hongbo Du, Zixin Lu, Jiaming QuLarge Language Model Uncertainty

  3. AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows

    Jun 16, 2026Jiahui Niu, Huizi Yu, Wenkong Wang +11Large Language Model EvaluationReal-World Clinical Workflows