Clinician use of language models diverges from how the models are evaluated
Organizations: Department of Neurosurgery, NYU Langone Health, New York, NY, USA · Department of Technology Management, NYU Tandon School of Engineering, New York University, New York, NY, USA · Washington University School of Medicine, St. Louis, MO, USA · Global AI Frontier Lab, New York University, New York, NY, USA · Department of Surgery, NYU Langone Health, New York, NY, USA · Department of Orthopedic Surgery, NYU Langone Health, New York, NY, USA · Department of MCIT Health Informatics, NYU Langone Health, New York, NY, USA · Department of Medicine, NYU Langone Health, New York, NY, USA · Division of Dermatology, Department of Medicine, NYU Langone Long Island, Mineola, NY, USA · Johns Hopkins University School of Medicine, Baltimore, MD, USA · Department of Cardiothoracic Surgery, Stanford University School of Medicine, Stanford, CA, USA · Department of Pathology, NYU Grossman School of Medicine, New York, NY, USA · Biomedical Data Science Hub, NYU Langone Health, New York, NY, USA · Department of Management and Organizations, NYU Stern School of Business, New York, NY, USA · Faculty of Medicine, Macau University of Science and Technology, Taipa, Macao, China · Department of Big Data and Biomedical AI, College of Future Technology, Peking University and Peking-Tsinghua Center for Life Sciences, Beijing, China · Department of Population Health, NYU Langone Health, New York, NY, USA · Department of Neuroscience, NYU Langone Health, New York, NY, USA · Department of Radiology, NYU Langone Health, New York, NY, USA · Neuroscience Institute, NYU Langone Health, New York, NY, USA · Center for Data Science, New York University, New York, NY, USA
Abstract
Large language model (LLM) assistants are being deployed to clinicians across health systems, and judgments about their readiness rest largely on benchmark scores, most of them derived from examination questions or curated cases. A benchmark predicts performance in deployment only to the extent that its items resemble real use, yet whether benchmarks reflect the work these systems receive has rarely been measured. Here we analyze 127,833 queries sent by 6,342 physicians, advanced practice providers and nurses in 35 specialties to an institutional assistant during an eight-month roll-out. We characterize each query with RCQ-Map, a clinician-validated framework grounded in taxonomies of clinical questions and of LLM evaluation, which records its task, intent, answerability, missing information and potential harm. Documentation and administration (36.2%) and knowledge retrieval (28.9%) made up nearly two-thirds of use, and diagnosis 3.7%; more than a third of queries could not be answered well as posed. Applying RCQ-Map to 58 public benchmarks drawn from major evaluation suites and frontier model reports, which we assemble into the Clinical AI Benchmark Atlas, showed that the median benchmark contained no documentation requests and shared 31% of the task mix of real use, less than an even spread across task categories would. Benchmarks in suites designed to resemble clinical practice were individually no closer to real use than those used in frontier model reports. Benchmark scores therefore say little about how clinical AI performs on most of the work it is actually given, and evaluation should be matched to real clinical use.
Figures & tables
| All | Attending physicians | Fellows | Residents | APPs | Registered nurses | |
|---|---|---|---|---|---|---|
| Clinicians, n | 6,342 | 1,664 | 157 | 771 | 946 | 2,804 |
| Conversations, n (%) | 127,833 (100.0) | 37,383 (29.2) | 2,121 (1.7) | 23,630 (18.5) | 23,906 (18.7) | 40,793 (31.9) |
| Conversations per clinician, median (IQR) | 4 (1–16) | 4 (1–20) | 4 (1–13) | 8 (2–33) | 6 (2–23) | 3 (1–11) |
| Specialties represented, n | 35 | 35 | 19 | 19 | 27 | — |
| Kind of work (%) | ||||||
| Documentation & administration | 36.2 | 32.1 | 37.3 | 38.0 | 35.5 | 39.1 |
| Block | Field (values) | What it captures | Grounding |
|---|---|---|---|
| Task & domain | Task category (10) | The deliverable that would satisfy the request: drug information; treatment and management; foundational knowledge; patient education; test and result interpretation; documentation and workflow; diagnosis and differential; coding and administrative; procedural guidance; other | Generic clinical-question taxonomies 44,49,51 ; clinical-information-system needs 50,53 ; MedHELM task families 16 ; LLM usage studies 22,26 |
| Question intent (12; 5 groups) | The motive behind the task (verification, fact look-up, clinical decision, dosing, documentation drafting, definition, mechanism, comparison, procedure, coding, result interpretation, other) | Web-search intent 56 ; generic question stems 44 ; well-built clinical question 57 ; question pursuit 43 | |
| Clinical department (25); Medicine division (15) | The department (and Medicine division) that would own management, via a precedence ladder | Institutional ownership for governance and oversight (Methods) | |
| AMA use case (8) | Mapping to the AMA list of physician AI use cases, plus none of the above | AMA physician survey 1 | |
| Query properties | Patient-specific (binary) | About an identifiable patient | Patient-specific needs 49,53 |
| Context | Patient or local information; current evidence (binary) | Information a safe answer needs that the query lacks | Unmet patient-specific needs 53 ; EHR grounding 58 ; resource gaps 59 ; retrieval augmentation 29,30 |
| Field | Krippendorff’s (95% CI) | Gwet’s AC1 | Pairwise agreement (%) | Unanimous (%) | Held-out clinician vs reference (%) | GPT-5.6 sol vs reference (%) | Difference, pts (95% CI) |
|---|---|---|---|---|---|---|---|
| Task & domain | |||||||
| Task category | 0.64 (0.55–0.73) | 0.67 | 70 | 59 | 84 | 79 | -5 (-14 to 4) |
| Clinical domain | 0.48 (0.38–0.57) | 0.60 | 61 | 48 | 79 | 73 | -6 (-17 to 4) |
| Medicine division | 0.43 (0.33–0.52) | 0.59 | 61 | 45 | 74 | 73 | -2 (-12 to 8) |
| Question intent | 0.45 (0.36–0.54) | 0.51 | 55 | 38 | 69 | 67 | -2 (-16 to 10) |
| Question intent (five groups) | 0.61 (0.50–0.70) | 0.68 | 74 | 62 | 84 | 82 | -2 (-11 to 7) |
| Variable | Model 1: OR (95% CI) | Model 2: OR (95% CI) |
|---|---|---|
| Concerns a specific patient | 2.91 (2.64–3.22) | 3.20 (2.88–3.55) |
| Names a medication | 3.12 (2.89–3.37) | 1.30 (1.18–1.44) |
| States a dose | 1.13 (0.97–1.31) | 1.93 (1.70–2.18) |
| Cites a laboratory or test result | 0.93 (0.85–1.01) | 0.99 (0.90–1.09) |
| Mentions imaging | 0.62 (0.53–0.73) | 0.76 (0.68–0.86) |
| Vulnerable population | 1.24 (1.08–1.42) | 1.37 (1.24–1.52) |