A Living Benchmark for Information Retrieval from Electronic Health Records
Organizations: Department of Biomedical Data Science, Stanford University, Stanford, CA · Department of Pathology, Stanford University, Stanford, CA · Department of Anesthesiology, Perioperative and Pain Medicine, Stanford University, Stanford, CA · Department of Medicine, Stanford University, Stanford, CA · Department of Surgery, Stanford University, Stanford, CA · Department of Radiation Oncology, Stanford Cancer Center, Palo Alto, CA, USA · Weill Cancer Hub West · Center for Clinical Excellence Research, Stanford School of Medicine, Stanford, CA, USA · Department of Computer Science, Stanford University, Stanford, CA
Abstract
Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset. Across nine LLMs and five inference strategies, state-of-the-art systems frequently omit clinically important information, particularly for questions requiring synthesis across multiple documents and encounters. Because the generator itself is validated, BRIE supports evaluations that static benchmarks cannot, including the generation of multiple answers that reflect variation in clinician reasoning for robust performance assessment and continuously refreshing benchmark content to guard against leakage. Our results demonstrate that scalable benchmark generation enables rigorous, up-to-date evaluation of clinical LLMs as they are deployed in rapidly evolving healthcare settings.
Figures & tables
| Stage | Count | Token Count |
| Raw Note | ||
| Fact | ||
| Refined Facts |
| Category | n | Tokens |
| General function words | 113 | a, an, the, and, or, but, in, on, at, to, for, of, with, by, from, up, about, into, through, during, before, after, above, below, between, out, off, over, under, then, here, there, when, where, how, all, both, each, more, most, other, some, such, no, nor, not, only, same, so, than, too, very, can, will, just, now, this, that, these, those, is, are, was, were, be, been, being, have, has, had, do, does, did, would, could, should, may, might, must, shall, also, its, it, he, she, they, we, i, my, his, her, their, our, your, which, who, whom, as, if, while, although, however, therefore, thus, since, because, well, within, without, including, via, per, see |
| Clinical narrative & charting | 45 | patient, pt, px, history, hx, assessment, plan, reported, reports, noted, notes, follow, followup, discharge, admission, admitted, presents, presenting, presented, new, old, previous, prior, current, recent, stable, unchanged, medical, surgical, social, family, review, reviewed, discussed, visit, appointment, clinic, office, normal, abnormal, negative, positive, right, left, bilateral |
| Medication & dosing | 22 | tablet, tablets, tab, capsule, capsules, cap, oral, daily, prn, mg, ml, mcg, mgs, mls, kg, dose, doses, dosing, prescribed, unit, units, cc |
| Temporal & measurement | 11 | year, years, month, months, week, weeks, day, days, date, time, status |
| Characteristic | n | % | Ref. % | |
| Gender | ||||
| Male | 44 | 58.7 | 49.9 | |
| Female | 31 | 41.3 | 50.1 | |
| Ethnicity and race | ||||
| Not Hispanic or Latino | White | 40 | 53.3 | 49.7 |
| Hispanic or Latino | No matching concept | 14 | 18.7 | 13.6 |
| Filtering stage | Remaining | Removed | % of original |
| Original | 750 | — | 100.0 |
| Two annotators | 675 | 75 | 90.0 |
| Consistent | 660 | 15 | 88.0 |
| Relevant questions | 550 | 110 | 73.3 |
| One annotator accepted the answer | 508 | 42 | 67.7 |
| Both annotators accepted the answer | 308 | 200 | 41.1 |
| Model | Official context | Buffer | Tiktoken limit |
| Gemini 2.5 Pro | 1,000,000 | 50,000 | 950,000 |
| Gemini 2.5 Flash Lite | 1,000,000 | 50,000 | 950,000 |
| Claude Opus 4.7 | 1,000,000 | 75,000 | 925,000 |
| Claude Haiku 4.5 | 200,000 | 50,000 | 150,000 |
| GPT 5.4 | 272,000 | 10,000 | 262,000 |
| GPT 5.4 Nano | 272,000 | 10,000 | 262,000 |
| Model | Mean | SD | Min | 25% | Median | 75% | Max |
| Claude Opus | 4.17 | 3.31 | 0 | 2 | 3 | 5 | 25 |
| Claude Haiku | 6.18 | 4.97 | 0 | 2 | 4 | 9 | 35 |
| Model | Tool | Mean | SD | Min | 25% | Median | 75% | Max |
| Claude Opus | search_notes | 0.57 | 0.13 | 0.00 | 0.50 | 0.50 | 0.67 | 0.93 |
| get_note | 0.34 | 0.20 | 0.00 | 0.17 | 0.50 | 0.50 | 0.83 | |
| summarize_notes | 0.09 | 0.17 | 0.00 | 0.00 | 0.00 | 0.00 | 0.67 | |
| Claude Haiku | search_notes | 0.57 | 0.17 | 0.00 | 0.50 | 0.50 | 0.67 | 1.00 |
| get_note | 0.42 | 0.17 | 0.00 | 0.33 | 0.50 | 0.50 | 0.86 | |
| summarize_notes | 0.01 | 0.05 | 0.00 | 0.00 | 0.00 | 0.00 | 0.50 |
| Inference type | Rate | 95% CI |
| Recent | 0.99 | [0.99, 1.00] |
| Recent-200K | 0.99 | [0.98, 1.00] |
| Dense | 0.99 | [0.97, 1.00] |
| BM25 | 0.99 | [0.99, 1.00] |
| Agent | 0.99 | [0.94, 1.00] |
| Price ($/1M) | Cost per run ($) | |||||
| Model | In | Out | Full ctx | 180K | 50K | 25K |
| Claude Haiku 4.5 | 1.00 | 5.00 | 0.2013 | 0.1813 | 0.0513 | 0.0263 |
| Claude Opus 4.7 | 5.00 | 25.00 | 5.0013 | 0.9013 | 0.2513 | 0.1263 |
| Gemini Flash Lite 2.5 | 0.10 | 0.30 | 0.1013 | 0.0193 | 0.0063 | — |
| Gemini Pro 2.5 | 2.50 | 15.00 | 2.5013 | 0.4513 | 0.1263 | — |
| GPT-5.4 | 5.00 | 22.50 | 1.3613 | 0.9013 | 0.2513 | — |