Sep 13, 2026 · cs.CLJ/K move · Enter open · S save
Suzannah E McKinney, Phuc Vu, Samuel A Justice, Christopher Humphries+6
Mass General Brigham AI, Mass General Brigham, Boston, MA, USA · Generative AI Laboratory, School of Informatics, University of Edinburgh, Edinburgh, UK · University of Edinburgh Centre for Cardiovascular Science, Edinburgh, UK · Experimental Medicine Division, Nuffield Department of Medicine, University of Oxford, Oxford, UK · Department of Biomedical Data Science, Stanford University School of Medicine, Stanford University, Stanford, CA, USA
Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.