cs.AIJun 3, 2026

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison

Authors: Alejandro LozanoKeiko IharaPing-Hao YangCarrie E. RobertsonJennifer SternAllan PurdyHsiangkuo YuanPengfei Zhang+8 more

Organizations: Stanford University, Palo Alto, CA, USA · Department of Neurology, Mayo Clinic, Rochester, MN, USA · Department of Neurology, Dalhousie University, Halifax, Canada · Jefferson Headache Center, Department of Neurology, Thomas Jefferson University, PA, USA · Beth Israel Deaconess Medical Center, Boston, MA, USA · Department of Neurology, University of Florida, Gainesville, FL, USA · University of Colorado School of Medicine, Department of Pediatrics, Division of Child Neurology, Aurora, CO, USA · Department of Medicine, Mount Sinai Hospital, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Department of Neurology, Mayo Clinic, Scottsdale, AZ, USA · Harvard Medical School, Boston, MA, USA · Department of Neurology, Mount Sinai Hospital, Icahn School of Medicine at Mount Sinai, New York, NY, USA

Abstract

Summarizing the latest medical literature to guide clinical decision-making is essential for evidence-based medicine and high-quality patient care. Yet clinicians face increasing challenges due to limited time with patients and a rapidly growing volume of published articles. Although retrieval-augmented large language models (LLMs) have shown promise in clinical summarization, human evaluations of their effectiveness in synthesizing broader scientific literature and direct comparisons to expert-written syntheses remain scarce. We constructed a RAG-based agentic AI framework using three state-of-the-art LLMs: Sonnet, GPT-4o, and Llama 3.1. A headache specialist created 13 questions, three for prompt optimization and ten for evaluation. Ten headache specialists across the United States and Canada each wrote a summary for one question, yielding four summaries per question (expert, Sonnet, GPT-4o, and Llama). The experts, blinded to authorship, critically evaluated the summaries, excluding the topic for which they wrote a summary, based on correctness, completeness, conciseness, and clinical utility, scoring each from 1 to 10 using standardized rubrics. They also ranked the summaries by preference and indicated whether they believed each summary was written by an expert or an LLM. Our study, comparing LLM- and expert-written literature summaries evaluated by headache specialists, showed that expert-written summaries were preferred, although experts sometimes found it challenging to distinguish between human- and AI-generated summaries. We also identified key expert-valued features beyond standard evaluation metrics that can guide future refinement of both human and AI literature summarization pipelines.

Explore similar work

CardsList