Large language models improve physician accuracy but lead to false reliance
Organizations: Division of Digital Prevention, Diagnostics and Therapy Guidance, German Cancer Research Center (DKFZ), Heidelberg, Germany · Medical Faculty, University Heidelberg, Heidelberg, Germany · Skin Cancer Unit, German Cancer Research Center (DKFZ), Heidelberg, Germany · Department of Dermatology, Venereology and Allergology, University Medical Center Mannheim, Ruprecht-Karl University of Heidelberg, Mannheim, Germany · DKFZ Hector Cancer Institute at the University Medical Center Mannheim, Mannheim, Germany · Department of Dermatology, Medical University of Vienna, Vienna, Austria · Department of Dermatology, Escuela de Medicina, Pontificia Universidad Católica de Chile, Santiago, Chile · Clinic and Polyclinic for Dermatology, Venereology and Allergology, University Medical Center Rostock, Rostock, Germany · Department of Medical Oncology, National Center for Tumor Diseases (NCT), Heidelberg University Hospital, Heidelberg, Germany
Abstract
Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.