Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
Organizations: National Institute on Drug Abuse Intramural Research Program, National Institutes of Health, Baltimore, MD, USA · National Library of Medicine, National Institutes of Health, Bethesda, MD, USA · School of Information Sciences, University of Illinois Urbana-Champaign, Champaign, IL, USA
Abstract
Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant studies, yet the quality of retrieved evidence and factors influencing study selection remain unclear. We evaluated three general-purpose LLM chatbots (Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5) using 20 clinical questions adapted from 2026 Cochrane reviews. We simulated patient, clinician, and evidence-synthesis researcher roles and obtained four independent responses for each chatbot-role-question combination, yielding 720 responses (3 chatbots 3 user roles 4 repetitions 20 review questions). Chatbots were asked to support their answers with primary clinical citations, which were benchmarked against the included and excluded study sets of the corresponding Cochrane reviews. On average, a single response retrieved 39.2% 29.8% of the corresponding Cochrane included-study set and 5.0% 9.4% of the excluded-study set. Recall of included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% 29.5% vs. 37.0% 23.8% vs. 17.3% 13.1%; blocked permutation test, ), and the researcher role yielded higher recall than the clinician or patient roles (42.8% 30.8% vs. 38.6% 28.9% vs. 36.1% 29.3%; ). Controlling for publication year, citations per year, and open-access status, sample size was the only significant predictor of retrieval: each doubling of sample size was associated with 50% higher odds of retrieval (odds ratio 1.50, 95% CI 1.24-1.81). These findings show that LLM chatbots can retrieve studies identified by expert reviewers, but retrieval varies substantially across models and user roles and favors larger clinical trials.