cs.IRAug 13, 2026

Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

Authors: Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu

Organizations: National Institute on Drug Abuse Intramural Research Program, National Institutes of Health, Baltimore, MD, USA · National Library of Medicine, National Institutes of Health, Bethesda, MD, USA · School of Information Sciences, University of Illinois Urbana-Champaign, Champaign, IL, USA

Abstract

Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant studies, yet the quality of retrieved evidence and factors influencing study selection remain unclear. We evaluated three general-purpose LLM chatbots (Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5) using 20 clinical questions adapted from 2026 Cochrane reviews. We simulated patient, clinician, and evidence-synthesis researcher roles and obtained four independent responses for each chatbot-role-question combination, yielding 720 responses (3 chatbots ×\times 3 user roles ×\times 4 repetitions ×\times 20 review questions). Chatbots were asked to support their answers with primary clinical citations, which were benchmarked against the included and excluded study sets of the corresponding Cochrane reviews. On average, a single response retrieved 39.2% ±\pm 29.8% of the corresponding Cochrane included-study set and 5.0% ±\pm 9.4% of the excluded-study set. Recall of included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% ±\pm 29.5% vs. 37.0% ±\pm 23.8% vs. 17.3% ±\pm 13.1%; blocked permutation test, p=2.0×10−5p=2.0\times10^{-5}), and the researcher role yielded higher recall than the clinician or patient roles (42.8% ±\pm 30.8% vs. 38.6% ±\pm 28.9% vs. 36.1% ±\pm 29.3%; p=2.0×10−5p=2.0\times10^{-5}). Controlling for publication year, citations per year, and open-access status, sample size was the only significant predictor of retrieval: each doubling of sample size was associated with 50% higher odds of retrieval (odds ratio 1.50, 95% CI 1.24-1.81). These findings show that LLM chatbots can retrieve studies identified by expert reviewers, but retrieval varies substantially across models and user roles and favors larger clinical trials.

Explore similar work

CardsList
  1. When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG

    Jun 2, 2026Erfan Nourbakhsh, Rocky Slavin, Ke Yang +1

  2. Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency

    May 18, 2026Matthew L. Smith, Jonathan P. Shock, Samuel T. Segun +2Large Language Model ReliabilityFactual Recall

  3. When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering

    May 20, 2026Doeun Lee, Muge Zhang, Yi Yu +11Clinical Reasoning TrainingQuestion-Answering Benchmarks