cs.CVJul 31, 2026

Performance of large language models in the optical diagnosis of colorectal polyps

Authors: Joshua C. VencesWilliam T. TranNikko GimpayaCatharine M. WalshRishad J. KhanRobert BecharaAsher C. WigginsCeline N. Rousan+10 more

Organizations: 1. Scarborough Health Network Research Institute, Toronto, Ontario, Canada · 2. Division of Gastroenterology, Hepatology and Nutrition, and The SickKids Learning and Research Institutes, The Hospital for Sick Children, Toronto, Ontario, Canada · 3. Department of Paediatrics, Temerty Faculty of Medicine, University of Toronto, Toronto, Ontario, Canada · 4. Division of Gastroenterology, University of Calgary · 5. Division of Gastroenterology, Department of Medicine, Queen’s University and Kingston Health Sciences Centre, Kingston, Ontario, Canada · 6. Department of Medicine, University of Toronto, Toronto, Ontario, Canada · 7. Scarborough Health Network, Division of Hematology, Toronto, Ontario, Canada · 8. Université de Montréal, Department of Medicine, Montreal, Quebec, Canada · 9. Centre de Recherche du Centre hospitalier de l’Université de Montréal (CRCHUM), Montreal, Quebec, Canada · 10. Interventional and Experimental Endoscopy (InExEn), Department of Internal Medicine II, University Hospital Würzburg, Würzburg, Germany · 11. Division of Gastroenterology and Hepatology, Mayo Clinic, Rochester, Minnesota, United States · 12. Université de Sherbrooke, Faculty of Medicine and Health Sciences, Sherbrooke, Quebec, Canada · 13. Centre Hospitalier Universitaire de Sherbrooke, Division of Gastroenterology, Sherbrooke, Quebec, Canada · 14. Division of Gastroenterology and Hepatology, University of Toronto, Toronto, Ontario, Canada

Abstract

Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology. Methods: We conducted a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging (NBI) images. We evaluated Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic (NICE), and predicted histology, we calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran's Q and McNemar's Test were used to determine differences between predicted values of each MLLM. Results: The F1 scores among MLLMs were >0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification. Conclusions: Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment.

Explore similar work

CardsList
  1. Understanding Model Behavior in Monocular Polyp Sizing

    May 19, 2026Xinqi Xiong, Andrea Dunn Beltran, Junmyeong Choi +3Polyp SegmentationColonoscopy