q-bio.QMAug 7, 2026

Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications

Authors: Kaela KokkasHairong WangRichard KleinNazir A. IsmailNatalie IrwinMohammad Z. MoonsamyKubendran NaidooJeremy Nel+6 more

Organizations: Department of Clinical Microbiology and Infectious Diseases, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa · School of Computer Science and Applied Mathematics, University of the Witwatersrand, Johannesburg, South Africa · Wits Machine Intelligence and Neural Discovery (MIND) Institute, University of the Witwatersrand, Johannesburg, South Africa · Department of Clinical Microbiology and Infectious Diseases, National Health Laboratory Service and Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa · Division of Medical Oncology, Department of Internal Medicine, University of the Witwatersrand, Johannesburg, South Africa · Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa · South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa · South African Medical Research Council Wits Antiviral Gene Therapy Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa · National Health Laboratory Service, Johannesburg, South Africa · Department of Molecular Medicine and Haematology, School of Pathology, University of the Witwatersrand, Johannesburg, South Africa · Wits Research Institute for Malaria, Faculty of Health Sciences, National Health Laboratory Service, University of the Witwatersrand, Johannesburg, South Africa · Division of Infectious Diseases, School of Clinical Medicine, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa · Department of Surgery, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa · Division of Virology, University of the Witwatersrand & National Health Laboratory Service, Johannesburg, South Africa · OncoVectra, London, United Kingdom · Department of Global Health, Rollins School of Public Health, Emory University, Atlanta, GA United States

Abstract

Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will reduce disease burdens. However, relevant evidence is dispersed and infeasible for humans to comprehensively synthesize. LLMs may enable scalable, expert-level systematic evidence synthesis to identify microbe-cancer pairs; however, such capabilities have not yet been demonstrated. Domain experts were recruited to create a dataset to benchmark LLM performance (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano) on 24 research papers using MMTV-LV and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal, consisting of MCQ, Likert-scale, multi-select, and free-text question types (77 items across 24 papers). Agreement between (1) experts and (2) experts and each LLM was determined per question instance using novel metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement distributions to determine whether LLMs behaved as additional experts by increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively. Across all question types, LLM responses aligned closely with experts, with GPT-5 and GPT-5 Nano achieving score distributions indistinguishable from experts. Gemini models behaved similarly but were significantly more lenient in applying microbial oncogenesis criteria. Hallucinations were rare. Methodological appraisal and identification of contradictions within full-texts were the most persistent LLM vulnerabilities. GPT-5 and GPT-5 Nano were indistinguishable from experts on structured domain research paper evaluation tasks. This supports use of LLMs for automated systematic evidence synthesis. However, methodological appraisal tasks and contradiction identification in full-texts remain weaknesses requiring strengthening.

Explore similar work

CardsList
  1. Agentic systems for breast cancer treatment recommendations

    Jul 13, 2026Vinicius Anjos de Almeida, Nícolas Henrique Borges, Leonardo Vicenzi +5Clinical Decision SupportAgentic Systems