q-bio.GNSep 17, 2026

Large Language Model Agents for Evidence Based Genetic Disease Severity Classification

Authors: Tohid Ghasemnejad, Ahmadreza Argha, Mark Grosser, John Wang, Min Yang, Thantrira Porntaveetus, Tony Roscioli, Nigel H. Lovell, +2 more

Organizations: UNSW BioMedical Machine Learning Lab (BML), School of Biomedical Engineering, UNSW Sydney, Sydney, NSW 2052, Australia · School of Biomedical Engineering, UNSW Sydney, NSW 2052, Australia · 23Strands, Pyrmont, Australia · Division of Genetic and Genomic Medicine, Department of Pediatrics, University of Pittsburgh, School of Medicine, Pittsburgh, PA, USA · Shenzhen Institute of Advanced Technology, Shenzhen, China · Center of Excellence in Precision Medicine and Digital Health, Department of Physiology, Faculty of Dentistry, Chulalongkorn University, Bangkok, Thailand · New South Wales Health Pathology Genomics, Prince of Wales Hospital, Randwick, NSW, Australia · Neuroscience Research Australia (NeuRA), University of New South Wales Sydney, NSW, Australia · Medical Genetics and Genomics Laboratories, University of Pittsburgh Medical Center, Pittsburgh, PA, USA · Departments of Pathology, and Obstetrics, Gynecology, and Reproductive Sciences, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA · Visiting Scholar (Collaborative Projects), Center of Excellence in Precision Medicine and Digital Health, Chulalongkorn University, Bangkok, Thailand

Abstract

Disease severity classification for genetic conditions is subjective and labor-intensive, creating bottlenecks in genomic screening, where commercial panels vary widely in size and overlap. We developed an autonomous AI agent integrating Reasoning and Acting (ReAct) with Retrieval-Augmented Generation (RAG) to classify 10,211 Human Phenotype Ontology terms. It uses American College of Medical Genetics (ACMG)-endorsed severity guidelines and American College of Obstetricians and Gynecologists (ACOG) quality-of-life criteria to retrieve PubMed literature, generate interpretable reasoning chains, and independently verify claims. At the phenotype level, using expert-curated cohorts, the agent achieved 93.55% accuracy (MCC 0.9237) with 82.6% to 91.4% of claims supported by direct evidence or valid inferences. Gene-level severity was aggregated across 8,738 pairs, identifying 3,283 autosomal recessive pairs with severe or profound presentations. External validation showed 95.2% concordance with Mackenzie's Mission gene list. This system enables standardized panel design by providing reliable, automated classification supported by direct evidence.

Explore similar work

CardsList
  1. Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)

    Jul 6, 2026Manuela Del Castillo Suero, Arnault-Quentin Vermillet, Nicole Sonne Heckmann +2Calibrated SeverityElectronic Health Records