MonteRET: AI Agent Enhancing Multimodal LLMs with Multi-granularity Knowledge Retrieval for Chest CT Report Generation
Organizations: Department of Population Health Sciences, Weill Cornell Medicine, New York, USA · School of Physics, Mathematics and Computing, University of Western Australia, Crawley, Australia · Department of Radiology, Weill Cornell Medicine, New York, USA · Biostatistics and Health Data Science, School of Medicine, Indiana University, Indianapolis, USA · Regenstrief Institute, Indianapolis, USA · Department of Biomedical Informatics and Data Science, Yale School of Medicine, Yale University, New Haven, USA · Department of Computational Medicine, University of California, Los Angeles, USA
Abstract
Automated chest CT report generation remains challenging because clinically faithful reporting requires both whole-volume understanding and accurate description of localized anatomical findings. Here we developed and retrospectively evaluated MonteRET, a region-aware retrieval-enhanced framework for generating chest CT findings sections. MonteRET integrates global CT features with region-level anatomical representations, retrieves clinically relevant knowledge using predicted medical conditions and region-level vision-language alignment, and refines initial reports through a knowledge-guided report rewriting agent. We trained our model on a public cohort with 24,128 CT scans from RadGenome-ChestCT. We evaluated MonteRET on the public RadGenome-ChestCT test set of 1,564 CT scans and an external cohort of 82 CT scans from NewYork-Presbyterian/Weill Cornell Medical Center. MonteRET improved report quality, semantic similarity, and clinical efficacy compared with a matched baseline and several state-of-the-art methods. Gains were most pronounced for recall, suggesting fewer omitted findings. Human expert evaluation by radiology residents also favored MonteRET.