The autoPET3 Challenge: Automated Lesion Segmentation in Whole-Body PET/CT Multitracer Multicenter Generalization
Organizations: Department of Radiology, LMU University Hospital, LMU Munich, Munich, Germany · Munich Center for Machine Learning (MCML), Munich, Germany · Comprehensive Pneumology Center (CPC-M), Member of the German Center for Lung Research (DZL), Munich, Germany · relAI – Konrad Zuse School of Excellence in Reliable AI, Munich, Germany · Faculty of Mathematics and Computer Science, Heidelberg University, Heidelberg, Germany · Helmholtz Imaging, DKFZ, Heidelberg, Germany · Pattern Analysis and Learning Group, Department of Radiation Oncology, Heidelberg University Hospital, Heidelberg, Germany · Department of Nuclear Medicine, University Hospital Essen (AöR), Essen, Germany · HIDSS4Health - Helmholtz Information and Data Science School for Health, Karlsruhe/Heidelberg, Germany · Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE · Department of Computer Science and Engineering, The Chinese University of Hong Kong, Shatin, Hong Kong SAR · Department of Electronic Engineering, The Chinese University of Hong Kong, Shatin, Hong Kong SAR · University Hospital Tübingen, Department of Radiology, Tübingen, Germany · Cluster of Excellence iFIT (EXC 2180) "Image Guided and Functionally Instructed Tumor Therapies", University of Tübingen, Tuebingen, Germany · Department of Radiology, Stanford University, Stanford, USA · Department of Nuclear Medicine, LMU University Hospital, LMU Munich, Munich, Germany
Abstract
We report the design and results of the third autoPET challenge (MICCAI 2024), which benchmarked automated lesion segmentation in whole-body PET/CT under a compositional generalization setting. Training data comprised 1,014 [18F]-FDG PET/CT studies from the University Hospital Tübingen and 597 [18F]/[68Ga]-PSMA PET/CT studies from the LMU University Hospital Munich, constituting the largest publicly available annotated PSMA PET/CT dataset to date. The held-out test set of 200 studies covered four tracer-center combinations, two of which represented unseen compositional pairings. A complementary data-centric award category isolated the contribution of data handling strategies by restricting participants to a fixed baseline model. Seventeen teams submitted 27 algorithms, predominantly nnU-Net-based 3D networks with PET/CT channel concatenation. The top-ranked algorithm achieved a mean DSC of 0.66, FNV of 3.18 mL, and FPV of 2.78 mL across all four test conditions, improving DSC by 8% and reducing the false-negative volume by 5 mL relative to the provided baseline. Ranking was stable across bootstrap resampling and alternative ranking schemes for the top tier. Beyond the benchmark, we provide an in-depth analysis of segmentation performance at the patient and lesion level. Three main conclusions can be drawn: (1) in-domain multitracer PET/CT segmentation is sufficient and probably approaching reader agreement; (2) compositional generalization to unseen tracer-center combinations remains an open problem mainly driven by systematic volume overestimation; (3) heterogeneity and case difficulty drive performance variation substantially more than the choice of algorithm among top-ranked teams.