Three-Phase Scribble-Adaptive Curriculum Learning for autoPETV Grand Challenge
Organizations: Duke University
Abstract
This report describes Libo Zhang's algorithmic solution to autoPETV Grand Challenge on interactive lesion segmentation in whole-body PET/CT. Interaction is encoded as two additional input channels that rasterize the accumulated foreground and background scribbles, and a residual-encoder U-Net of about 140 million parameters is trained with a three-phase curriculum over 4000 epochs: the network first learns fully automatic segmentation with silent interaction channels, then observes ground-truth-derived scribbles under randomly sampled visibility modes, and finally adapts to its own mistakes through online simulation of up to five error-driven correction steps. Training draws on 1811 autoPET and DeepPSMA studies, and the submission ensembles the best and final checkpoints of five folds by logit averaging. In interactive five-fold cross-validation with six interaction steps, the final checkpoints reach a mean AUC-Dice of 3.836 and a mean AUC-DMM of 3.869, improving monotonically in every fold, with roughly half of the total gain delivered by the first corrective scribble. Our code and trained model checkpoints are available on https://github.com/Libo1023/autoPETV-Curriculum.
Figures & tables
| Team name | LiboZhang |
| Algorithm name | autoPETV-Curriculum |
| Data pre-processing | DeepPSMA CT linearly resampled to the PET grid; all channels resampled to the median spacing of mm (spline order 3 for images, linear for labels); CT clipped to and z-scored (global mean 126.7, standard deviation 286.7); PET and both interaction channels z-scored per image; scribbles rasterized as two binary heatmaps |
| Data post-processing | Arg-max of the ensemble-averaged logits |
| Training data augmentation | Default nnU-Net pipeline on all four channels (rotation , scaling 0.7 to 1.4, Gaussian noise and blur, brightness and contrast 0.75 to 1.25, simulated low resolution, gamma, mirroring on all axes); three-phase interaction curriculum of Sect. 2.3 |
| Standardized framework | nnU-Net v2 (2.6.0) on PyTorch 2.6.0 with CUDA 12.4 |
| Network architecture | 3D residual-encoder U-Net at a 40 GB compute budget: 7 stages with 32 to 320 features, encoder depths , deep supervision, about 140M parameters, each checkpoint’s size is roughly 1.1 GB |
| AUC-Dice | AUC-DMM | ||||||
|---|---|---|---|---|---|---|---|
| Fold ( ) | Ckpt | All | FDG | PSMA | All | FDG | PSMA |
| 0 (207) | B | 3.8739 | 4.0525 | 3.7070 | 3.9011 | 3.6764 | 4.1110 |
| F | 3.9087 | 4.0936 | 3.7359 | 3.9581 | 3.7940 | 4.1114 | |
| 1 (215) | B | 3.7320 | 4.0672 | 3.4565 | 3.8039 | 3.6299 | 3.9470 |
| F | 3.7532 | 4.0726 | 3.4907 | 3.8490 | 3.7181 | 3.9565 | |
| 2 (213) | B | 3.7793 | 4.1725 | 3.4111 | 3.7941 | 3.7117 | 3.8712 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.