Organizations: Department of Electrical and Computer Engineering, Vanderbilt University, Nashville, TN 37235, USA · Department of Radiology, Weill Cornell Medicine, New York, NY 10021, USA · Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37235, USA · Department of Pathology, Microbiology and Immunology, Vanderbilt University Medical Center, Nashville, TN 37235, USA · Department of Computer Science, Vanderbilt University, Nashville, TN 37235, USA
Accurate nuclei instance segmentation is essential for quantitative renal pathology, yet general-purpose models often struggle with low contrast, dense nuclei, complex morphology, and strong background staining. In this work, we extended a human-in-the-loop framework by combining 5,901 foundation-model-generated pseudo-labels from well-segmented cases (Easy), 860 newly expert-annotated unresolved challenging cases (Medium), and 198 expert-annotated consensus failure cases (Hard). These annotations, spanning different levels of segmentation difficulty, enabled the systematic evaluation of seven single-source and mixed-source fine-tuning strategies across nine cell segmentation model configurations. Fine-tuning improved all models, with Medium data included in seven of the nine best-performing strategies. LSP-DETR achieved the highest F1 score of 0.8725 with Hard-only fine-tuning, while StarDist showed the largest improvement, increasing from 0.7380 to 0.8332 with Medium-only fine-tuning. These findings show that annotations spanning multiple difficulty levels support effective model adaptation, although the optimal annotation composition remains model dependent.
Figures & tables
Figure 1 : Overall framework. (A) Expert annotation of previously unresolved Medium-quality kidney pathology patches to enrich the existing Easy and Hard fine-tuning data. (B) Expanded evaluation of cell nuclei segmentation models released from 2022 to 2026, including both previously evaluated and newly incorporated models. (C) Fine-tuning with seven single-source and mixed-source combinations of Easy, Medium, and Hard datasets to evaluate model-specific adaptation performance.
Figure 2 : Representative examples of prediction-rating categories. Foundation-model predictions were categorized as Good, Medium, or Bad according to the proportion of identifiable nuclei captured in each image patch. Green contours denote nucleus instances; the top row shows model predictions, whereas the bottom row shows the corresponding ground-truth annotations.
Figure 3 : Human-in-the-loop annotation and construction of multi-source fine-tuning data. (A) Pathology experts screened unresolved Medium-quality kidney patches, selected challenging samples, and manually annotated nucleus instances from the original images. (B) The final fine-tuning data consisted of three annotation sources: prior foundation-model-generated pseudo-labels from Easy cases, prior expert-annotated consensus failure cases (Hard), and newly human-in-the-loop annotated unresolved cases (Medium).
Dataset
Count
Annotation Source
Easy
5,901
Foundation-model-generated pseudo-labels
Medium
860
Unresolved Medium-quality patches annotated by experts
Hard
198
Prior foundation-model failure patches annotated by experts
Table 2 : Summary of the annotation sources used for model fine-tuning.
Fine-tuning Strategies
Samples
Original Ratio (%)
Expected Ratio (%)
Easy
5,901
100.0
100.0
Medium
860
100.0
100.0
Hard
198
100.0
100.0
Easy + Medium
6,761
87.3 / 12.7
66.5 / 33.5 ( γ=0.85 )
Easy + Hard
6,099
96.8 / 3.2
84.5 / 15.5 ( γ=0.85 )
Medium + Hard
1,058
81.3 / 18.7
72.8 / 27.2 ( γ=0.55 )
Table 3 : Fine-tuning data composition and expected sampling ratios for the seven fine-tuning strategies
Figure 4 : Comparison of F1 scores across seven fine-tuning strategies. The nine nuclei instance-segmentation model configurations are grouped into (a) the Cellpose family, (b) the CellViT family, and (c) other segmentation models. For each model, performance is shown for the baseline and seven fine-tuning strategies: Easy (E) , Medium (M) , Hard (H) , Easy + Hard (E+H) , Easy + Medium (E+M) , Medium + Hard (M+H) and Easy + Medium + Hard (E+M+H) .
Cell Segmentation Models
Baseline Performance
Fine-tuned Performance
Fine-tuning Strategy
StarDist
0.7380
0.8332
Medium
Cellpose 3.0
0.6748
0.7378
Easy + Medium + Hard
Cellpose-SAM
0.7816
0.8657
Medium
CellViT (HIPT-256)
0.7838
0.8299
Medium
CellViT++ (Virchow)
0.7995
0.8338
Medium + Hard
CellViT (SAM-H)
0.8108
0.8423
Medium
Table 4 : Best fine-tuning strategy and corresponding F1 performance for each model.
Figure 5 : Qualitative comparison of baseline and fine-tuned models. Green contours denote nuclei instances. Dashed boxes highlight recovered false negatives (FN) and improved instance separation.
Research Center for Medical Image Analysis and Artificial Intelligence, Department of Medicine, Faculty of Medicine and Dentistry, Danube Private University, Krems an der Donau, Austria · Department of Pathology, Stanford University School of Medicine, USA · Center for Artificial Intelligence in Medicine & Imaging, Stanford University, USA