Organizations: Department of Electrical and Computer Engineering, Vanderbilt University, Nashville, TN 37235, USA · Department of Radiology, Weill Cornell Medicine, New York, NY 10021, USA · Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37235, USA · Department of Pathology, Microbiology and Immunology, Vanderbilt University Medical Center, Nashville, TN 37235, USA · Department of Computer Science, Vanderbilt University, Nashville, TN 37235, USA
Accurate nuclei instance segmentation is essential for quantitative renal pathology, yet general-purpose models often struggle with low contrast, dense nuclei, complex morphology, and strong background staining. In this work, we extended a human-in-the-loop framework by combining 5,901 foundation-model-generated pseudo-labels from well-segmented cases (Easy), 860 newly expert-annotated unresolved challenging cases (Medium), and 198 expert-annotated consensus failure cases (Hard). These annotations, spanning different levels of segmentation difficulty, enabled the systematic evaluation of seven single-source and mixed-source fine-tuning strategies across nine cell segmentation model configurations. Fine-tuning improved all models, with Medium data included in seven of the nine best-performing strategies. LSP-DETR achieved the highest F1 score of 0.8725 with Hard-only fine-tuning, while StarDist showed the largest improvement, increasing from 0.7380 to 0.8332 with Medium-only fine-tuning. These findings show that annotations spanning multiple difficulty levels support effective model adaptation, although the optimal annotation composition remains model dependent.
Figures & tables
Figure 1 : Overall framework. (A) Expert annotation of previously unresolved Medium-quality kidney pathology patches to enrich the existing Easy and Hard fine-tuning data. (B) Expanded evaluation of cell nuclei segmentation models released from 2022 to 2026, including both previously evaluated and newly incorporated models. (C) Fine-tuning with seven single-source and mixed-source combinations of Easy, Medium, and Hard datasets to evaluate model-specific adaptation performance.
Figure 2 : Representative examples of prediction-rating categories. Foundation-model predictions were categorized as Good, Medium, or Bad according to the proportion of identifiable nuclei captured in each image patch. Green contours denote nucleus instances; the top row shows model predictions, whereas the bottom row shows the corresponding ground-truth annotations.
Figure 3 : Human-in-the-loop annotation and construction of multi-source fine-tuning data. (A) Pathology experts screened unresolved Medium-quality kidney patches, selected challenging samples, and manually annotated nucleus instances from the original images. (B) The final fine-tuning data consisted of three annotation sources: prior foundation-model-generated pseudo-labels from Easy cases, prior expert-annotated consensus failure cases (Hard), and newly human-in-the-loop annotated unresolved cases (Medium).
Dataset
Count
Annotation Source
Easy
5,901
Foundation-model-generated pseudo-labels
Medium
860
Unresolved Medium-quality patches annotated by experts
Hard
198
Prior foundation-model failure patches annotated by experts
Table 2 : Summary of the annotation sources used for model fine-tuning.
Fine-tuning Strategies
Samples
Original Ratio (%)
Expected Ratio (%)
Easy
5,901
100.0
100.0
Medium
860
100.0
100.0
Hard
198
100.0
100.0
Easy + Medium
6,761
87.3 / 12.7
66.5 / 33.5 ( γ=0.85 )
Easy + Hard
6,099
96.8 / 3.2
84.5 / 15.5 ( γ=0.85 )
Medium + Hard
1,058
81.3 / 18.7
72.8 / 27.2 ( γ=0.55 )
Table 3 : Fine-tuning data composition and expected sampling ratios for the seven fine-tuning strategies
Figure 4 : Comparison of F1 scores across seven fine-tuning strategies. The nine nuclei instance-segmentation model configurations are grouped into (a) the Cellpose family, (b) the CellViT family, and (c) other segmentation models. For each model, performance is shown for the baseline and seven fine-tuning strategies: Easy (E) , Medium (M) , Hard (H) , Easy + Hard (E+H) , Easy + Medium (E+M) , Medium + Hard (M+H) and Easy + Medium + Hard (E+M+H) .
Cell Segmentation Models
Baseline Performance
Fine-tuned Performance
Fine-tuning Strategy
StarDist
0.7380
0.8332
Medium
Cellpose 3.0
0.6748
0.7378
Easy + Medium + Hard
Cellpose-SAM
0.7816
0.8657
Medium
CellViT (HIPT-256)
0.7838
0.8299
Medium
CellViT++ (Virchow)
0.7995
0.8338
Medium + Hard
CellViT (SAM-H)
0.8108
0.8423
Medium
Table 4 : Best fine-tuning strategy and corresponding F1 performance for each model.
Figure 5 : Qualitative comparison of baseline and fine-tuned models. Green contours denote nuclei instances. Dashed boxes highlight recovered false negatives (FN) and improved instance separation.
Background and Objective: Precise and scalable instance segmentation of cell nuclei is a fundamental prerequisite for computational pathology, yet gigapixel whole-slide images (WSIs) pose significant computational challenges. While patch-based processing is standard during training, existing methods are often limited to small tile sizes during inference due to architectural bottlenecks or reliance on computationally expensive post-processing for instance separation. We introduce a faster, scalable, and end-to-end framework capable of processing large-scale image tiles while accurately modeling biologically realistic overlapping nuclei. Methods: We propose LSP-DETR (Local Star Polygon DEtection TRansformer). The model represents nuclei as star-convex polygons and employs a lightweight transformer with linear complexity, enabling the processing of high-resolution images in a single forward pass. A novel radial distance loss accommodates annotation uncertainty, allowing the segmentation of overlapping nuclei to emerge naturally without explicit overlap labels. Results: LSP-DETR achieves state-of-the-art efficiency, with an inference time of 0.45 s/mm^2, a 3.2x speedup over StarDist, the next-fastest method. On PanNuke, the model achieves competitive accuracy (67.5 bPQ), while yielding an F1-score of 0.964 in polygon overlap when evaluated against consensus annotations from two expert pathologists. Furthermore, it outperforms larger models such as LKCell in generalization robustness, reaching an F1-score of 85.0 on MoNuSeg. Conclusions: LSP-DETR bridges the gap between high-fidelity segmentation and practical clinical requirements by eliminating heuristic post-processing. By providing a scalable, linear-complexity solution that naturally handles overlaps between nuclei, this framework sets a new direction for efficient high-throughput WSI analysis in digital pathology.
Matěj Pekár, Vít Musil, Rudolf Nenutil +2
Masaryk University, Faculty of Informatics, Botanická 68a, Brno, 602 00, Czech Republic · Masaryk Memorial Cancer Institute, Žlutý kopec 7, Brno, 656 53, Czech Republic · BBMRI-ERIC, Neue Stiftingtalstraße 2/B/6, Graz, 8010, Austria
In computational pathology, nuclear instance segmentation is a fundamental task with many downstream clinical applications. With the advent of deep learning, many approaches, including convolutional neural networks (CNNs) and vision transformers (ViTs), have been proposed for this task, along with both machine learning-based and non-machine learning-based pre- and post-processing techniques to further boost performance. However, one fundamental aspect that has received less attention is the evaluation pipeline. In this study, we identify four key issues associated with nuclear instance segmentation evaluation and propose corresponding solutions. Our proposed modifications, namely handling vague regions, score normalization, overlapping instances, and border uncertainty, are integrated into a unified framework called NucEval, which enables robust evaluation of nuclear instance segmentation. We evaluate this pipeline using the NuInsSeg dataset, which provides unique characteristics that make it particularly suitable for this study, as well as two additional external datasets, with three CNN- and ViT-based nuclear instance segmentation models, to demonstrate the impact of these modifications on instance segmentation metrics. The code, along with complete guidelines and illustrative examples, is publicly available at: https://github.com/masih4/nuc_eval.
Amirreza Mahbod, Ramona Woitek, Jeanne Shen
Research Center for Medical Image Analysis and Artificial Intelligence, Department of Medicine, Faculty of Medicine and Dentistry, Danube Private University, Krems an der Donau, Austria · Department of Pathology, Stanford University School of Medicine, USA · Center for Artificial Intelligence in Medicine & Imaging, Stanford University, USA
Accurate classification of nuclei subtypes in histopathology images is critical for downstream tasks including tumor grading, immune infiltrate quantification, and prognosis prediction. Existing approaches rely on either convolutional or transformer-based encoders in isolation, limiting their ability to simultaneously capture fine-grained local texture and long-range spatial context. We present AMN (Adaptive Multi-Scale Nuclei Network), a dual-encoder segmentation framework that jointly leverages a Swin Transformer and a ResNet-50 feature pyramid, fused via a learned per-channel gating mechanism that dynamically weighs each encoder's contribution at every scale. AMN is trained with a multi-objective loss combining class-weighted focal loss, boundary-aware loss with positive-pixel emphasis, and a novel uncertainty-modulated classification term that suppresses overconfident erroneous predictions. Evaluated on the CoNIC benchmark across seven nuclei classes, AMN achieves a mean Dice of 0.82 and mean F1 of 0.68, with an F1 of 0.67 on the diagnostically challenging lymphocyte class. AMN outperforms eight baseline models spanning pure-CNN, pure-transformer, and recent hybrid architectures: U-Net, ResU-Net, DeepLabV3+, SegNet, ViT-Small, HmsU-Net, ConvFormer-UNet, and BEFUnet. Cross-dataset evaluation on MoNuSeg demonstrates strong generalization without retraining and validating the domain robustness of the learned representations.
Spoorthi M, Suja Palaniswamy
Department of Computer Science & Engineering, Amrita School of Computing, Bengaluru, Amrita Vishwa Vidyapeetham, India