Hybrid Approach for Enhancing Lesion Segmentation in Fundus Images
Organizations: Department of Electrical & Software Engineering, University of Calgary, Canada. · Department of Computer Science and Physics, Wilfrid Laurier University, Waterloo, Canada. · Dept of Ophthalmology & Visual Sciences, University of Alberta, Edmonton, Canada. · Wills Eye Hospital, Philadelphia, USA. · Department of Surgery, University of Calgary, Canada.
Abstract
Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical to improving survival rates, but misdiagnosis or delayed diagnosis can lead to poor outcomes. Despite advancements in AI-based image analysis, diagnosing choroidal nevi in colour fundus images remains challenging, particularly for clinicians without specialized expertise. Existing datasets often suffer from low resolution and inconsistent labelling, limiting the effectiveness of segmentation models. This paper addresses the challenge of achieving precise segmentation of fundus lesions, a critical step toward developing robust diagnostic tools. While deep learning models like U-Net have demonstrated effectiveness, their accuracy heavily depends on the quality and quantity of annotated data. Previous mathematical/clustering segmentation methods, though accurate, required extensive human input, making them impractical for medical applications. This paper proposes a novel approach that combines mathematical/clustering segmentation models with insights from U-Net, leveraging the strengths of both methods. This hybrid model improves accuracy, reduces the need for large-scale training data, and achieves significant performance gains on high-resolution fundus images. The proposed model achieves a Dice coefficient of 89.7% and an IoU of 80.01% on 1024*1024 fundus images, outperforming the Attention U-Net model, which achieved 51.3% and 34.2%, respectively. It also demonstrated better generalizability on external datasets. This work forms a part of a broader effort to develop a decision support system for choroidal nevus diagnosis, with potential applications in automated lesion annotation to enhance the speed and accuracy of diagnosis and monitoring.
Figures & tables
| Ref# | Method | Segmented Part | Dataset |
| ( Eshragh et al., 2024 ) | Ensemble UNet | CN | Private |
| ( Biglarbeiki et al., 2024a ) | Application (YOLO) | CN | Private |
| ( Jiang et al., 2020 ) | Multi-Path Recurrent UNet | Vessels & Optic | DRIVE ( Niemeijer et al., 2009 ) |
| ( Zhao et al., 2021 ) | UNet & Transfer Learning | Disk & Cup | Drishti-GS ( Sivaswamy et al., 2014 ) |
| Component | Setting | How chosen |
| Guide network input | RGB; scaled to | Design (Sec. 7.5) |
| Guide output threshold | 0.5 on sigmoid map | Fixed |
| Data split | 90/10 train/test; 80/20 train/validation; seed 101112 | Fixed |
| (Eq. 13) | 5 | Largest training lesion |
| Calibration factor (Eq. 14) | 0.85 | Validation grid search |
| Default for empty mask | 700 | Smallest training lesion |
| Model | IOU | Dice | Hausdorff (px) | Specificity | Volumetric Similarity |
| Attention UNet | 47.05% | 60.57% | 77.91 | 98.77% | 72.37% |
| Swin UNet (Transformer-based) | 62.17% | 72.97% | 104.30 | 99.42% | 88.90% |
| nnU-Net | 73.04% | 83.23% | 49.22 | 98.89% | 90.03% |
| Hybrid Model | 87.89% | 90.93% | 21.6 | 99.92% | 95.32% |
| Model | Train Time (min) | Energy (Wh) |
| Attention UNet | 134.2 | 671.0 |
| Swin UNet (Transformer-based) | 160.7 | 803.5 |
| nnU-Net | 102.0 | 510.0 |
| Hybrid Model (one training run of the 128 128 guide network) | 3.16 | 15.8 |
| Model | Mean Dice (%) | 95% CI (bootstrap) | Median (%) | SD | Images with Dice 80% |
| Attention UNet | 60.57 | [50.63, 69.22] | 65.89 | 24.07 | 6/25 |
| Swin UNet | 72.97 | [63.06, 81.60] | 85.14 | 24.60 | 14/25 |
| nnU-Net | 83.23 | [77.73, 87.72] | 86.96 | 13.20 | 20/25 |
| Hybrid Model | 90.93 | [81.27, 97.68] | 98.96 | 21.88 | 21/25 |
| Hybrid Model vs. | Mean Dice [95% CI] | Median Dice | Wins / losses | Wilcoxon ; (Holm) | Permutation | Effect size ; |
| nnU-Net | +7.70 [0.93, 14.06] | +8.66 | 21 / 4 | 51; | 0.69; 0.45 | |
| Swin UNet | +17.96 [8.89, 27.63] | +10.60 | 21 / 4 | 41; | 0.75; 0.74 | |
| Att UNet | +30.36 [20.30, 40.42] | +31.45 | 23 / 2 | 18; | 0.89; 1.16 |
| Prediction | Measurement Method | Results |
| Attention UNet | IOU | 42.41% |
| Dice | 54.07% | |
| Swin UNet | IOU | 47.27% |
| Dice | 59.28% | |
| nnU-Net | IOU | 59.49% |
| Dice | 69.31% |
| Model | Mean Dice (%) | 95% CI (bootstrap) | Median (%) | SD | Images with Dice 80% |
| Attention UNet | 54.07 | [48.19, 59.88] | 57.87 | 28.84 | 22/90 |
| Swin UNet | 59.28 | [53.56, 64.93] | 63.19 | 27.54 | 26/90 |
| nnU-Net | 69.31 | [63.14, 75.06] | 81.46 | 29.07 | 46/90 |
| Hybrid Model | 76.79 | [69.94, 82.99] | 91.95 | 31.63 | 64/90 |
| Hybrid Model vs. | Mean Dice [95% CI] | Median Dice | Wins / losses / ties | Wilcoxon ; (Holm) | Permutation | Effect size ; |
| nnU-Net | +7.47 [2.81, 12.16] | +3.26 | 72 / 15 / 3 | 725; | 0.62; 0.32 | |
| Swin UNet | +17.51 [11.92, 22.95] | +14.83 | 73 / 13 / 4 | 508; | 0.73; 0.66 | |
| Att UNet | +22.72 [16.35, 29.16] | +18.67 | 75 / 12 / 3 | 467; | 0.76; 0.73 |
| Covariate | Internal test ( ) | External ( ) | Mann–Whitney | SMD |
| Brightness ( mean) | 31.43 4.96 | 32.09 5.25 | ||
| Contrast ( SD) | 8.36 1.35 | 8.15 2.05 | ||
| Colour cast ( mean) | 26.69 7.13 | 27.03 6.80 | ||
| Colour cast ( mean) | 27.28 6.47 | 27.92 5.91 | ||
| Sharpness (Laplacian variance) | 1.18 0.48 | 1.27 0.48 | ||
| Lesion area (% of FOV) | 1.64 2.03 | 4.97 8.52 |
| Covariate | Failures (Dice 0.5, ) | Non-failures ( ) | Mann–Whitney | Spearman with Dice |
| Brightness ( mean) | 31.47 | 31.86 | ( ) | |
| Contrast ( SD) | 7.72 | 7.76 | ( ) | |
| Colour cast ( mean) | 26.28 | 27.97 | ( ) | |
| Colour cast ( mean) | 26.69 | 27.66 | ( ) | |
| Sharpness (Laplacian variance) | 1.13 | 1.18 | ( ) | |
| Lesion area (% of FOV) | 8.36 | 2.00 | ( ) |
| Variant | Configuration | Mean Dice (%) [95% CI] | Median (%) | Full model better on | Wilcoxon (Holm) |
| A | Full model: automated + probability selection | 90.93 [81.27, 97.68] | 98.96 | – | – |
| B | Fixed + probability selection | 69.12 [57.84, 79.86] | 85.94 | 21 / 25 | |
| B | Fixed + probability selection | 45.70 [34.99, 56.80] | 39.42 | 23 / 25 | |
| B | Fixed + probability selection | 27.66 [20.25, 35.68] | 22.21 | 24 / 25 | |
| B ∗ | Oracle: best fixed per image | 80.37 | 90.34 | – | – |
| C | Automated + centroid selection | 90.93 [81.27, 97.68] | 98.96 | 0 / 25 (identical) | – |
| Model | Stage | 1024 1024 (s per image) | 3900 3900 (s per image) |
| Hybrid Model | Guide network: read + resize to 128 128 | 0.07 0.01 | 0.07 0.01 |
| Guide network: forward pass (CPU) | 0.49 0.03 | 0.49 0.03 | |
| Resize guide mask, compute (Eqs. 11–14) | 0.01 | 0.19 0.01 | |
| SLIC (Eq. 15) | 1.00 0.06 | 14.16 1.12 | |
| Superpixel selection (Eqs. 17–18) | 0.01 | 0.14 0.01 | |
| Morphological closing | 0.01 | 0.02 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Test image | Attention UNet | Swin UNet | nnU-Net | Hybrid Model |
| 0 | 74.81 | 55.31 | 85.86 | 99.67 |
| 1 | 73.12 | 84.08 | 82.56 | 79.36 |
| 2 | 58.23 | 92.19 | 91.62 | 98.16 |
| 3 | 19.18 | 46.58 | 79.75 | 98.96 |
| 4 | 46.89 | 92.38 | 79.19 | 78.39 |
| 5 | 87.81 | 74.51 | 80.05 | 49.80 |
| Img | Att | Swin | nnU | Hyb | Img | Att | Swin | nnU | Hyb | Img | Att | Swin | nnU | Hyb |
| 0 | 49.1 | 20.2 | 51.6 | 0.0 | 30 | 54.3 | 49.0 | 48.5 | 50.8 | 60 | 62.0 | 74.5 | 93.9 | 97.6 |
| 1 | 63.4 | 96.5 | 75.2 | 98.7 | 31 | 95.5 | 88.4 | 94.9 | 94.9 | 61 | 5.1 | 51.2 | 55.2 | 87.5 |
| 2 | 5.4 | 1.2 | 1.2 | 21.2 | 32 | 61.3 | 49.7 | 75.6 | 88.1 | 62 | 52.3 | 60.2 | 80.5 | 96.8 |
| 3 | 62.7 | 62.1 | 95.6 | 97.9 | 33 | 89.1 | 94.7 | 95.3 | 97.8 | 63 | 42.7 | 77.1 | 84.0 | 73.6 |
| 4 | 76.6 | 94.1 | 86.9 | 84.5 | 34 | 93.4 | 95.3 | 96.8 | 97.1 | 64 | 82.5 | 61.7 | 84.5 | 90.6 |
| 5 | 45.2 | 59.4 | 91.3 | 97.0 | 35 | 68.8 | 91.6 | 94.5 | 97.0 | 65 | 12.5 | 35.3 | 63.5 | 91.0 |