Organizations: Zhengzhou Police University, Zhengzhou, China · School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing, China · Qilu University of Technology and Jinan Supercomputing Center, Jinan, China
Contour initialization and numerical stopping can jointly affect the evaluation of active-contour segmentation. We examine their interaction using the open-source scikit-image Chan-Vese implementation on a resized ISIC 2017 mirror. A fixed development set of 100 images selects a common input channel; all 600 images in the repository's held-out partition are then evaluated. Otsu thresholding is compared with checkerboard-, disk-, and Otsu-initialized contours under default and tighter level-set tolerances. At the default tolerance, Otsu initialization increases mean image Dice from 0.6011 to 0.6660 relative to checkerboard initialization, a paired difference of 0.0649 (95% image-bootstrap interval [0.0452, 0.0860]). Otsu thresholding alone achieves 0.6897. The default disk initializer stops after one iteration on 471 images. Tightening the tolerance reduces the Otsu-seed advantage over checkerboard initialization to 0.0197, with most runs reaching the 500-iteration limit. The default-tolerance advantage also reverses between small- and large-lesion strata. These findings show that an improvement over a generic initializer can coexist with deterioration relative to the threshold baseline. Evaluations should retain the unrefined mask as a comparator and report the initial-field definition, stopping tolerance, and observed iteration counts together.
Figures & tables
Configuration
ϵ
Dice
IoU
BF 2
Low Dice
Iter.
CPU (ms)
Otsu only
–
0.6897
0.6011
0.3079
135
0
3.8
CV: checkerboard
10−3
0.6011
0.5049
0.1873
210
54
127.2
CV: disk
10−3
0.3933
0.3113
0.0843
385
1
14.4
CV: Otsu seed
10−3
0.6660
0.5654
0.2024
137
199
359.8
CV: checkerboard
10−5
0.6404
0.5491
0.2399
181
500
908.4
CV: disk
10−5
0.4069
0.3314
0.1062
366
500
919.6
Table 1: Results for all 600 held-out image IDs. BF 2 : boundary F1 within two pixels. Low Dice: count with Dice <0.5 . Iteration count and CPU time are medians; timing excludes loading and scoring.
Figure 1: Overlap and computation under the two stopping tolerances. Dashed lines in (a,b) mark the Otsu-only baseline; error bars in (a) show 95% image-bootstrap intervals. The dotted line in (c) marks the 500-iteration limit.
Figure 2: Cumulative distributions of per-image Dice differences. Positive values favor Otsu-seeded over checkerboard-initialized Chan–Vese.
Configuration
<10%
10–30%
>30%
Images
225
182
193
Otsu
0.583
0.747
0.760
Checkerboard D
0.299
0.757
0.806
Disk D
0.080
0.363
0.787
Otsu seed D
0.534
0.733
0.756
Checkerboard T
0.402
0.772
0.795
Table 2: Mean image Dice by reference foreground fraction. D: ϵ=10−3 ; T: ϵ=10−5 .
Figure 3: Examples nearest the 10th, 50th, and 90th percentiles of the default Otsu-seed minus checkerboard Dice difference, shown from top to bottom. Green contours: reference; magenta contours: prediction. Values below panels are per-image Dice.
The annotation used to select a segmentation threshold is part of the evaluation protocol, yet its effect is easily conflated with model quality. We examine this choice for retinal vessel segmentation using all 28 CHASE DB1 images and both human annotations. A fixed seven-fold protocol keeps both eyes of each of the 14 subjects together. Random forests and Extra Trees are fitted against observer 1 with three random seeds, yielding 42 fits. Five threshold policies share identical score maps: fixed 0.50, observer-1 tuning, observer-2 tuning, mean-observer tuning, and maximin tuning of the per-image lower observer Dice. For random forests, maximin changes the threshold in 19 of 21 fits, but worst-observer Dice decreases from 70.53 percent to 70.45 percent. The paired difference is -0.073 percentage points, with a conditional subject-bootstrap 95 percent interval of [-0.384, 0.238]. Extra Trees shows the same direction. Identical observer-1-tuned random-forest masks score 73.66 percent against observer 1 and 71.06 percent against observer 2. The results support explicit reporting of both the threshold-selection reference and evaluation reference; they do not support an accuracy benefit from maximin tuning in this cohort. All splits, raw predictions, metrics and code are supplied. AI assistance is disclosed.
Wenhao Xu, Yixian Kong, Ting Pan +2
Zhengzhou Police University, Zhengzhou, China · School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing, China · Qilu University of Technology and Jinan Supercomputing Center, Jinan, China
The segmentation of multiple degradations has been a challenging problem in the field of image segmentation. Existing level set approaches commonly adopt a length regularization term to constrain the geometric shape of the segmentation contour. However, the introduction of the length term often results in numerical instability and high computational cost. In this paper, we show that the length term is not essential under certain smoothness constraints, and theoretically prove that the presence of the length term affects the property of ∣∇φ∣=1. Based on the finding, we define a class of smooth images, construct the grayscale level set, and propose a fast segmentation framework for degraded images, such as heavily noisy images and intensity inhomogeneous images. The framework transforms PDE evolution into one-dimensional threshold search, which has significant advantages in computational speed, especially on large-scale images. Experiments validate the segmentation performance of the proposed framework on various degraded images.
Xingkai Li, Jiebao Sun, Fanghui Song +1
School of Mathematics, Harbin Institute of Technology, Harbin, 150001, Heilongjiang, China.
Multilevel image thresholding is widely used for segmentation in applications ranging from medical imaging to remote sensing. Classical objective functions, such as Otsu's between-class variance and Kapur's entropy, are often optimized using metaheuristic algorithms, with performance evaluated via metrics like Structural Similarity Index (SSIM) and Peak Signal-to-Noise Ratio (PSNR). These evaluations implicitly assume that SSIM and PSNR provide unbiased measures of segmentation quality. In this study, we examine this assumption by analyzing the correlation between thresholding objective functions and quality metrics across all possible thresholds for images in the BSDS500 dataset. Results show that Otsu's criterion consistently exhibits high correlation with both SSIM and PSNR, while Kapur's entropy demonstrates weaker and more variable correlation. Otsu outperforms Kapur in correlation with PSNR for all images and with SSIM for over 91%. Our findings reveal an inherent metric-objective-function bias. This work highlights the need for more neutral evaluation frameworks and motivates extending the analysis to additional thresholding criteria and domains. Source code of this paper can be found at https://w3id.org/met-dp/icpr26-95