On the Relaxation of Conditional Independence Assumption for Image Segmentation
Organizations: Department of Statistics and Data Science The Chinese University of Hong Kong
Abstract
In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Assumption (CIA), which ignores crucial label correlations and therefore degrades performance in ambiguous or low-contrast scenarios. However, accounting for full label dependence is computationally prohibitive, requiring time. To address this, we replace the CIA with a Spatially Localized Dependence (SLD) structure that captures local label correlations while keeping the dependence model tractable. We further overcome the remaining computational bottleneck via a Reciprocal Moment Approximation coupled with a novel fixed-point optimization strategy that eliminates exhaustive search. The proposed algorithm achieves a highly practical complexity and consistently outperforms conventional argmax and CIA-based RankSEG across diverse segmentation benchmarks. Improvements are significant in low-contrast or small-object scenarios, where label dependence offers valuable signals complementary to image information for accurate segmentation. The code of experiments is available at https://github.com/ZixunWang/RankSEG-DEP.
Figures & tables
| Model | Prediction | KiTS | LiTS | ||
| IoU | Dice | IoU | Dice | ||
| UNet | Argmax-prob | 51.00 | 57.36 | 38.45 | 47.58 |
| CIA-RankSEG | 53.54 | 60.07 | 40.70 | 50.07 | |
| Ours | 53.92 | 60.48 | 40.77 | 50.16 | |
| DeepLabV3+ | Argmax-prob | 54.19 | 61.16 | 38.34 | 47.38 |
| CIA-RankSEG | 56.22 | 63.56 | 40.09 | 49.50 | |
| Model | Prediction | ADE20K | Cityscapes | DeepGlobe | |||
| mIoU | mDice | mIoU | mDice | mIoU | mDice | ||
| PSPNet | Argmax-prob | 51.32 | 58.66 | 73.07 | 80.45 | 57.24 | 66.23 |
| CIA-RankSEG | 51.57 | 59.17 | 73.72 | 81.14 | 58.20 | 67.82 | |
| Ours | 51.92 | 59.49 | 73.84 | 81.27 | 58.47 | 68.14 | |
| DeepLabV3+ | Argmax-prob | 52.53 | 59.57 | 73.37 | 80.59 | 57.75 | 66.73 |
| CIA-RankSEG | 52.64 | 59.95 | 73.92 | 81.24 | 58.38 | 67.94 | |
| Quantile | 0.2 | 0.4 | 0.6 |
|---|---|---|---|
| Argmax-prob | +14.37 % | +7.46 % | +3.97 % |
| CIA-RankSEG | +2.51 % | +1.04 % | +0.48 % |
| Group | small | medium | large |
|---|---|---|---|
| Argmax-prob | +5.47 % | +2.92 % | +2.01 % |
| CIA-RankSEG | +1.19 % | +0.82 % | +0.42 % |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | LiTS | KiTS | ADE20K | Cityscapes | DeepGlobe |
|---|---|---|---|---|---|
| #Steps | 63.36% | 88.52% | 9.30% | 2.80% | 40.49% |
| #Steps | 36.42% | 11.41% | 90.55% | 97.20% | 50.51% |
| #Steps | 0.22% | 0.07% | 0.15% | 0.00% | 0.00% |
| Dataset | KiTS | LiTS | ||
|---|---|---|---|---|
| Metric | IoU | Dice | IoU | Dice |
| Fixed-point Optimization | 56.70 | 64.07 | 40.48 | 49.88 |
| Brute-force Search | 56.73 | 64.10 | 40.48 | 49.88 |
| Method | CRF | Argmax-prob | CIA-RankSEG | Ours | ||||
|---|---|---|---|---|---|---|---|---|
| mIoU | mDice | mIoU | mDice | mIoU | mDice | mIoU | mDice | |
| DeepGlobe | 58.61 | 67.40 | 59.49 | 68.60 | 60.19 | 69.57 | 60.52 | 69.91 |
| Cityscapes | 73.26 | 80.35 | 75.66 | 82.61 | 76.17 | 83.21 | 76.25 | 83.31 |
| ADE20K | 56.05 | 62.61 | 56.94 | 63.98 | 57.67 | 64.92 | 57.94 | 65.19 |