When Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-Labeling
Organizations: Graduate School of Information, Production and Systems, Waseda University, Kitakyushu, Japan · School of Software, Dalian University of Technology, Dalian, China · School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China
Abstract
Pseudo-labeling has become a cornerstone of learning from unlabeled data in semantic segmentation. Yet its effectiveness drops sharply in real-world scenarios where strong imaging noise and long-tailed class distributions occur together. We trace this failure to a vicious cycle of pseudo-label degradation. Imaging noise entangles foreground and background features, lowering prediction confidence across all classes, while long-tailed distributions leave tail classes with far fewer training samples and inherently lower confidence. Under fixed high-threshold filtering, these tail-class predictions are systematically filtered out, so they receive no supervision from unlabeled data and thus features keep degrading in subsequent iterations. Critically, noise and long-tail are not independent obstacles but mutually amplifying ones, and addressing either alone is insufficient. To break this cycle, we propose FTC-Seg, a Feature-Threshold dual-Calibration framework built on a standard teacher-student framework. At the feature level, Orthogonal Prototype Reconstruction (OPR) uses a set of learnable orthogonal prototypes to residually purify pixel-wise features, widening the margin between weak foreground targets and noisy backgrounds. At the threshold level, Adaptive Threshold Calibration (ATC) dynamically adjusts class-specific thresholds based on learning difficulty and prediction-distribution bias, rescuing low-confidence pseudo-labels of tail classes from systematic exclusion. Extensive experiments on four public benchmarks spanning three distinct noise modalities show that FTC-Seg achieves strong performance against state-of-the-art methods, with particularly substantial gains on tail classes. Our results establish that jointly calibrating features and thresholds is essential for robust pseudo-labeling under compounded noise and class imbalance.
Figures & tables
| Method | Venue | Encoder | FLSMD | FSSG | ||||
| 2% | 5% | 10% | 2% | 5% | 10% | |||
| Labeled Only | – | RN-101 | 46.47 1.17 | 55.15 0.47 | 61.94 0.83 | 17.64 1.35 | 32.23 0.66 | 52.27 1.01 |
| Labeled Only | – | DINOv2-S | 51.08 0.50 | 61.02 0.95 | 68.48 0.70 | 35.61 0.89 | 41.67 1.42 | 54.58 0.32 |
| AEL [ 10 ] | NIPS’21 | RN-101 | 52.70 1.54 | 64.89 1.24 | 70.73 0.88 | 53.51 1.02 | 57.84 0.98 | 63.08 0.46 |
| Dual Teacher [ 15 ] | NIPS’23 | MiT-B2 | 56.23 0.70 | 66.09 0.56 | 69.61 0.71 | 58.31 0.80 | 60.16 0.93 | 65.91 1.01 |
| UniMatch [ 25 ] | CVPR’23 | RN-101 | 50.70 1.47 | 56.14 0.68 | 61.66 1.34 | 58.39 0.88 | 60.87 0.70 | 63.15 1.27 |
| Method | Encoder | SUIM | ACDC | ||||
| 2% | 5% | 10% | 2% | 5% | 10% | ||
| Labeled Only | RN-101 | 35.64 1.44 | 45.47 1.52 | 52.68 0.97 | 36.62 1.13 | 44.43 0.91 | 48.47 1.27 |
| Labeled Only | DINOv2-S | 56.61 0.80 | 62.56 1.29 | 64.26 1.15 | 52.60 0.45 | 56.57 0.62 | 66.98 0.97 |
| UniMatch | RN-101 | 47.86 2.17 | 49.17 1.98 | 55.01 1.71 | 45.09 2.36 | 50.81 1.49 | 51.66 1.20 |
| SemiVL | RN-101 | 55.99 0.79 | 58.43 0.60 | 60.15 1.32 | 45.13 0.57 | 50.65 0.71 | 52.97 0.97 |
| FTC-Seg (Ours) | RN-101 | 54.71 0.53 | 59.05 1.46 | 62.85 0.97 | 45.56 1.35 | 51.80 0.89 | 55.73 1.04 |
| Method | Component | Pseudo-Label Acc. (%) | Global | |||
| Res. | Global | Tail | mIoU | |||
| Baseline | – | – | – | 67.03 | 56.56 | 58.75 |
| w/o Residual | ✗ | ✓ | ✓ | 68.69 | 89.61 | 56.93 |
| w/o | ✓ | ✗ | ✓ | 71.30 | 81.04 | 59.26 |
| w/o | ✓ | ✓ | ✗ | 64.32 | 83.16 | 59.84 |
| OPR | ✓ | ✓ | ✓ | 72.33 | 87.80 | 62.55 |
| Method | Pseudo-Label Acc (%) | Filtered-Rate (%) | Correctly Kept GT Rate (%) | ||||||
| Global | Head | Tail | Global | Head | Tail | Global | Head | Tail | |
| Baseline | 67.03 | 73.06 | 56.56 | 13.15 | 9.50 | 17.83 | 58.22 | 67.04 | 46.47 |
| + OPR | 72.33 | 74.08 | 87.80 | 9.04 | 7.04 | 11.24 | 63.98 | 71.34 | 77.93 |
| + ATC | 69.73 | 79.60 | 88.65 | 11.63 | 10.38 | 10.39 | 64.80 | 68.13 | 79.44 |
| FTC-Seg (ours) | 73.83 | 84.13 | 89.03 | 8.26 | 6.75 | 8.43 | 67.15 | 78.21 | 80.61 |
| OPR Layers | mIoU | Params (M) | FLOPs (G) | Infer. Time (ms) | Peak Mem. (MB) |
| 0 (Baseline) | 58.75 | 24.79 | 56.88 | 12.97 | 14279 |
| 1 | 61.92 | 24.80 | 56.90 | 13.15 | 14348 |
| 2 | 62.55 | 24.81 | 56.92 | 13.27 | 14398 |
| 3 | 61.47 | 24.82 | 56.93 | 13.43 | 14449 |
| 4 | 60.98 | 24.84 | 56.95 | 13.58 | 14497 |
| Prototypes 1–4 | Prototypes 5–8 | Prototypes 9–12 | Prototypes 13–16 |
| : SP SP TI | : CC CC DV | : CC UR CC | : SP SP CG |
| : CC DV BL | : SF SF UR | : SF SF PO | : PO PO CC |
| : SQ UR DV | : PO PO DV | : DV DV PO | : SF TI CC |
| : CC BL CC | : PO PO SP | : SQ SQ PO | : TI TI PO |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Labels | M3 | M4 | M5 | Holm | |
| FLSMD | 2% | 59.80 | 60.17 | 63.07 | +2.90 | 0.0436 |
| 5% | 66.72 | 66.49 | 68.84 | +2.12 | 0.0225 | |
| 10% | 71.24 | 70.99 | 73.49 | +2.25 | 0.0356 | |
| FSSG | 2% | 60.87 | 61.24 | 63.20 | +1.96 | 0.0448 |
| 5% | 62.73 | 62.46 | 66.31 | +3.58 | 0.0160 | |
| 10% | 63.97 | 64.16 | 68.01 | +3.85 | 0.0086 |
| Method | mIoU | Tail mIoU | Pseudo-label mIoU |
| EMA baseline | 58.78 | 5.33 | 53.49 |
| Original OPR | 62.55 | 17.19 | 73.86 |
| Class-anchored OPR | 54.87 | 8.93 | 58.25 |
| Setting | mIoU | Diver IoU | ||
| Correct | 62.55 | – | 17.50 | |
| Add background | 59.80 | 10.11 | ||
| Omit Diver | 58.42 | 11.49 | ||
| Add background + omit Diver | 53.09 | 2.47 |
| Dataset / corruption | Mild: parameter / | Medium: parameter / | Severe: parameter / |
| FLSMD / speckle | / | / | / |
| SUIM / scattering | / | / | / |
| ACDC / fog | / | / | / |
| Method | Long-tail only | Long-tail + noise | Noise-induced change |
| Baseline | 70.76 | 63.18 | |
| FTC-Seg | 75.36 | 69.11 | |
| FTC-Seg gain | – |
| Method | Params (M) | FLOPs (G) | FPS | mIoU (%) |
| Method comparison | ||||
| Labeled only | 24.79 | 56.94 | 122.62 | 35.61 |
| UniMatch V2 | 24.79 | 56.94 | 122.62 | 58.78 |
| FTC-Seg | 24.81 | 56.96 | 120.93 | 63.20 |
| Component ablation | ||||
| M2: EMA | 24.79 | 56.88 | 122.62 | 58.75 |