Prompt and Refinement: Asymmetric Mutual Learning for Infrared Small Target Detection with Noisy Labels
Authors: Yimin Fu, Songbo Wang, Lizhuo Liu, Baicheng Pan, Zhunga Liu, Michael K. Ng
Organizations: Department of Mathematics, Hong Kong Baptist University, Hong Kong, China · School of Automation, Northwestern Polytechnical University, Xi’an, 710072, China · School of Electrical and Control Engineering, Xi’an University of Science and Technology, Xi’an, 710054, China
Existing data-driven infrared small target detection (ISTD) methods typically require large-scale datasets with accurate pixel-level annotations for model training. However, such labor-intensive requirements are difficult to satisfy in real-world applications due to the heavy reliance on expert knowledge and the inherently weak distinctiveness of infrared small targets. Consequently, the presence of noisy labels during model training is inevitable, which can severely mislead the learning of target perception toward spurious patterns. To address this challenge, we propose Prompt and Refinement (PAR), a label-noise-robust asymmetric mutual learning paradigm for ISTD. Specifically, PAR comprises a pretrained Segment Anything Model (SAM) and an ISTD-specific detector trained from scratch, which learn collaboratively through a peer-teaching scheme. Coupled with local contrast regularity, the predictions of the two asymmetric peer models are mutually exploited as rectification cues for the supervisory masks of their counterparts. The interaction between complementary inductive biases effectively prevents the label correction process from degenerating into the self-confirmation loop of a single model, enabling progressive refinement of the annotations toward intrinsic target characteristics. In addition, the detector predictions are utilized as corrective mask prompts to facilitate task-specific adaptation of the vision foundation model. Moreover, an evidential uncertainty estimation strategy is introduced into the optimization process to further alleviate the adverse effects of noisy labels. Extensive experiments under diverse noisy label scenarios on three ISTD datasets demonstrate that PAR consistently achieves state-of-the-art performance.
Figures & tables
Fig. 1: The illustration of different types of noisy labels for infrared small targets. Pixels colored in red and green denote false-positive and false-negative annotations, respectively.
Fig. 2: The architecture overview of our proposed PAR for ISTD under noisy-label supervision.
Method
Venue
Dilation & Erosion
Shrinkage
Expansion
IoU ↑
F1↑
Pd↑
Fa↓
IoU ↑
F1↑
Pd↑
Fa↓
IoU ↑
F1↑
Pd↑
Fa↓
Tophat [ 6 ]
OE96
43.13
32.27
83.49
35.31
43.13
32.27
83.49
35.31
43.13
32.27
83.49
35.31
MPCM [ 7 ]
PR16
24.41
34.65
88.99
17.84
24.41
34.65
88.99
17.84
24.41
34.65
88.99
17.84
ACMNet [ 10 ]
WACV21
48.19
65.03
87.61
54.87
52.23
68.61
91.74
22.42
48.55
65.36
91.42
56.23
DNANet [ 11 ]
TIP22
58.52
73.83
95.66
61.31
61.96
76.51
94.60
12.78
63.74
77.85
94.17
31.91
UIUNet [ 12 ]
TIP22
54.01
70.14
88.78
18.38
68.18
81.08
96.61
16.04
60.08
75.06
94.49
25.07
TABLE I: Comparison of results under different noise types on the NUDT-SIRST dataset in terms of IoU (%), F1 (%), Pd (%), and Fa ( 10−6 ). The best results are highlighted in bold.
Method
Venue
Dilation & Erosion
Shrinkage
Expansion
IoU ↑
F1↑
Pd↑
Fa↓
IoU ↑
F1↑
Pd↑
Fa↓
IoU ↑
F1↑
Pd↑
Fa↓
Tophat [ 6 ]
OE96
39.64
36.39
78.41
167.20
39.64
36.39
78.41
167.20
39.64
36.39
78.41
167.20
MPCM [ 7 ]
PR16
12.80
16.62
66.24
207.90
12.80
16.62
66.24
207.90
12.80
16.62
66.24
207.90
ACMNet [ 10 ]
WACV21
57.41
72.94
93.57
9.05
59.51
74.62
91.74
23.24
54.91
70.89
94.49
117.98
DNANet [ 11 ]
TIP22
63.01
77.31
94.49
10.82
70.56
82.74
99.08
14.90
58.88
74.12
95.41
29.81
UIUNet [ 12 ]
TIP22
63.51
77.69
98.16
40.10
69.83
82.24
96.33
10.64
52.56
68.91
99.08
50.38
TABLE II: Comparison of results under different noise types on the NUAA-SIRST dataset in terms of IoU (%), F1 (%), Pd (%), and Fa ( 10−6 ). The best results are highlighted in bold.
Fig. 3: ROC curves of different methods on the NUAA-SIRST and NUDT-SIRST datasets under different label noise scenarios.
Method
Venue
MSHNet [ 39 ]
MTUNet [ 13 ]
L2SKNet [ 40 ]
IoU ↑
F1↑
Pd↑
Fa↓
IoU ↑
F1↑
Pd↑
Fa↓
IoU ↑
F1↑
Pd↑
Fa↓
baseline
—
45.39
62.43
84.65
29.86
43.10
60.23
76.47
32.39
43.72
60.84
66.45
29.46
+ URN [ 28 ]
AAAI22
46.88
63.84
71.31
23.26
44.21
61.31
69.51
22.91
42.14
59.29
58.54
19.06
+ RMD [ 60 ]
TMI23
41.67
58.83
76.46
17.45
46.72
63.69
79.74
29.83
40.69
57.84
72.24
16.12
+ AIO2 [ 70 ]
TGRS24
43.24
60.37
62.28
18.75
44.11
61.22
56.11
26.86
40.67
57.82
55.48
25.51
+ DSR [ 59 ]
TNNLS25
45.48
62.52
69.52
17.88
47.75
64.64
72.46
15.17
49.69
66.39
71.63
14.85
TABLE III: Comparison of results under different noise types on the NUDT-Sea dataset in terms of IoU (%), F1 (%), Pd (%), and Fa ( 10−6 ). The best results are highlighted in bold.
Fig. 4: Visualizations of IRSTD results of different methods. Boxes in red, blue, and yellow refer to detected targets, missed detections, and false alarms. Zoomed-in regions of the detected targets are provided in the corner for clear visualization.
AML
JAR
EUE
IoU ↑
F1↑
Pd↑
Fa↓
✘
✘
✘
53.26
69.50
90.79
24.61
✘
✘
✔
57.20
72.77
94.17
20.11
✘
✔
✘
65.51
79.16
93.75
17.28
✔
✘
✘
62.52
76.94
95.56
12.11
✔
✘
✔
66.69
80.01
95.66
15.85
✔
✔
✘
67.71
80.74
95.23
11.35
TABLE IV: Ablation experiments about contributions of each component to noisy ISTD. Best results are highlighted in bold.
Fig. 5: Ablation experiments for different configurations of architectural and objective asymmetry.
Strategy
Shrinkage ←θ→ Expansion
0.3
0.4
0.5
0.6
0.7
Baseline
58.18
66.73
68.47
60.94
55.28
Entropy
57.87
66.84
67.60
62.58
55.08
Variance
58.34
65.60
67.70
64.27
55.20
Sensitivity
59.93
67.68
69.89
65.76
56.76
Evidence
60.98
68.52
70.90
67.59
58.01
TABLE V: Ablation experiments for different uncertainty estimation strategies. Best results are highlighted in bold.
Method
Params (M)
Clean labels
Noisy labels
IoU ↑
Pd↑
IoU ↑
Pd↑
MTUNet
12.75
61.43
84.54
43.10
76.47
+ RMD
25.50
57.80
72.25
46.72
79.74
+ RoCoT
25.50
55.32
77.01
46.02
62.39
+ PAR
22.05
61.11
84.71
56.55
81.43
TABLE VI: Analysis of computational complexity. Best results are highlighted in bold.
Fig. 6: The performance of PAR with different combinations of τp and τc .
Single-frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel-wise annotations are expensive. Point supervision with online label evolution reduces annotation cost; however, lightweight CNN detectors often lack sufficient semantics, leading to noisy pseudo-masks and unstable optimization. To address this, we propose a hierarchical VFM-driven knowledge distillation framework that uses a frozen Vision Foundation Model (VFM) during training. We formulate point-supervised learning as a bilevel optimization process: the inner loop adapts a VFM-embedded teacher on reweighted training samples, while the outer loop transfers validation-guided knowledge to a lightweight student to mitigate pseudo-label noise and training-set bias. We further introduce Semantic-Conditioned Affine Modulation (SCAM) to inject VFM semantics into CNN features at multiple layers. In addition, a dynamic collaborative learning strategy with cluster-level sample reweighting enhances robustness to imperfect pseudo-masks. Experiments on diverse challenging cases across multiple ISTD backbones demonstrate consistent improvements in detection accuracy and training stability. Our code is available at https://github.com/yuanhang-yao/semantic-prior.
Yuanhang Yao, Ping Qian, Zhu Liu +2
School of Software Technology, Dalian University of Technology, China
Infrared small target detection (IRSTD) in high-resolution images is crucial for unmanned aerial vehicle (UAV) surveillance and UAV-based ground monitoring. However, small target size, weak features, and interference from complex dynamic backgrounds make IRSTD challenging. Existing methods incur redundant computation in non-target background regions and insufficiently exploit target context, limiting detection performance. To address these issues, we propose ECFNet, an efficient coarse-to-fine IRSTD framework with attention prior-guided knowledge distillation. In the coarse stage, we design a region binary classification network (RBCN) on grid-based multi-scale feature maps to efficiently identify target-containing context region proposals. A new denoising-assisted training strategy incorporates noisy ground-truth (GT) masks into RBCN feature maps and trains the network to reconstruct the original GT masks. This auxiliary task encourages explicit learning of target-background context to better distinguish target proposals from background regions. In the fine stage, we customize a lightweight target detector to the coarse-stage region proposals to balance accuracy and efficiency. Furthermore, we introduce a knowledge distillation strategy guided by a teacher-student cross-attention prior. This strategy directs the student to focus on critical target regions, enhancing discriminative feature representations for infrared small targets. Extensive experiments on three real infrared datasets demonstrate that ECFNet outperforms existing single-stage and two-stage approaches while maintaining high real-time processing efficiency. Code: https://github.com/IVPLabs/ECFNet.
Houzhang Fang, Ruixuan Huang, Qiuhuan Chen +3
Xidian University, Xi’an, China · Huazhong University of Science and Technology, Wuhan, China
Point supervision has become a scalable solution to address dense annotation for infrared small target detection, but its performance is limited by two coupled bottlenecks: unstable pseudo-label evolution in cluttered, low-contrast infrared imagery and severe sample-distribution imbalance. In this paper, we present a more adaptive and stable framework to address these issues. Leveraging the intrinsic consistency between thermal radiation patterns and heat diffusion, we propose a physics-induced annotation strategy that expands single-point labels into reliable pseudo-masks. To further enhance supervision and alleviate sample imbalance, we develop a bi-level dual-update framework that jointly optimizes detector weights, sample weights, and diffusion parameters. A meta-classifier dynamically predicts sample-wise loss weights, while a differentiable diffusion module refines pseudo-labels with detection feedback, enabling adaptive interaction between training and hyperparameter optimization. Extensive experiments across multiple datasets demonstrate five-fold annotation acceleration, superior detection accuracy, and comparable performance with 30% of the training data, validating the efficiency and practicality of our approach. Our code is available at https://github.com/yuanhang-yao/diffuse-to-detect.
Zhu Liu, Yuanhang Yao, Ping Qian +2
School of Software Technology, Dalian University of Technology, Dalian, China.