cs.CVOct 5, 2026

Prompt and Refinement: Asymmetric Mutual Learning for Infrared Small Target Detection with Noisy Labels

Authors: Yimin Fu, Songbo Wang, Lizhuo Liu, Baicheng Pan, Zhunga Liu, Michael K. Ng

Organizations: Department of Mathematics, Hong Kong Baptist University, Hong Kong, China · School of Automation, Northwestern Polytechnical University, Xi’an, 710072, China · School of Electrical and Control Engineering, Xi’an University of Science and Technology, Xi’an, 710054, China

Abstract

Existing data-driven infrared small target detection (ISTD) methods typically require large-scale datasets with accurate pixel-level annotations for model training. However, such labor-intensive requirements are difficult to satisfy in real-world applications due to the heavy reliance on expert knowledge and the inherently weak distinctiveness of infrared small targets. Consequently, the presence of noisy labels during model training is inevitable, which can severely mislead the learning of target perception toward spurious patterns. To address this challenge, we propose Prompt and Refinement (PAR), a label-noise-robust asymmetric mutual learning paradigm for ISTD. Specifically, PAR comprises a pretrained Segment Anything Model (SAM) and an ISTD-specific detector trained from scratch, which learn collaboratively through a peer-teaching scheme. Coupled with local contrast regularity, the predictions of the two asymmetric peer models are mutually exploited as rectification cues for the supervisory masks of their counterparts. The interaction between complementary inductive biases effectively prevents the label correction process from degenerating into the self-confirmation loop of a single model, enabling progressive refinement of the annotations toward intrinsic target characteristics. In addition, the detector predictions are utilized as corrective mask prompts to facilitate task-specific adaptation of the vision foundation model. Moreover, an evidential uncertainty estimation strategy is introduced into the optimization process to further alleviate the adverse effects of noisy labels. Extensive experiments under diverse noisy label scenarios on three ISTD datasets demonstrate that PAR consistently achieves state-of-the-art performance.

Figures & tables

Explore similar work

May 14, 2026cs.CV

Learning with Semantic Priors: Stabilizing Point-Supervised Infrared Small Target Detection via Hierarchical Knowledge Distillation

Single-frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel-wise annotations are expensive. Point supervision with online label evolution reduces annotation cost; however, lightweight CNN detectors often lack sufficient semantics, leading to noisy pseudo-masks and unstable optimization. To address this, we propose a hierarchical VFM-driven knowledge distillation framework that uses a frozen Vision Foundation Model (VFM) during training. We formulate point-supervised learning as a bilevel optimization process: the inner loop adapts a VFM-embedded teacher on reweighted training samples, while the outer loop transfers validation-guided knowledge to a lightweight student to mitigate pseudo-label noise and training-set bias. We further introduce Semantic-Conditioned Affine Modulation (SCAM) to inject VFM semantics into CNN features at multiple layers. In addition, a dynamic collaborative learning strategy with cluster-level sample reweighting enhances robustness to imperfect pseudo-masks. Experiments on diverse challenging cases across multiple ISTD backbones demonstrate consistent improvements in detection accuracy and training stability. Our code is available at https://github.com/yuanhang-yao/semantic-prior.
Jun 20, 2026cs.CV

Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation

Infrared small target detection (IRSTD) in high-resolution images is crucial for unmanned aerial vehicle (UAV) surveillance and UAV-based ground monitoring. However, small target size, weak features, and interference from complex dynamic backgrounds make IRSTD challenging. Existing methods incur redundant computation in non-target background regions and insufficiently exploit target context, limiting detection performance. To address these issues, we propose ECFNet, an efficient coarse-to-fine IRSTD framework with attention prior-guided knowledge distillation. In the coarse stage, we design a region binary classification network (RBCN) on grid-based multi-scale feature maps to efficiently identify target-containing context region proposals. A new denoising-assisted training strategy incorporates noisy ground-truth (GT) masks into RBCN feature maps and trains the network to reconstruct the original GT masks. This auxiliary task encourages explicit learning of target-background context to better distinguish target proposals from background regions. In the fine stage, we customize a lightweight target detector to the coarse-stage region proposals to balance accuracy and efficiency. Furthermore, we introduce a knowledge distillation strategy guided by a teacher-student cross-attention prior. This strategy directs the student to focus on critical target regions, enhancing discriminative feature representations for infrared small targets. Extensive experiments on three real infrared datasets demonstrate that ECFNet outperforms existing single-stage and two-stage approaches while maintaining high real-time processing efficiency. Code: https://github.com/IVPLabs/ECFNet.
May 20, 2026cs.CV

Diffuse to Detect: Bi-Level Sample Rebalancing with Pseudo-Label Diffusion for Point-Supervised Infrared Small-Target Detection

Point supervision has become a scalable solution to address dense annotation for infrared small target detection, but its performance is limited by two coupled bottlenecks: unstable pseudo-label evolution in cluttered, low-contrast infrared imagery and severe sample-distribution imbalance. In this paper, we present a more adaptive and stable framework to address these issues. Leveraging the intrinsic consistency between thermal radiation patterns and heat diffusion, we propose a physics-induced annotation strategy that expands single-point labels into reliable pseudo-masks. To further enhance supervision and alleviate sample imbalance, we develop a bi-level dual-update framework that jointly optimizes detector weights, sample weights, and diffusion parameters. A meta-classifier dynamically predicts sample-wise loss weights, while a differentiable diffusion module refines pseudo-labels with detection feedback, enabling adaptive interaction between training and hyperparameter optimization. Extensive experiments across multiple datasets demonstrate five-fold annotation acceleration, superior detection accuracy, and comparable performance with 30% of the training data, validating the efficiency and practicality of our approach. Our code is available at https://github.com/yuanhang-yao/diffuse-to-detect.