Organizations: School of Mechatronic Engineering, Xi’an Technological University, Xi’an, 710021, Shaanxi, China · School of Electronic Information Engineering, Xi’an Technological University, Xi’an, 710021, Shaanxi, China · Shaanxi Shanhua Coal Chemical Co., Ltd., Weinan, 714104, Shaanxi, China
Infrared gas leak detection is important for industrial safety and environmental monitoring, but automatic detection remains challenging because gas plumes are often faint, small, semi-transparent, and weakly bounded. This study proposes an Edge-Aware and Content-Adaptive Feature Fusion Detector (ECAF-Det) for infrared gas leak detection in weak-plume and cluttered thermal scenes. The main methodological contributions of ECAF-Det comprise three task-oriented components. A local--global feature enhancement block preserves fine boundary cues and long-range plume continuity. A multi-scale edge perception module transforms directional-gradient and Gabor-response cues into hierarchical boundary-sensitive structural priors. A content-adaptive sparse routing path aggregation network dynamically regulates multi-scale feature propagation and limits the contribution of less informative cross-scale responses. Experiments on the IIG dataset show that ECAF-Det improves overall and small-plume detection while maintaining moderate computational complexity. On this dataset, ECAF-Det achieves an average precision (AP) of 29.8%, an AP at an IoU threshold of 0.5 AP50 of 84.3%, and a small-object AP of 25.3%. Compared with the Real-Time Detection Transformer with a ResNet-18 backbone (RT-DETR-R18), these values represent improvements of 3.0, 6.5, and 5.4 percentage points, respectively. The model requires 43.7 giga floating-point operations (GFLOPs) and 14.3 M parameters. On the LangGas dataset, ECAF-Det achieves an AP of 36.3% and an AP50 of 68.5%. The AI contribution lies in edge-aware representation learning and content-adaptive sparse feature routing for weak infrared plume perception. The engineering application is automated infrared gas leak detection for industrial safety monitoring, early warning, and remote inspection.
Figures & tables
Figure 1: Performance and computational-complexity comparison of ECAF-Det and representative detectors on the IIG dataset. Panels (a)–(c) show AP, AP 50 , and AP S , respectively. In each panel, the line represents detection performance, while the bars indicate GFLOPs and parameter counts. Numerical values are annotated for all compared methods to facilitate direct performance–complexity comparison.
Figure 2: Overall architecture of the proposed Edge-Aware and Content-Adaptive Feature Fusion Detector (ECAF-Det) for infrared gas leak detection. The framework consists of a multi-scale edge perception module (MSEPM), a plume-oriented local–global backbone, a content-adaptive sparse routing path aggregation network (CASR-PAN), and a decoder-based detection head. MSEPM provides hierarchical boundary-sensitive structural priors to support plume representation, while CASR-PAN uses an importance estimator to generate content-dependent routing weights for adaptive multi-scale feature aggregation. In AIMM-F and AIMM-S, β denotes a fixed positive offset in the variation-dependent modulation term, while γ denotes the fixed identity-preservation coefficient.
Figure 3: Structure of the plume-oriented local–global backbone. MSEPM generates hierarchical boundary-sensitive structural priors, which are fused with the corresponding backbone features through scale-specific feature-fusion operations. The local–global plume blocks then combine a local branch for preserving fine-grained plume structures with a frequency-domain global branch for capturing broader contextual continuity.
Figure 4: Architectures of the proposed structural-prior generation and content-adaptive routing components. (a) Multi-scale edge perception module (MSEPM) and gradient–Gabor edge operator (GGEO). GGEO combines four fixed 3×3 Sobel-like directional-gradient kernels with a depthwise 3×3 Gabor-response branch to generate an initial structural representation E0 . MSEPM then constructs hierarchical structural priors through progressive 2×2 max-pooling and 1×1 projection. (b) Importance estimator (IE) in CASR-PAN. The estimator integrates global importance, local importance, and feature-diversity cues to generate content-dependent importance maps for adaptive multi-scale feature routing.
Model
AP (%)
AP 50 (%)
AP 75 (%)
AP S (%)
AP M (%)
AP L (%)
GFLOPs (G)
Params (M)
YOLOv10n ( Wang et al., 2024a )
27.0
70.7
10.6
17.6
34.6
38.4
6.5
2.265
YOLOv10m
25.2
65.8
11.2
14.8
34.2
23.6
58.9
15.313
YOLOv11n ( Khanam and Hussain, 2024 )
24.7
72.3
6.7
16.7
31.1
42.8
6.3
2.582
YOLOv11m
25.8
70.4
10.4
14.7
34.5
31.5
67.6
20.030
YOLOv12n ( Tian et al., 2025 )
25.7
75.2
7.2
19.0
30.9
29.2
6.3
2.555
YOLOv12m
25.2
70.7
9.8
18.2
33.3
12.7
67.1
20.105
Table 1: Detection performance comparison of ECAF-Det and representative detectors on the IIG dataset.
Model
AP (%)
AP 50 (%)
AP 75 (%)
AP S (%)
AP M (%)
AP L (%)
GFLOPs (G)
Params (M)
YOLOv10n ( Wang et al., 2024a )
32.3
61.7
30.2
0.1
11.4
39.8
6.5
2.265
YOLOv10m
28.8
56.9
26.8
0.8
14.1
34.9
58.9
15.313
YOLOv11n ( Khanam and Hussain, 2024 )
33.5
65.1
30.1
1.7
14.7
40.5
6.3
2.582
YOLOv11m
32.6
65.1
28.1
1.6
14.5
39.8
67.6
20.030
YOLOv12n ( Tian et al., 2025 )
34.4
65.7
31.0
1.1
15.3
41.3
6.3
2.555
YOLOv12m
32.8
63.6
28.7
1.7
13.8
39.7
67.1
20.105
Table 2: Detection performance comparison of ECAF-Det and representative detectors on the LangGas dataset.
Model
Precision (%)
Recall (%)
F1-score (%)
Missed detection rate (%)
Model inference latency (ms/image)
End-to-end FPS
AP 50 (%)
RT-DETR-R18
86.97
70.60
77.94
29.40
12.53
75.60
77.8
ECAF-Det
86.31
79.44
82.73
20.56
15.46
61.81
84.3
Table 3: Detection-oriented performance comparison between RT-DETR-R18 and ECAF-Det on the IIG test set. Model-inference latency and end-to-end throughput were measured on an NVIDIA GeForce RTX 5070 Ti GPU with a batch size of 4. End-to-end FPS includes preprocessing, model inference, and postprocessing.
Metric
Value
Matched pairs
1062
Pred./GT area ratio, median (IQR)
0.912 (0.788–1.136)
Pred./GT width ratio, median (IQR)
0.960 (0.891–1.065)
Pred./GT height ratio, median (IQR)
0.960 (0.860–1.089)
Normalized center displacement, median (IQR)
0.080 (0.051–0.116)
Area ratio >1.25 (%)
17.23
Table 4: Bounding-box geometry statistics for ECAF-Det on the IIG test set. Statistics are computed for one-to-one matched detections with IoU ≥0.5 .
Figure 5: Qualitative comparison of detection results on the IIG dataset. All examples are original thermal infrared samples from the released IIG dataset. Each row shows the results of a different detector. ECAF-Det detects more weak and small plume regions in several challenging cases, but extremely faint and distant plumes remain difficult to localize accurately.
Condition
Model
AP (%)
AP 50 (%)
AP 75 (%)
AR 100 (%)
Plume visibility
Weak
RT-DETR-R18
24.8
75.4
4.5
51.2
ECAF-Det
29.5
88.0
4.9
53.8
Moderate
RT-DETR-R18
26.9
79.1
9.3
53.9
ECAF-Det
32.1
87.2
11.9
56.5
Strong
RT-DETR-R18
28.4
74.5
12.1
49.8
Table 5: Challenge-oriented detection performance on the IIG test set.
Local–Global Plume Block
MSEPM
AP
AP 50
AP S
AP M
AP L
GFLOPs (G)
Params (M)
×
×
26.8
77.8
19.9
32.0
37.1
56.9
19.873
✓
×
28.5
80.7
22.2
32.8
37.3
51.9
15.965
×
✓
27.9
80.4
24.0
30.6
38.1
59.8
21.119
✓
✓
29.5
81.0
24.3
33.5
38.9
54.8
17.211
Table 6: Ablation study of the plume-oriented local–global backbone on the IIG dataset. The local–global plume block denotes the proposed block for local detail preservation and frequency-domain contextual aggregation. MSEPM denotes the multi-scale edge perception module using GGEO to generate hierarchical boundary-sensitive structural priors.
Plume-Oriented Local–Global Backbone
CASR-PAN
AP
AP 50
AP S
AP M
AP L
GFLOPs (G)
Params (M)
×
×
26.8
77.8
19.9
32.0
37.1
56.9
19.873
✓
×
29.5
81.0
24.3
33.5
38.9
54.8
17.211
×
✓
29.4
79.1
23.0
33.6
42.2
45.8
16.939
✓
✓
29.8
84.3
25.3
32.5
42.6
43.7
14.277
Table 7: Ablation study of the complete ECAF-Det on the IIG dataset.
Configuration
AP
AP 50
AP S
AP M
AP L
GFLOPs
Params
All BasicBlocks baseline
26.8
77.8
19.9
32.0
37.1
56.9
19.873
P2–P5 local–global plume blocks
27.4
77.9
21.1
31.1
35.7
47.1
15.725
P4–P5 local–global plume blocks
28.5
80.7
22.2
32.8
37.3
51.9
15.965
P5 local–global plume block
27.0
78.6
22.4
30.2
42.8
54.5
16.743
Table 8: Stage-wise placement of local–global plume blocks in the ResNet-18 backbone on the IIG dataset.
Figure 6: Effective receptive field (ERF) visualization comparison between (a) the proposed plume-oriented local–global backbone and (b) the original ResNet-18 backbone under the same training and inference settings.
Figure 7: Comparison of high-contribution area ratios between the proposed plume-oriented local–global backbone and the original ResNet-18 backbone under different cumulative contribution thresholds.
Figure 8: Performance comparison of different backbone architectures on the IIG dataset. The marked point denotes the proposed plume-oriented local–global backbone.
Figure 9: Visualization results of the proposed gradient–Gabor edge operator (GGEO). (a) Original infrared image; (b) directional-gradient response extracted by the fixed Sobel-like operators; (c) Gabor-response cue obtained from the depthwise Gabor branch; (d) fused structural response combining directional-gradient and Gabor cues.
Figure 10: Visualization of the hierarchical structural representations involved in the multi-scale edge perception module (MSEPM). (a) Original infrared image; (b) E0 : the initial structural representation generated by GGEO before downsampling; (c) E1 : the representation after one 2×2 max-pooling operation; (d) E2 : the representation after two successive max-pooling operations; and (e) E3 : the representation after three successive max-pooling operations. The downsampled representations E1 , E2 , and E3 are subsequently projected through scale-specific 1×1 convolutions to form the hierarchical structural priors used by the corresponding backbone stages.
Gradient Directions
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs (G)
Params (M)
0∘
26.7
79.8
6.3
21.4
31.1
33.8
59.7
21.117
0∘ , 90∘
28.8
81.1
9.4
22.1
33.1
42.3
59.7
21.118
0∘ , 45∘ , 90∘ , 135∘
27.9
80.4
10.3
24.0
30.6
38.1
59.8
21.119
Table 9: Ablation study of gradient directions in GGEO on the IIG dataset.
Figure 11: Effect of the fusion coefficient α in GGEO. Performance comparisons under different fixed values of α and learnable α are reported in terms of (a) AP, (b) AP 50 , and (c) AP S .
Figure 12: Training evolution of the learnable fusion coefficient α in GGEO over 200 epochs. After an initial increasing phase, α fluctuates within a narrow range, indicating a stable fusion preference during training.
Operator
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs
Params
Sobel ( Heath et al., 1998 )
27.9
79.7
8.1
21.5
32.9
33.6
59.4
21.109
Canny ( Canny, 1986 )
25.8
80.0
7.2
22.1
27.7
31.2
60.1
21.110
Laplacian ( Wang, 2007 )
26.7
79.4
8.2
21.9
30.1
29.7
59.8
21.110
GGEO
27.9
80.4
10.3
24.0
30.6
38.1
59.8
21.119
Table 10: Comparison of edge operators and controlled edge-prior replacements within MSEPM on the IIG dataset.
Model
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs (G)
Params (M)
Routing-path ablation
Naive additive fusion
25.9
75.9
8.5
23.6
28.0
38.3
45.8
16.939
Without deep-to-mid fusion
27.5
79.5
7.6
22.5
31.1
36.8
45.8
16.939
Without deep-to-shallow fusion
25.7
76.2
6.8
19.2
30.5
40.6
45.8
16.939
Without shallow-to-mid fusion
25.7
75.8
6.2
19.8
29.7
40.2
45.8
16.939
Without mid-level self-fusion
28.3
80.3
8.3
22.3
31.9
43.5
45.8
16.939
Table 11: Ablation study of CASR-PAN components on the IIG dataset. The first group evaluates the adaptive routing paths, while the second group examines the importance-estimator branches and branch-weighting strategy.
Analysis
β
γ
AP
AP 50
AP 75
AP S
AP M
AP L
β sensitivity ( γ=1.0 )
β variation
0.00
1.00
27.5
78.4
8.6
21.6
31.1
42.6
β variation
0.25
1.00
28.2
77.8
10.3
21.9
32.2
35.8
Default
0.50
1.00
29.4
79.1
12.3
23.0
33.6
42.2
β variation
0.75
1.00
25.7
76.5
6.8
21.3
28.5
39.0
β variation
1.00
1.00
25.9
76.9
7.4
20.7
29.8
36.3
Table 12: Sensitivity analysis of the fixed modulation coefficients β and γ in CASR-PAN on the IIG dataset. One coefficient is varied while the other is fixed at its default value.
Model
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs (G)
Params (M)
PANet ( Liu et al., 2018 )
27.3
80.9
8.0
22.1
30.0
39.2
103.8
21.463
BiFPN ( Tan et al., 2020 )
27.0
76.6
8.3
20.8
31.3
38.9
64.3
20.300
NAS-FPN ( Ghiasi et al., 2019 )
25.1
76.4
7.5
19.3
29.9
26.1
93.8
19.756
CASR-PAN
29.4
79.1
12.3
23.0
33.6
42.2
45.8
16.939
Table 13: Comparison of different neck structures on the IIG dataset.
Global-context branch
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs (G)
Params (M)
31×31 large-kernel convolution
25.3
74.5
7.3
17.8
31.2
31.6
55.1
18.095
Non-local module
26.6
77.9
8.6
20.3
31.1
33.9
53.6
17.274
Swin shifted-window attention
26.8
76.9
9.6
22.2
29.6
39.3
72.8
31.743
Proposed DCT-based frequency branch
28.5
80.7
8.8
22.2
32.8
37.3
51.9
15.965
Table 14: Comparison of alternative global-context modeling strategies within the local–global plume block on the IIG dataset.
Abbr./Symbol
Meaning
Abbr./Symbol
Meaning
RT-DETR
Real-Time Detection Transformer
ECAF-Det
Edge-Aware and Content-Adaptive Feature Fusion Detector
Single-point supervised infrared small target detection (IRSTD) drastically reduces dense annotation costs. Current state-of-the-art (SOTA) methods achieve high precision by recovering mask supervision through explicit, offline pseudo-label construction, such as multi-stage active learning and physics-driven mask generation. In this paper, we study a minimalist alternative: generating point-to-mask supervision online through in-batch, point-anchored feature-affinity propagation. We instantiate this paradigm as GSACP, an end-to-end testbed that directly supervises the detector using hard-margin feature affinity gated by local image priors, entirely eliminating external label-evolution loops. This compact design, however, exposes an optimization bottleneck. Because the affinity target is generated from the same feature representation being optimized, training forms a self-referential loop. We theoretically formalize this as \emph{Self-Referential Propagation Drift}, a representation-supervision entanglement that can sharpen true boundaries or distort the feature space to satisfy its own targets. To systematically isolate these failure modes, we apply a protocolized single-variable ablation procedure spanning local EMA teacher decoupling, hard-background contrastive separation, and adaptive support geometry. On the SIRST3 dataset, GSACP-Final establishes a new ultra-low false-alarm operating regime, achieving a highly competitive 0.6674 mIoU while demonstrating a 38%relativereductioninfalse−positiveartifacts(\mathrm{Fa}$) compared with PAL. By systematically deconstructing the end-to-end paradigm, we map its performance boundaries and show that in-batch feature propagation provides a compact alternative for deployment scenarios where false-alarm suppression is paramount.
Robust object detection under adverse visual conditions remains a long-standing challenge for multi-modal perception systems. Existing fusion-based methods typically require both RGB and infrared (IR) inputs, and treat them equally during both training and inference, which compromises their robustness when the RGB modality becomes unreliable or unavailable. In this case, we propose \textbf{InfraNet}, an IR-centric quality-aware framework that regulates RGB guidance during training and supports flexible RGB--IR or IR-only deployment. InfraNet employs an asymmetric architecture where the primary IR pathway extracts multi-scale infrared features for predictions, while the auxiliary RGB pathway provides reliability-controlled supervisory signals. The core of InfraNet is \textbf{QualGate}, a quality-aware fusion module that learns a task-oriented control signal to suppress unreliable RGB guidance and compensate IR features during cross-modal training. Built upon InfraNet, we design two architectural variants: a lightweight IR-only architecture InfraNet-IR and an RGB--IR architecture InfraNet-RGB-IR. Our method is evaluated through extensive experiments on four benchmark datasets (LLVIP, FLIR-Aligned, M3FD, and DroneVehicle), showing strong or competitive accuracy in challenging low-light and adverse weather conditions. Notably, InfraNet maintains high efficiency in IR-only inference, making it both accurate and computationally efficient.
Zichao Feng, Haodong Zhu, Jingying Yang +8
Beihang University, Beijing, China · Zhongguancun Academy, Beijing, China · Communication University of China, Beijing, China +1
Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak targets in infrared imagery usually occupy only a few pixels and are easily submerged by cloud clutter, ground edges, and bright noise, making it difficult for lightweight segmentation-based methods to preserve local target structures while suppressing background interference. To address these challenges, we propose LCMamNet, a lightweight cross-scale Mamba network that progressively enhances local target structures, interacts cross-scale context in a latent space, and restores spatial details with background suppression. Specifically, a compact hierarchical encoder with cross-shaped directional bottleneck residual (CDBR) blocks strengthens direction-sensitive target structures under a small computation budget. A latent dense cross-scale fusion (LDCF) module then performs dense all-level interaction through bidirectional Mamba modeling and reorganizes the interacted features into stable hierarchical semantics. Finally, a progressive decoder selectively recovers shallow spatial details while suppressing irrelevant background textures. Extensive experiments on IRSTD-1k, NUAA-SIRST, and NUDT-SIRST show that the proposed network achieves mIoU scores of 71.25%, 79.60%, and 95.58%, respectively, with only 1.175M parameters and 6.91 GFLOPs. It also runs with a mean inference latency of 6.62 ms, and deployment results on an NVIDIA Jetson Orin NX 16G SUPER further demonstrate its practical potential for real-time edge inference. The code and checkpoints are publicly available at https://github.com/Haoyu096/LCMamNet.
Yuhao Fan, Le Hui, Yuchao Dai
School of Electronics and Information, Northwestern Polytechnical University, Xi’an 710072, China