Organizations: School of Mechatronic Engineering, Xi’an Technological University, Xi’an, 710021, Shaanxi, China · School of Electronic Information Engineering, Xi’an Technological University, Xi’an, 710021, Shaanxi, China · Shaanxi Shanhua Coal Chemical Co., Ltd., Weinan, 714104, Shaanxi, China
Infrared gas leak detection is important for industrial safety and environmental monitoring, but automatic detection remains challenging because gas plumes are often faint, small, semi-transparent, and weakly bounded. This study proposes an Edge-Aware and Content-Adaptive Feature Fusion Detector (ECAF-Det) for infrared gas leak detection in weak-plume and cluttered thermal scenes. The main methodological contributions of ECAF-Det comprise three task-oriented components. A local--global feature enhancement block preserves fine boundary cues and long-range plume continuity. A multi-scale edge perception module transforms directional-gradient and Gabor-response cues into hierarchical boundary-sensitive structural priors. A content-adaptive sparse routing path aggregation network dynamically regulates multi-scale feature propagation and limits the contribution of less informative cross-scale responses. Experiments on the IIG dataset show that ECAF-Det improves overall and small-plume detection while maintaining moderate computational complexity. On this dataset, ECAF-Det achieves an average precision (AP) of 29.8%, an AP at an IoU threshold of 0.5 AP50 of 84.3%, and a small-object AP of 25.3%. Compared with the Real-Time Detection Transformer with a ResNet-18 backbone (RT-DETR-R18), these values represent improvements of 3.0, 6.5, and 5.4 percentage points, respectively. The model requires 43.7 giga floating-point operations (GFLOPs) and 14.3 M parameters. On the LangGas dataset, ECAF-Det achieves an AP of 36.3% and an AP50 of 68.5%. The AI contribution lies in edge-aware representation learning and content-adaptive sparse feature routing for weak infrared plume perception. The engineering application is automated infrared gas leak detection for industrial safety monitoring, early warning, and remote inspection.
Figures & tables
Figure 1: Performance and computational-complexity comparison of ECAF-Det and representative detectors on the IIG dataset. Panels (a)–(c) show AP, AP 50 , and AP S , respectively. In each panel, the line represents detection performance, while the bars indicate GFLOPs and parameter counts. Numerical values are annotated for all compared methods to facilitate direct performance–complexity comparison.
Figure 2: Overall architecture of the proposed Edge-Aware and Content-Adaptive Feature Fusion Detector (ECAF-Det) for infrared gas leak detection. The framework consists of a multi-scale edge perception module (MSEPM), a plume-oriented local–global backbone, a content-adaptive sparse routing path aggregation network (CASR-PAN), and a decoder-based detection head. MSEPM provides hierarchical boundary-sensitive structural priors to support plume representation, while CASR-PAN uses an importance estimator to generate content-dependent routing weights for adaptive multi-scale feature aggregation. In AIMM-F and AIMM-S, β denotes a fixed positive offset in the variation-dependent modulation term, while γ denotes the fixed identity-preservation coefficient.
Figure 3: Structure of the plume-oriented local–global backbone. MSEPM generates hierarchical boundary-sensitive structural priors, which are fused with the corresponding backbone features through scale-specific feature-fusion operations. The local–global plume blocks then combine a local branch for preserving fine-grained plume structures with a frequency-domain global branch for capturing broader contextual continuity.
Figure 4: Architectures of the proposed structural-prior generation and content-adaptive routing components. (a) Multi-scale edge perception module (MSEPM) and gradient–Gabor edge operator (GGEO). GGEO combines four fixed 3×3 Sobel-like directional-gradient kernels with a depthwise 3×3 Gabor-response branch to generate an initial structural representation E0 . MSEPM then constructs hierarchical structural priors through progressive 2×2 max-pooling and 1×1 projection. (b) Importance estimator (IE) in CASR-PAN. The estimator integrates global importance, local importance, and feature-diversity cues to generate content-dependent importance maps for adaptive multi-scale feature routing.
Model
AP (%)
AP 50 (%)
AP 75 (%)
AP S (%)
AP M (%)
AP L (%)
GFLOPs (G)
Params (M)
YOLOv10n ( Wang et al., 2024a )
27.0
70.7
10.6
17.6
34.6
38.4
6.5
2.265
YOLOv10m
25.2
65.8
11.2
14.8
34.2
23.6
58.9
15.313
YOLOv11n ( Khanam and Hussain, 2024 )
24.7
72.3
6.7
16.7
31.1
42.8
6.3
2.582
YOLOv11m
25.8
70.4
10.4
14.7
34.5
31.5
67.6
20.030
YOLOv12n ( Tian et al., 2025 )
25.7
75.2
7.2
19.0
30.9
29.2
6.3
2.555
YOLOv12m
25.2
70.7
9.8
18.2
33.3
12.7
67.1
20.105
Table 1: Detection performance comparison of ECAF-Det and representative detectors on the IIG dataset.
Model
AP (%)
AP 50 (%)
AP 75 (%)
AP S (%)
AP M (%)
AP L (%)
GFLOPs (G)
Params (M)
YOLOv10n ( Wang et al., 2024a )
32.3
61.7
30.2
0.1
11.4
39.8
6.5
2.265
YOLOv10m
28.8
56.9
26.8
0.8
14.1
34.9
58.9
15.313
YOLOv11n ( Khanam and Hussain, 2024 )
33.5
65.1
30.1
1.7
14.7
40.5
6.3
2.582
YOLOv11m
32.6
65.1
28.1
1.6
14.5
39.8
67.6
20.030
YOLOv12n ( Tian et al., 2025 )
34.4
65.7
31.0
1.1
15.3
41.3
6.3
2.555
YOLOv12m
32.8
63.6
28.7
1.7
13.8
39.7
67.1
20.105
Table 2: Detection performance comparison of ECAF-Det and representative detectors on the LangGas dataset.
Model
Precision (%)
Recall (%)
F1-score (%)
Missed detection rate (%)
Model inference latency (ms/image)
End-to-end FPS
AP 50 (%)
RT-DETR-R18
86.97
70.60
77.94
29.40
12.53
75.60
77.8
ECAF-Det
86.31
79.44
82.73
20.56
15.46
61.81
84.3
Table 3: Detection-oriented performance comparison between RT-DETR-R18 and ECAF-Det on the IIG test set. Model-inference latency and end-to-end throughput were measured on an NVIDIA GeForce RTX 5070 Ti GPU with a batch size of 4. End-to-end FPS includes preprocessing, model inference, and postprocessing.
Metric
Value
Matched pairs
1062
Pred./GT area ratio, median (IQR)
0.912 (0.788–1.136)
Pred./GT width ratio, median (IQR)
0.960 (0.891–1.065)
Pred./GT height ratio, median (IQR)
0.960 (0.860–1.089)
Normalized center displacement, median (IQR)
0.080 (0.051–0.116)
Area ratio >1.25 (%)
17.23
Table 4: Bounding-box geometry statistics for ECAF-Det on the IIG test set. Statistics are computed for one-to-one matched detections with IoU ≥0.5 .
Figure 5: Qualitative comparison of detection results on the IIG dataset. All examples are original thermal infrared samples from the released IIG dataset. Each row shows the results of a different detector. ECAF-Det detects more weak and small plume regions in several challenging cases, but extremely faint and distant plumes remain difficult to localize accurately.
Condition
Model
AP (%)
AP 50 (%)
AP 75 (%)
AR 100 (%)
Plume visibility
Weak
RT-DETR-R18
24.8
75.4
4.5
51.2
ECAF-Det
29.5
88.0
4.9
53.8
Moderate
RT-DETR-R18
26.9
79.1
9.3
53.9
ECAF-Det
32.1
87.2
11.9
56.5
Strong
RT-DETR-R18
28.4
74.5
12.1
49.8
Table 5: Challenge-oriented detection performance on the IIG test set.
Local–Global Plume Block
MSEPM
AP
AP 50
AP S
AP M
AP L
GFLOPs (G)
Params (M)
×
×
26.8
77.8
19.9
32.0
37.1
56.9
19.873
✓
×
28.5
80.7
22.2
32.8
37.3
51.9
15.965
×
✓
27.9
80.4
24.0
30.6
38.1
59.8
21.119
✓
✓
29.5
81.0
24.3
33.5
38.9
54.8
17.211
Table 6: Ablation study of the plume-oriented local–global backbone on the IIG dataset. The local–global plume block denotes the proposed block for local detail preservation and frequency-domain contextual aggregation. MSEPM denotes the multi-scale edge perception module using GGEO to generate hierarchical boundary-sensitive structural priors.
Plume-Oriented Local–Global Backbone
CASR-PAN
AP
AP 50
AP S
AP M
AP L
GFLOPs (G)
Params (M)
×
×
26.8
77.8
19.9
32.0
37.1
56.9
19.873
✓
×
29.5
81.0
24.3
33.5
38.9
54.8
17.211
×
✓
29.4
79.1
23.0
33.6
42.2
45.8
16.939
✓
✓
29.8
84.3
25.3
32.5
42.6
43.7
14.277
Table 7: Ablation study of the complete ECAF-Det on the IIG dataset.
Configuration
AP
AP 50
AP S
AP M
AP L
GFLOPs
Params
All BasicBlocks baseline
26.8
77.8
19.9
32.0
37.1
56.9
19.873
P2–P5 local–global plume blocks
27.4
77.9
21.1
31.1
35.7
47.1
15.725
P4–P5 local–global plume blocks
28.5
80.7
22.2
32.8
37.3
51.9
15.965
P5 local–global plume block
27.0
78.6
22.4
30.2
42.8
54.5
16.743
Table 8: Stage-wise placement of local–global plume blocks in the ResNet-18 backbone on the IIG dataset.
Figure 6: Effective receptive field (ERF) visualization comparison between (a) the proposed plume-oriented local–global backbone and (b) the original ResNet-18 backbone under the same training and inference settings.
Figure 7: Comparison of high-contribution area ratios between the proposed plume-oriented local–global backbone and the original ResNet-18 backbone under different cumulative contribution thresholds.
Figure 8: Performance comparison of different backbone architectures on the IIG dataset. The marked point denotes the proposed plume-oriented local–global backbone.
Figure 9: Visualization results of the proposed gradient–Gabor edge operator (GGEO). (a) Original infrared image; (b) directional-gradient response extracted by the fixed Sobel-like operators; (c) Gabor-response cue obtained from the depthwise Gabor branch; (d) fused structural response combining directional-gradient and Gabor cues.
Figure 10: Visualization of the hierarchical structural representations involved in the multi-scale edge perception module (MSEPM). (a) Original infrared image; (b) E0 : the initial structural representation generated by GGEO before downsampling; (c) E1 : the representation after one 2×2 max-pooling operation; (d) E2 : the representation after two successive max-pooling operations; and (e) E3 : the representation after three successive max-pooling operations. The downsampled representations E1 , E2 , and E3 are subsequently projected through scale-specific 1×1 convolutions to form the hierarchical structural priors used by the corresponding backbone stages.
Gradient Directions
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs (G)
Params (M)
0∘
26.7
79.8
6.3
21.4
31.1
33.8
59.7
21.117
0∘ , 90∘
28.8
81.1
9.4
22.1
33.1
42.3
59.7
21.118
0∘ , 45∘ , 90∘ , 135∘
27.9
80.4
10.3
24.0
30.6
38.1
59.8
21.119
Table 9: Ablation study of gradient directions in GGEO on the IIG dataset.
Figure 11: Effect of the fusion coefficient α in GGEO. Performance comparisons under different fixed values of α and learnable α are reported in terms of (a) AP, (b) AP 50 , and (c) AP S .
Figure 12: Training evolution of the learnable fusion coefficient α in GGEO over 200 epochs. After an initial increasing phase, α fluctuates within a narrow range, indicating a stable fusion preference during training.
Operator
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs
Params
Sobel ( Heath et al., 1998 )
27.9
79.7
8.1
21.5
32.9
33.6
59.4
21.109
Canny ( Canny, 1986 )
25.8
80.0
7.2
22.1
27.7
31.2
60.1
21.110
Laplacian ( Wang, 2007 )
26.7
79.4
8.2
21.9
30.1
29.7
59.8
21.110
GGEO
27.9
80.4
10.3
24.0
30.6
38.1
59.8
21.119
Table 10: Comparison of edge operators and controlled edge-prior replacements within MSEPM on the IIG dataset.
Model
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs (G)
Params (M)
Routing-path ablation
Naive additive fusion
25.9
75.9
8.5
23.6
28.0
38.3
45.8
16.939
Without deep-to-mid fusion
27.5
79.5
7.6
22.5
31.1
36.8
45.8
16.939
Without deep-to-shallow fusion
25.7
76.2
6.8
19.2
30.5
40.6
45.8
16.939
Without shallow-to-mid fusion
25.7
75.8
6.2
19.8
29.7
40.2
45.8
16.939
Without mid-level self-fusion
28.3
80.3
8.3
22.3
31.9
43.5
45.8
16.939
Table 11: Ablation study of CASR-PAN components on the IIG dataset. The first group evaluates the adaptive routing paths, while the second group examines the importance-estimator branches and branch-weighting strategy.
Analysis
β
γ
AP
AP 50
AP 75
AP S
AP M
AP L
β sensitivity ( γ=1.0 )
β variation
0.00
1.00
27.5
78.4
8.6
21.6
31.1
42.6
β variation
0.25
1.00
28.2
77.8
10.3
21.9
32.2
35.8
Default
0.50
1.00
29.4
79.1
12.3
23.0
33.6
42.2
β variation
0.75
1.00
25.7
76.5
6.8
21.3
28.5
39.0
β variation
1.00
1.00
25.9
76.9
7.4
20.7
29.8
36.3
Table 12: Sensitivity analysis of the fixed modulation coefficients β and γ in CASR-PAN on the IIG dataset. One coefficient is varied while the other is fixed at its default value.
Model
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs (G)
Params (M)
PANet ( Liu et al., 2018 )
27.3
80.9
8.0
22.1
30.0
39.2
103.8
21.463
BiFPN ( Tan et al., 2020 )
27.0
76.6
8.3
20.8
31.3
38.9
64.3
20.300
NAS-FPN ( Ghiasi et al., 2019 )
25.1
76.4
7.5
19.3
29.9
26.1
93.8
19.756
CASR-PAN
29.4
79.1
12.3
23.0
33.6
42.2
45.8
16.939
Table 13: Comparison of different neck structures on the IIG dataset.
Global-context branch
AP
AP 50
AP 75
AP S
AP M
AP L
GFLOPs (G)
Params (M)
31×31 large-kernel convolution
25.3
74.5
7.3
17.8
31.2
31.6
55.1
18.095
Non-local module
26.6
77.9
8.6
20.3
31.1
33.9
53.6
17.274
Swin shifted-window attention
26.8
76.9
9.6
22.2
29.6
39.3
72.8
31.743
Proposed DCT-based frequency branch
28.5
80.7
8.8
22.2
32.8
37.3
51.9
15.965
Table 14: Comparison of alternative global-context modeling strategies within the local–global plume block on the IIG dataset.
Abbr./Symbol
Meaning
Abbr./Symbol
Meaning
RT-DETR
Real-Time Detection Transformer
ECAF-Det
Edge-Aware and Content-Adaptive Feature Fusion Detector