Organizations: School of Electronic and Optical Engineering, Nanjing University of Science and Technology, Nanjing 210094, China · 28th Institute of China Electronics Technology Group Corporation, Nanjing 210007, China · School of Electronic Science and Technology, Shanghai Institute of Technical Physics, Chinese Academy of Sciences, Shanghai 200083, China · State Key Laboratory of Extreme Environment Optoelectronic Dynamic Measurement Technology and Instrument, Nanjing University of Science and Technology, Nanjing 210094, China
In aerial RGB--IR object detection, effectively exploiting complementary information across modalities is critical for robust perception under complex illumination and environmental conditions. Existing multimodal detectors mainly focus on spatial-domain interaction or frequency-specific feature enhancement, while the cross-modal interaction patterns of different frequency components remain insufficiently explored. Moreover, spectral discrepancy itself may contain both useful complementary cues and unreliable modality-specific responses, making indiscriminate frequency fusion suboptimal. To address these issues, we propose FoCal, a frequency-oriented framework for aerial RGB--IR object detection. First, a Frequency-Aware Dual-Domain Calibration (FADC) module is developed to explicitly model frequency-dependent cross-modal interaction. Low-frequency components are collaboratively consolidated into a shared structural consensus, whereas high-frequency components preserve modality-specific information through selective cross-modal exchange. The resulting frequency-aware cues are further transferred to the original feature domain to regulate cross-modal calibration. Second, we introduce a Discrepancy-Guided Spectral Modulation (DGSM) module, which characterizes cross-modal spectral imbalance using confidence-weighted relative amplitude discrepancy and transforms it into a bounded signed gate for adaptive enhancement, preservation, or attenuation of the joint multimodal spectrum. Extensive experiments on DroneVehicle, ESCVehicle, and ATR-UMOD demonstrate the effectiveness of FoCal, yielding mAP50 values of 83.5%, 54.8%, and 64.6%, respectively. Meanwhile, with only 3.0M parameters, FoCal achieves 113.6 FPS while preserving leading detection accuracy, highlighting a favorable accuracy--efficiency trade-off. Code is available at {https://github.com/universeliang/FoCal.
Figures & tables
Fig. 3: Overall architecture of the proposed FoCal. The upper part presents the macroscopic detection pipeline, which consists of a Backbone for feature extraction, a Path Aggregation Network (PAN) for multi-scale feature fusion, and a Detection Head.
Fig. 4: Overall architecture of the proposed DGSM
Methods
Pub.
Modality
mAP 50
mAP
Car
Truck
Bus
Van
Freight car
Param#
S 2 ANet [ 35 ]
TGRS’21
RGB
61.0
31.4
80.0
54.2
84.9
43.8
42.2
70.7M
YOLO11n † [ 34 ]
Ultralytics’24
RGB
69.1
47.9
93.6
62.0
90.4
50.4
49.2
2.7M
S 2 ANet [ 35 ]
TGRS’21
IR
67.5
40.4
89.9
54.5
88.9
48.4
55.8
70.7M
YOLO11n † [ 34 ]
Ultralytics’24
IR
79.2
63.3
98.3
76.8
95.0
60.4
65.7
2.7M
YOLO11n-Dual † [ 34 ]
Ultralytics’24
RGB+IR
80.8
65.2
98.5
80.5
95.8
62.2
67.0
4.0M
C 2 Former [ 36 ]
TGRS’24
RGB+IR
74.2
47.5
90.2
68.3
89.8
58.5
64.4
118.5M
TABLE I: Comparison of FoCal with other methods on the DroneVehicle test set. The terms “RGB” and “IR” denote detection outcomes from visible images and infrared images, respectively, while “RGB+IR” reflects combined detection outcomes from fused infrared and visible images. The best and second-best results in each category are highlighted in red and blue , respectively. † denotes the results reproduced by ourselves under the same training and evaluation settings.
Method
Modality
mAP 50
mAP
Param#
FLOPs
RoI Transformer [ 43 ]
RGB
37.2
19.3
87.2M
148.7G
S 2 ANet [ 35 ]
RGB
35.5
17.5
70.7M
120.6G
Oriented R-CNN [ 44 ]
RGB
39.3
20.6
73.2M
134.8G
KLD [ 45 ]
RGB
32.9
16.2
41.9M
131.2G
RoI Transformer [ 43 ]
IR
22.0
10.5
87.2M
148.7G
S 2 ANet [ 35 ]
IR
21.1
9.7
70.7M
120.6G
TABLE II: Comparison of FoCal with other methods on the ESCVehicle dataset.
Detectors
Modality
CR
SV
VN
BS
FC
TK
TT
TR
CE
ER
ME
mAP 50
S 2 ANet [ 35 ]
RGB
34.2
44.9
45.9
69.2
24.4
37.4
5.6
22.5
49.5
31.2
25.6
35.5
ReDet [ 48 ]
36.9
52.5
51.6
74.8
33.5
48.1
16.7
40.7
61.4
36.8
32.9
41.1
RoI Transformer [ 43 ]
37.2
53.3
51.9
71.5
30.1
46.8
18.2
36.3
58.9
38.3
25.3
42.5
Oriented R-CNN [ 44 ]
36.9
52.5
51.6
74.8
33.5
48.1
16.7
40.7
61.7
36.8
32.9
44.2
YOLOv5s [ 34 ]
45.8
60.7
57.5
75.2
41.6
52.1
18.2
42.3
68.7
47.5
47.4
50.7
S 2 ANet [ 35 ]
IR
50.2
35.9
31.8
59.9
35.5
24.3
31.4
16.0
10.8
1.0
32.0
29.9
TABLE III: Comparison of FoCal with other methods on the ATR-UMOD dataset. All methods perform localization and classification using oriented bounding box (OBB) heads. The categories, car, SUV, van, bus, freight car, truck, motorcycle, trailer, excavator, crane, and tank truck, are abbreviated as CR, SV, VN, BS, FC, TK, ME, TR, ER, CE, and TT, respectively.
RR
FADC
DGSM
mAP 50
mAP 75
mAP
Param#
FLOPs
✗
✗
✗
80.8
76.6
65.2
4.01M
9.5G
✓
✗
✗
81.9
78.5
67.3
2.63M
6.8G
✓
✓
✗
82.9
79.5
68.2
2.75M
7.3G
✓
✗
✓
82.6
79.1
68.0
2.84M
7.9G
✓
✓
✓
83.5
79.9
68.6
2.96M
8.4G
TABLE IV: Ablation study of individual components in FoCal on the DroneVehicle dataset.
Variant
HF
LF
mAP 50
mAP 75
mAP
Param#
FLOPs
Sym.
Ex.
Ex.
82.3
78.8
67.6
2.71M
7.1G
Sym.
Con.
Con.
82.7
79.2
68.0
2.79M
7.5G
Reverse
Con.
Ex.
82.4
78.9
67.7
2.75M
7.3G
Ours
Ex.
Con.
82.9
79.5
68.2
2.75M
7.3G
TABLE V: Ablation study of different frequency interaction strategies on the DroneVehicle dataset. “Ex.” and “Con.” denote selective exchange and structural consensus, respectively.
Fig. 5: Qualitative comparison of the proposed method against other competing methods under complex aerial scenarios.
Fig. 6: Visualization of the FADC frequency pathway at P3. The top and bottom rows correspond to RGB and IR. From left to right, the panels show the input images, raw low-frequency (LF) components, the shared LF representation after cooperative interaction, raw high-frequency (HF) components, HF components after competitive interaction, and the recomposed LF+HF features. One shared LF map is produced, whereas the HF and recomposed features remain modality-specific.
Fig. 7: Comparison of heatmap visualizations across different input modalities. From top to bottom: visible and infrared modalities. From left to right: baseline model and the proposed FoCal.
Method
Param#
FLOPs
mAP 50
Latency (ms/pair) ↓
FPS ↑
CrossWeaver [ 42 ]
5.4M
19.5G
80.6
32.1
31.2
COMO [ 41 ]
6.5M
25.5G
81.2
15.5
64.5
FoCal (Ours)
3.0M
8.4G
83.5
8.8
113.6
TABLE VI: Comparison of model complexity and inference efficiency. Latency and FPS are measured using the trained .pt models in PyTorch on an NVIDIA RTX 4090 with a batch size of 1.
Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object detection. EGM-Det employs a dual-stream architecture to preserve modality-specific representations and introduces an Entropy Offset Gate Fusion module for adaptive multi-scale fusion. The module derives shallow entropy priors from input intensity, local entropy, and cross-modal discrepancy, and uses them to guide local offset alignment and spatial-channel gated fusion. It therefore selectively aggregates reliable RGB and infrared cues instead of uniformly combining heterogeneous features. We further introduce cross-modal distillation to regularize the learned fusion gates and reduce fusion degradation. Each student branch extracts complementary knowledge from the cross-modality teacher branch matched to the main branch, while entropy-adaptive supervision emphasizes uncertain modality decisions. Experiments on DroneVehicle, LLVIP, and VEDAI demonstrate state-of-the-art performance across all three benchmarks; in particular, EGM-Det outperforms prior approaches by more than 10 percentage points on VEDAI.
Cunzheng Fan, Dawei Yan, Guanlin Wang +4
School of Cybersecurity, Northwestern Polytechnical University, Xi’an 710072, China · School of Automation and Software Engineering, Shanxi University, Taiyuan 030006, China
Infrared-visible object detection improves detection performance by combining complementary features from multispectral images. Existing backbone-specific and backbone-shared approaches still suffer from the problems of severe bias of modality-shared features and the insufficiency of modality-specific features. To address these issues, we propose a novel detection framework WD-FQDet that explicitly decouples modality-shared and modality-specific information from infrared and visible modalities in the new view of low- and high-frequency domains, allowing fusion strategies tailored to their frequency characteristics. Specifically, a low-frequency homogeneity alignment module is proposed to align modality-shared features across modalities via a cross-modal attention mechanism, and a high-frequency specificity retention module is proposed to preserve modality-specific features through the multi-scale gradient consistency loss. To reinforce the feature representation in the frequency domain, we propose a hybrid feature enhancement module that incorporates spatial cues. Furthermore, considering that the contributions of homogeneous and modality-specific features to object detection vary across scenarios, we propose a frequency-aware query selection module to dynamically regulate their contributions. Experimental results on the FLIR, LLVIP, and M3FD datasets demonstrate that WD-FQDet achieves state-of-the-art performance across multiple evaluation metrics.
Chunjin Yang, Xiwei Zhang, Yiming Xiao +1
University of Electronic Science and Technology of China,Chengdu, Sichuan 611731, China
Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rely on implicit feature adaptation to handle weakly misaligned scenarios, while large-offset geometric discrepancies remain insufficiently addressed. In this paper, we propose a Joint Feature-domain Registration and Detection network (JFRDet), an end-to-end visible-infrared oriented object detector tailored for severely cross-modal geometric discrepancies. JFRDet introduces a Cross-Modal Affine Alignment (CMAA) module to estimate an image-level affine transformation for explicit multi-level feature alignment. Note that illumination changes directly affect the reliability of RGB cues, an Illumination-Guided Complementary Fusion (IGCF) module adaptively exploits modality reliability under varying illumination conditions for cross-modal fusion. Then, an Alignment Quality-Consistency Gating (AQCG) strategy stabilizes joint optimization by modulating detection supervision according to alignment reliability and gradient consistency. We further construct DroneVehicle Misaligned (DVMA), a benchmark for evaluating visible-infrared oriented object detection under severe cross-modal geometric misalignment. The proposed JFRDet achieves 69.7% mAP50 on DVMA, which represents state-of-the-art (SOTA) performance. The code and dataset will be available on GitHub.
Qi Ming, Yuyang Wang, Mingjing Zhao +7
Beijing University of Technology, China · Central South University, China · Beijing Electronics Science & Technology Institute +4