Organizations: School of Electronic and Optical Engineering, Nanjing University of Science and Technology, Nanjing 210094, China · 28th Institute of China Electronics Technology Group Corporation, Nanjing 210007, China · School of Electronic Science and Technology, Shanghai Institute of Technical Physics, Chinese Academy of Sciences, Shanghai 200083, China · State Key Laboratory of Extreme Environment Optoelectronic Dynamic Measurement Technology and Instrument, Nanjing University of Science and Technology, Nanjing 210094, China
In aerial RGB--IR object detection, effectively exploiting complementary information across modalities is critical for robust perception under complex illumination and environmental conditions. Existing multimodal detectors mainly focus on spatial-domain interaction or frequency-specific feature enhancement, while the cross-modal interaction patterns of different frequency components remain insufficiently explored. Moreover, spectral discrepancy itself may contain both useful complementary cues and unreliable modality-specific responses, making indiscriminate frequency fusion suboptimal. To address these issues, we propose FoCal, a frequency-oriented framework for aerial RGB--IR object detection. First, a Frequency-Aware Dual-Domain Calibration (FADC) module is developed to explicitly model frequency-dependent cross-modal interaction. Low-frequency components are collaboratively consolidated into a shared structural consensus, whereas high-frequency components preserve modality-specific information through selective cross-modal exchange. The resulting frequency-aware cues are further transferred to the original feature domain to regulate cross-modal calibration. Second, we introduce a Discrepancy-Guided Spectral Modulation (DGSM) module, which characterizes cross-modal spectral imbalance using confidence-weighted relative amplitude discrepancy and transforms it into a bounded signed gate for adaptive enhancement, preservation, or attenuation of the joint multimodal spectrum. Extensive experiments on DroneVehicle, ESCVehicle, and ATR-UMOD demonstrate the effectiveness of FoCal, yielding mAP50 values of 83.5%, 54.8%, and 64.6%, respectively. Meanwhile, with only 3.0M parameters, FoCal achieves 113.6 FPS while preserving leading detection accuracy, highlighting a favorable accuracy--efficiency trade-off. Code is available at {https://github.com/universeliang/FoCal.
Figures & tables
Fig. 3: Overall architecture of the proposed FoCal. The upper part presents the macroscopic detection pipeline, which consists of a Backbone for feature extraction, a Path Aggregation Network (PAN) for multi-scale feature fusion, and a Detection Head.
Fig. 4: Overall architecture of the proposed DGSM
Methods
Pub.
Modality
mAP 50
mAP
Car
Truck
Bus
Van
Freight car
Param#
S 2 ANet [ 35 ]
TGRS’21
RGB
61.0
31.4
80.0
54.2
84.9
43.8
42.2
70.7M
YOLO11n † [ 34 ]
Ultralytics’24
RGB
69.1
47.9
93.6
62.0
90.4
50.4
49.2
2.7M
S 2 ANet [ 35 ]
TGRS’21
IR
67.5
40.4
89.9
54.5
88.9
48.4
55.8
70.7M
YOLO11n † [ 34 ]
Ultralytics’24
IR
79.2
63.3
98.3
76.8
95.0
60.4
65.7
2.7M
YOLO11n-Dual † [ 34 ]
Ultralytics’24
RGB+IR
80.8
65.2
98.5
80.5
95.8
62.2
67.0
4.0M
C 2 Former [ 36 ]
TGRS’24
RGB+IR
74.2
47.5
90.2
68.3
89.8
58.5
64.4
118.5M
TABLE I: Comparison of FoCal with other methods on the DroneVehicle test set. The terms “RGB” and “IR” denote detection outcomes from visible images and infrared images, respectively, while “RGB+IR” reflects combined detection outcomes from fused infrared and visible images. The best and second-best results in each category are highlighted in red and blue , respectively. † denotes the results reproduced by ourselves under the same training and evaluation settings.
Method
Modality
mAP 50
mAP
Param#
FLOPs
RoI Transformer [ 43 ]
RGB
37.2
19.3
87.2M
148.7G
S 2 ANet [ 35 ]
RGB
35.5
17.5
70.7M
120.6G
Oriented R-CNN [ 44 ]
RGB
39.3
20.6
73.2M
134.8G
KLD [ 45 ]
RGB
32.9
16.2
41.9M
131.2G
RoI Transformer [ 43 ]
IR
22.0
10.5
87.2M
148.7G
S 2 ANet [ 35 ]
IR
21.1
9.7
70.7M
120.6G
TABLE II: Comparison of FoCal with other methods on the ESCVehicle dataset.
Detectors
Modality
CR
SV
VN
BS
FC
TK
TT
TR
CE
ER
ME
mAP 50
S 2 ANet [ 35 ]
RGB
34.2
44.9
45.9
69.2
24.4
37.4
5.6
22.5
49.5
31.2
25.6
35.5
ReDet [ 48 ]
36.9
52.5
51.6
74.8
33.5
48.1
16.7
40.7
61.4
36.8
32.9
41.1
RoI Transformer [ 43 ]
37.2
53.3
51.9
71.5
30.1
46.8
18.2
36.3
58.9
38.3
25.3
42.5
Oriented R-CNN [ 44 ]
36.9
52.5
51.6
74.8
33.5
48.1
16.7
40.7
61.7
36.8
32.9
44.2
YOLOv5s [ 34 ]
45.8
60.7
57.5
75.2
41.6
52.1
18.2
42.3
68.7
47.5
47.4
50.7
S 2 ANet [ 35 ]
IR
50.2
35.9
31.8
59.9
35.5
24.3
31.4
16.0
10.8
1.0
32.0
29.9
TABLE III: Comparison of FoCal with other methods on the ATR-UMOD dataset. All methods perform localization and classification using oriented bounding box (OBB) heads. The categories, car, SUV, van, bus, freight car, truck, motorcycle, trailer, excavator, crane, and tank truck, are abbreviated as CR, SV, VN, BS, FC, TK, ME, TR, ER, CE, and TT, respectively.
RR
FADC
DGSM
mAP 50
mAP 75
mAP
Param#
FLOPs
✗
✗
✗
80.8
76.6
65.2
4.01M
9.5G
✓
✗
✗
81.9
78.5
67.3
2.63M
6.8G
✓
✓
✗
82.9
79.5
68.2
2.75M
7.3G
✓
✗
✓
82.6
79.1
68.0
2.84M
7.9G
✓
✓
✓
83.5
79.9
68.6
2.96M
8.4G
TABLE IV: Ablation study of individual components in FoCal on the DroneVehicle dataset.
Variant
HF
LF
mAP 50
mAP 75
mAP
Param#
FLOPs
Sym.
Ex.
Ex.
82.3
78.8
67.6
2.71M
7.1G
Sym.
Con.
Con.
82.7
79.2
68.0
2.79M
7.5G
Reverse
Con.
Ex.
82.4
78.9
67.7
2.75M
7.3G
Ours
Ex.
Con.
82.9
79.5
68.2
2.75M
7.3G
TABLE V: Ablation study of different frequency interaction strategies on the DroneVehicle dataset. “Ex.” and “Con.” denote selective exchange and structural consensus, respectively.
Fig. 5: Qualitative comparison of the proposed method against other competing methods under complex aerial scenarios.
Fig. 6: Visualization of the FADC frequency pathway at P3. The top and bottom rows correspond to RGB and IR. From left to right, the panels show the input images, raw low-frequency (LF) components, the shared LF representation after cooperative interaction, raw high-frequency (HF) components, HF components after competitive interaction, and the recomposed LF+HF features. One shared LF map is produced, whereas the HF and recomposed features remain modality-specific.
Fig. 7: Comparison of heatmap visualizations across different input modalities. From top to bottom: visible and infrared modalities. From left to right: baseline model and the proposed FoCal.
Method
Param#
FLOPs
mAP 50
Latency (ms/pair) ↓
FPS ↑
CrossWeaver [ 42 ]
5.4M
19.5G
80.6
32.1
31.2
COMO [ 41 ]
6.5M
25.5G
81.2
15.5
64.5
FoCal (Ours)
3.0M
8.4G
83.5
8.8
113.6
TABLE VI: Comparison of model complexity and inference efficiency. Latency and FPS are measured using the trained .pt models in PyTorch on an NVIDIA RTX 4090 with a batch size of 1.
School of Cybersecurity, Northwestern Polytechnical University, Xi’an 710072, China · School of Automation and Software Engineering, Shanxi University, Taiyuan 030006, China