Organizations: Dalian Maritime University, Dalian, China · The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China · Hubei University of Economics, Wuhan, China
Low-light UAV-based RGB-infrared oriented small-vehicle detection is important for nighttime traffic monitoring, emergency response, and urban inspection. Illumination variations, headlight glare, local shadows, and thermal-response degradation cause spatially varying modality reliability, while the small visual extent of vehicles further weakens boundaries, orientation cues, and thermal responses. Accordingly, selecting trustworthy observations based on local modality reliability while further exploiting complementary discriminative information in regions with ambiguous modality preference is key to constructing effective multimodal representations. Based on this insight, we propose ReDiffNet, a reliability-conditioned differential representation network in which modality reliability guides both evidence selection and complementary recovery. Specifically, degradation-aware reliability learning estimates relative spatial reliability, uncertainty-guided differential recovery exploits cross-modal differences to recover complementary cues in ambiguous regions, and reliability-conditioned reconstruction integrates retained and recovered evidence into a unified representation. ReDiffNet achieves 85.3% and 73.9% mAP50 on DroneVehicle and VEDAI, respectively, supporting its effectiveness.
Figures & tables
Figure 1: Illustration of the three core modules in ReDiffNet: (a) DARL for modality reliability estimation, (b) UGDR for complementary cue recovery, and (c) RCR for reliable feature reconstruction.
Method
Pub. + Year
RGB
Infrared
mAP50
RetinaNet [ 10 ]
ICCV 2017
✓
×
47.1
R3Det [ 11 ]
AAAI 2021
✓
×
60.8
S2ANet [ 12 ]
TGRS 2021
✓
×
61.0
Faster R-CNN [ 13 ]
TPAMI 2017
✓
×
55.9
RoITransformer [ 14 ]
CVPR 2019
✓
×
61.6
Oriented R-CNN [ 15 ]
ICCV 2021
✓
×
60.8
Table 1: Comparison with state-of-the-art methods on DroneVehicle. RGB and Infrared indicate the input modalities used by each method. All mAP50 values are percentages.
Method
RGB
Infrared
mAP50
RetinaNet [ 10 ]
✓
×
20.7
S2ANet [ 12 ]
✓
×
44.5
Faster R-CNN [ 13 ]
✓
×
61.5
RoITransformer [ 14 ]
✓
×
65.4
Oriented R-CNN [ 15 ]
✓
×
66.4
RetinaNet [ 10 ]
×
✓
18.7
Table 2: Comparison on the VEDAI dataset. All values are mAP50 .
Method
Params (M)
FLOPs (G)
FPS
mAP50
TSFADet [ 5 ]
104.7
109.8
18.6
73.9
CAGTDet [ 22 ]
–
120.6
17.8
74.6
C 2 Former [ 6 ]
100.8
89.9
–
74.2
DMM+S 2 A-Net [ 8 ]
87.97
–
–
79.4
CoDAF [ 20 ]
67.3
224.9
58.1
78.6
ReDiffNet (Ours)
10.53
27.42
106.9
85.3
Table 3: Reported model size and detection performance on DroneVehicle. All mAP50 values are percentages.
Method
Precision
Recall
F1
mAP50
mAP50:95
DARL removed
80.18
77.95
79.05
81.71
67.30
UGDR removed
79.53
75.95
77.70
80.48
66.16
RCR removed
79.12
76.86
77.98
80.91
66.62
Detection objective only
79.12
78.50
78.80
83.10
68.40
Full model
81.79
80.84
81.31
85.30
70.48
Table 4: Ablation study on DroneVehicle. All variants use the same training and evaluation protocol.
Condition
Factor
mAP50
mAP50:95
Drop 50
Clean
–
85.21
70.19
–
RGB darkening
0.75
84.80
69.80
0.41
RGB darkening
0.50
84.00
68.89
1.21
RGB darkening
0.30
82.77
67.63
2.44
IR contrast reduction
0.80
85.10
70.05
0.12
IR contrast reduction
0.60
84.66
69.62
0.55
Table 5: ReDiffNet under synthetic modality degradation on the DroneVehicle validation set; Drop 50 is in percentage points.
Detecting small unmanned aerial vehicles from RGB-infrared remote-sensing pairs remains challenging due to tiny target scale, cluttered backgrounds, and spatial misalignment between heterogeneous sensors. Existing bimodal detectors often align or fuse features without assessing the reliability of local cross-sensor correspondence, allowing mismatch artifacts to propagate into the detection head. To address this issue, we propose LER-YOLO, a reliability-aware sparse mixture-of-experts framework for misaligned RGB-infrared UAV detection. LER-YOLO first introduces an Uncertainty-Aware Target Alignment module that resamples visible features toward the infrared reference and estimates a spatial reliability map. This reliability prior is then used by a Reliability-Guided Sparse MoE Fusion module to adaptively select k experts from RGB-dominant, infrared-dominant, and interactive fusion experts, enabling trustworthy cross-modal interaction while suppressing unreliable fusion. Experiments on the public MBU benchmark under a YOLOv5s-family protocol show that LER-YOLO achieves 89.7+/-0.2% AP50 over three independent seeds, with a best result of 89.9%. Extensive ablations, parameter-matched comparisons, synthetic-shift evaluations, and complexity analysis demonstrate that the gains mainly come from reliability-guided expert routing rather than increased model capacity.
Liming Hou, Yueping Peng, Hexiang Hao +6
Engineering University of PAP, Xi’an 710086, China · Unit Command Department, Officers College of PAP, Chengdu 610213, China
Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object detection. EGM-Det employs a dual-stream architecture to preserve modality-specific representations and introduces an Entropy Offset Gate Fusion module for adaptive multi-scale fusion. The module derives shallow entropy priors from input intensity, local entropy, and cross-modal discrepancy, and uses them to guide local offset alignment and spatial-channel gated fusion. It therefore selectively aggregates reliable RGB and infrared cues instead of uniformly combining heterogeneous features. We further introduce cross-modal distillation to regularize the learned fusion gates and reduce fusion degradation. Each student branch extracts complementary knowledge from the cross-modality teacher branch matched to the main branch, while entropy-adaptive supervision emphasizes uncertain modality decisions. Experiments on DroneVehicle, LLVIP, and VEDAI demonstrate state-of-the-art performance across all three benchmarks; in particular, EGM-Det outperforms prior approaches by more than 10 percentage points on VEDAI.
Cunzheng Fan, Dawei Yan, Guanlin Wang +4
School of Cybersecurity, Northwestern Polytechnical University, Xi’an 710072, China · School of Automation and Software Engineering, Shanxi University, Taiyuan 030006, China
Reliable UAV object detection requires robustness to illumination changes, motion blur, and scene dynamics that suppress RGB cues. Thermal long-wave infrared (LWIR) sensing preserves contrast in low light, and event cameras retain microsecond-level temporal edges, but integrating all three modalities in a unified detector has not been systematically studied. We present a tri-modal framework that processes RGB, thermal, and event data with a dual-stream hierarchical vision transformer. At selected encoder depths, a Modality-Aware Gated Exchange (MAGE) applies inter-sensor channel and spatial gating, and a Bidirectional Token Exchange (BiTE) module performs bidirectional token-level attention with depthwise-pointwise refinement, producing resolution-preserving fused maps for a standard feature pyramid and two-stage detector. We introduce a 10,489-frame UAV dataset with synchronized and pre-aligned RGB-thermal-event streams and 24,223 annotated vehicles across day and night flights. Through 61 controlled ablations, we evaluate fusion placement, mechanism (baseline MAGE+BiTE, CSSA, GAFF), modality subsets, and backbone capacity. Tri-modal fusion improves over all dual-modal baselines, with fusion depth having a significant effect and a lightweight CSSA variant recovering most of the benefit at minimal cost. This work provides the first systematic benchmark and modular backbone for tri-modal UAV-based object detection.