SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection
Authors: Yuting Zhao, Ziyi Zheng, Shuxiao Li
Organizations: Institute of Automation, Chinese Academy of Sciences, Beijing, China · School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China · Wuhan College, Wuhan, China
Camera-LiDAR fusion has become a prevailing paradigm for 3D object detection in autonomous driving. However, existing fusion detectors often establish strong inter-modality dependencies by decoding object queries from tightly coupled multimodal representations. Under corrupted driving conditions, such dependencies make the detector vulnerable to unreliable modalities, where degraded observations may interfere with reliable modality-specific evidence and lead to suboptimal predictions. Moreover, modality reliability can vary across both global driving scenes and individual object queries, requiring adaptive fusion decisions at a finer granularity. To bridge this gap, we reformulate robust camera-LiDAR fusion as a scene-aware branch routing problem and propose SARFusion, a robust 3D object detector. Instead of producing detections from a single fused representation, SARFusion decouples object-query decoding into three parallel reasoning branches: a camera branch, a LiDAR branch, and a camera-LiDAR fusion branch. Guided by a Scene Reliability Prior estimated from the global driving context, SARFusion further incorporates object-level evidence to route each query to the most suitable branch. This query-wise routing strategy alleviates harmful cross-modal interference while preserving the benefits of multimodal fusion when complementary cues are trustworthy. On the nuScenes test set, SARFusion achieves strong performance with 72.5 mAP and 74.4 NDS. Extensive analyses demonstrate its robustness under challenging conditions, including sensor corruptions and environmental changes.
Figures & tables
Figure 1 : Overall architecture of the proposed SARFusion.
Figure 2 : The illustration of SRP module
Figure 3 : The illustration of SABR module.
Figure 4 : Comparison of camera-LiDAR observations under degraded sensing conditions.
Modality
Method
Validation
Test
mAP
NDS
mAP
NDS
Camera
FCOS3D [ 43 ]
34.3
41.5
35.8
42.8
PETR [ 27 ]
37.0
44.2
39.1
45.5
BEVDet [ 15 ]
–
–
42.2
48.2
LiDAR
SECOND [ 49 ]
52.6
63.0
52.8
63.3
CenterPoint [ 51 ]
59.6
66.8
60.3
67.3
Table 1 : Comparison with representative 3D object detection methods on the nuScenes validation and test sets. All results are reported in percentage.
Figure 5 : Qualitative detection results under representative adverse weather conditions, including fog, snow, and rain.
Method
Clean
Snow
Rain
Fog
Strong Sunlight
FCOS3D [ 43 ]
23.9
2.0
13.0
13.5
17.2
DETR3D [ 45 ]
34.7
5.1
20.4
27.9
34.7
PointPillars [ 23 ]
27.7
27.6
27.7
24.5
23.7
CenterPoint [ 51 ]
59.3
55.9
56.1
43.8
54.2
BEVFusion [ 30 ]
68.5
62.8
66.1
54.1
64.4
TransFusion [ 2 ]
66.4
63.3
65.4
53.7
55.1
Table 2 : Robustness comparison on the nuScenes validation set under representative weather and lighting conditions.
Setting
mAP
NDS
Single fusion branch
70.7
73.1
Multi-branch w/o routing
52.4
62.8
+ Scene-Aware Branch Routing
70.8
73.1
+ Scene Reliability Prior
71.1
73.7
Table 3 : Component ablation of SARFusion on the nuScenes validation set.
Setting
mAP
NDS
w/o scene prior supervision
69.6
72.1
w/ scene prior supervision
71.1
73.7
Table 4 : Effect of scene prior supervision on the nuScenes validation set.
Routing Evidence
mAP
NDS
Local evidence only
70.8
73.1
Scene prior + local evidence
71.1
73.7
Table 5 : Ablation of routing evidence on the nuScenes validation set.
Training Strategy
mAP
NDS
Direct training
65.7
69.6
Staged training
71.1
73.7
Table 6 : Effect of training strategy on the nuScenes validation set.
Robust 3D object detection in adverse weather conditions is challenging due to sensor limitations. Although combining complementary modalities such as LiDAR and 4D RADAR has shown promise, the sparsity of these sensors becomes apparent in adverse weather with reduced reflections, leading to objects with few or no point cloud returns. To address this limitation, camera sensors provide visual cues even when LiDAR and RADAR signals are weakened. However, cameras themselves are also vulnerable to adverse weather, where some regions become unreliable due to snow or rain occluding the camera lens. While some camera-fusion methods designed for adverse weather learn to weigh image regions via confidence maps, these maps receive no direct supervision and are learned solely through the detection loss. We introduce Reliability-Aware Fusion (RAF), which explicitly supervises per-pixel reliability estimation and provides a direct learning signal for identifying and suppressing unreliable visual cues. Our framework leverages pretrained LiDAR-RADAR networks, keeping their backbones frozen while only training the added camera branch, BEV fusion encoder, and detection head. Extensive experiments on the K-Radar and VoD datasets demonstrate that integrating RAF consistently improves detection accuracy over LiDAR-RADAR baselines, achieving up to +6.5 APBEV and +7.4 AP3D gains. Code is available at https://github.com/parkie0517/RAF.
In autonomous driving, camera-radar fusion offers complementary sensing and low deployment cost. Existing methods perform fusion through input mixing, feature map mixing, or query-based feature sampling. We propose a new fusion paradigm, termed heterogeneous query interaction, and present ConFusion, a camera-radar 3D object detector. ConFusion combines image queries, radar queries, and learnable world queries distributed in 3D space to improve query initialization and object coverage. To encourage cross-type interaction among heterogeneous queries, we introduce heterogeneous query mixing (QMix), which performs dedicated cross-type attention after feature sampling to consolidate complementary object evidence. We further propose interactive query swap sampling (QSwap), which improves feature sampling by allowing related queries to exchange informative feature tokens under attention and geometric constraints. Experiments on the nuScenes dataset show that ConFusion achieves state-of-the-art performance, reaching 59.1 mAP and 65.6 NDS on the validation set, and 61.6 mAP and 67.9 NDS on the test set.
Jialong Wu, Yihan Wang, Matthias Rottmann
1Osnabrück University · 3Aptiv Services Deutschland GmbH · University of Wuppertal
Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. Typically, existing methods transform multimodal data from independent sensors, such as camera and LiDAR, into a unified bird's-eye view (BEV) representation for fusion. Although effective in ideal conditions, this strategy suffers from substantial performance deterioration when camera or LiDAR data are missing, corrupted, or noisy. To address this vulnerability, we develop a framework-agnostic fusion module for camera and LiDAR data that allows for handling cases when one of the two modalities is missing or corrupted. To demonstrate the effectiveness of our module, we instantiate it in BEVFusion [1], a well-established framework to combine camera and LiDAR data for 3D object detection. By means of quantitative experiments on the MultiCorrupt dataset, we demonstrate that our module achieves favorable performance improvements under scenarios of missing and corrupted modalities, substantially outperforming existing unified representation approaches across a wide range of sensor deterioration scenarios and reaching state-of-the-art performance in scenarios of corrupted modality due to extreme weather conditions and sensor failure.
Markus Essl, Marta Moscati, Mubashir Noman +4
Johannes Kepler University Linz, Austria · MBZUAI, UAE · Macquarie University, Sydney, Australia +1