Organizations: University of Science, VNU-HCM, Ho Chi Minh City 700000, Vietnam · Michigan State University, East Lansing, MI 48824, USA · Center for Environmental Intelligence, VinUniversity, Hanoi 100000, Vietnam · University of Arkansas, Fayetteville, AR 72701, USA
Camera-LiDAR detectors can continue to consume unreliable features even when both sensors remain present, synchronized, and calibrated. We introduce \emph{Boundary Feature Repair} (BFR), a frozen-host adaptation framework that learns task-supervised residual corrections at modality interfaces the detector already consumes. BFR-C repairs each camera feature level read by fusion, whereas BFR-L aligns host-conditioned LiDAR candidates to a selected boundary and routes site-wise innovations relative to the frozen anchor. Their jointly trained composition is BFR-CL. Zero-initialized per-channel scales make every variant an exact detector-level identity before optimization; only the repair modules train, while the encoders, fusion consumer, router, detection head, and host normalization statistics remain fixed. At inference, BFR requires neither clean references, corruption metadata, temporal history, nor online updates. Across the complete 20-corruption, five-severity KITTI-C grid, BFR-C reduces RCE from 14.07 to 11.92 on MVX-Net and from 14.29 to 11.00 on Focals Conv-F relative to their reproduced frozen baselines. On the latter host, BFR-L raises APcor from 73.65 to 74.48, while BFR-CL reaches 77.01 APcor and 10.46 RCE with 86.02 clean AP. On nuScenes-R, BFR-CL raises the reproduced MoME baseline's mAP robustness ratio from 80.1 to 81.4. These results establish boundary repair as a targeted retrofit for corrupted-but-present sensing without retraining the deployed detector.
Figures & tables
Fig. 1: Frozen-host boundary repair. BFR-C and BFR-L repair host-consumed camera and LiDAR features while the encoders, fusion consumer, router, and detection head (snowflakes) remain frozen.
Fig. 2: Boundary repair mechanisms. (a) BFR-C applies an independent project–spatial–restore residual Rl to every camera feature level consumed by the host, then gates the update with a zero-initialized channel scale γlC . (b) BFR-L aligns BEV-adapter, point-, and range-derived candidates Ba , Bp , and Br , routes their anchor-relative innovations with a learned-center heavy-tailed gate, and writes the mixture through γL . The sensor encoders remain frozen.
Corruption
LC Fusion
MVX-Net [ 2 ]
Focals Conv-F [ 16 ]
EPNet [ 14 ] †
LoGoNet [ 15 ] ∗
RoboFusion-L [ 8 ]
RoboFusion-B [ 8 ]
RoboFusion-T [ 8 ]
(repro.)
+ BFR-C
(repro.)
+ BFR-C
+ BFR-L
+ BFR-CL
None (AP clean )
82.72
85.04
88.04
87.87
87.60
78.59
79.33
85.92
85.84
86.17
86.02
Weather
Snow
34.58
51.45
85.29
84.70
84.60
27.29
37.91
34.56
54.65
38.77
54.70
Rain
36.27
55.80
86.48
85.54
84.79
43.85
48.95
41.99
61.70
46.82
62.17
Fog
44.35
67.53
85.53
84.00
84.17
63.03
68.65
45.08
61.25
47.19
60.26
S.L.
69.65
75.54
85.50
85.15
84.75
73.80
75.46
80.04
81.79
81.59
82.39
TABLE I: KITTI-C Car 3D AP 40 at moderate difficulty. BFR variants are compared with their reproduced frozen baselines; † and ∗ mark benchmark-reported and RoboFusion-author reproduced results, respectively.
Fig. 3: KITTI-C severity-2 Car examples for frozen Focals Conv-F (top) and joint BFR-CL (bottom). Black, green, and red denote ground truth, true positives, and false positives; dashed blue circles mark host misses detected by BFR-CL. Both use score ≥0.5 and center distance ≤1.5 m.
Corruption
LC Fusion
MoME [ 9 ]
BEVFusion [ 4 ] †
DeepInteraction [ 6 ] ∗
RoboFusion-L [ 8 ]
(repro.)
+ BFR-CL
None ( mAPclean )
68.45
69.90
69.91
71.19
70.65
Weather
Snow
62.84
62.36
67.12
64.00
64.72
Rain
66.13
66.48
67.58
68.63
68.22
Fog
54.10
54.79
67.01
68.50
68.06
S.L.
64.42
64.93
67.24
60.20
66.76
TABLE II: nuScenes-C mAP under common corruptions. LC-fusion results average five severities; the reproduced MoME baseline and BFR-CL use severity 2. † and ∗ mark published and author-reproduced results.
Method
Clean
Perf. ratio ( R )
LiDAR failures
Camera failures
Beam reduction 4 beams
LiDAR drop all
Limited FOV [−60,60]
Object failure rate =0.5
View drop 6 drops
Occlusion w/ obstacle
mAP
NDS
mAP
NDS
mAP
NDS
mAP
NDS
mAP
NDS
mAP
NDS
mAP
NDS
mAP
NDS
UniBEV [ 22 ] †
64.2
68.5
75.8
83.7
45.1
56.1
35.0
42.2
35.9
50.1
57.1
64.0
58.2
65.3
60.5
66.4
TransFusion [ 5 ]
66.9
70.9
54.4
66.8
–
–
0.0
0.0
20.3
45.8
34.6
53.6
61.6
67.4
65.5
70.0
MetaBEV [ 7 ]
68.0
71.5
75.4
82.5
–
57.7
39.0
42.6
–
57.7
–
67.6
63.6
69.2
–
70.0
CMT [ 17 ] †
70.3
72.9
78.4
84.4
54.9
62.2
38.3
44.7
43.9
54.0
66.7
70.4
61.7
68.1
65.0
69.8
TABLE III: nuScenes-R robustness under structured sensor failures. R is the mean corrupted-to-clean performance ratio across the displayed failures.
Gate kernel
16 corruptions
8-corruption subset
Constant
+1.10
+0.64
Softmax ( e−d2 )
+1.12
+0.68
Laplace ( e−d )
+1.19
+0.85
Polynomial tail
+1.27
+0.90
TABLE IV: Gate-kernel ablation for sparse BFR-L on Focals Conv-F. Entries are severity-2 mean Δ AP 40 relative to the reproduced frozen host.
Model / route
Params (M)
Latency (ms)
FPS
VRAM (MiB)
Plug-in GFLOPs
RoboFusion
RoboFusion-L [ 8 ]
97.54
–
3.1
–
–
RoboFusion-B [ 8 ]
81.01
–
3.5
–
–
RoboFusion-T [ 8 ]
13.94
–
6.0
–
–
Boundary Feature Repair
Focals Conv-F [ 16 ] (host)
47.49
43.5
23.0
1009
–
TABLE V: Model complexity and inference cost. Parameters are totals, and GFLOPs are incremental BFR overhead; dashes denote unavailable measurements.
Multimodal sensor fusion has demonstrated remarkable performance improvements over unimodal approaches in 3D object detection for autonomous vehicles. Typically, existing methods transform multimodal data from independent sensors, such as camera and LiDAR, into a unified bird's-eye view (BEV) representation for fusion. Although effective in ideal conditions, this strategy suffers from substantial performance deterioration when camera or LiDAR data are missing, corrupted, or noisy. To address this vulnerability, we develop a framework-agnostic fusion module for camera and LiDAR data that allows for handling cases when one of the two modalities is missing or corrupted. To demonstrate the effectiveness of our module, we instantiate it in BEVFusion [1], a well-established framework to combine camera and LiDAR data for 3D object detection. By means of quantitative experiments on the MultiCorrupt dataset, we demonstrate that our module achieves favorable performance improvements under scenarios of missing and corrupted modalities, substantially outperforming existing unified representation approaches across a wide range of sensor deterioration scenarios and reaching state-of-the-art performance in scenarios of corrupted modality due to extreme weather conditions and sensor failure.
Markus Essl, Marta Moscati, Mubashir Noman +4
Johannes Kepler University Linz, Austria · MBZUAI, UAE · Macquarie University, Sydney, Australia +1
Robust 3D object detection in adverse weather conditions is challenging due to sensor limitations. Although combining complementary modalities such as LiDAR and 4D RADAR has shown promise, the sparsity of these sensors becomes apparent in adverse weather with reduced reflections, leading to objects with few or no point cloud returns. To address this limitation, camera sensors provide visual cues even when LiDAR and RADAR signals are weakened. However, cameras themselves are also vulnerable to adverse weather, where some regions become unreliable due to snow or rain occluding the camera lens. While some camera-fusion methods designed for adverse weather learn to weigh image regions via confidence maps, these maps receive no direct supervision and are learned solely through the detection loss. We introduce Reliability-Aware Fusion (RAF), which explicitly supervises per-pixel reliability estimation and provides a direct learning signal for identifying and suppressing unreliable visual cues. Our framework leverages pretrained LiDAR-RADAR networks, keeping their backbones frozen while only training the added camera branch, BEV fusion encoder, and detection head. Extensive experiments on the K-Radar and VoD datasets demonstrate that integrating RAF consistently improves detection accuracy over LiDAR-RADAR baselines, achieving up to +6.5 APBEV and +7.4 AP3D gains. Code is available at https://github.com/parkie0517/RAF.
The structural vulnerabilities of point cloud-based 3D object detectors remain poorly understood. Prior work has studied adversarial robustness primarily on isolated 3D object models, while recent LiDAR spoofing attacks target richer and more realistic driving scenes but focus mainly on physical realizability rather than understanding detector behavior or attack efficiency. In this work, we investigate how LiDAR-based detectors rely on spatial evidence in complex scenes and whether these reliance patterns can be exploited to induce failures more efficiently. To this end, we propose an explainability-guided adversarial analysis methodology. We introduce the Saliency-LiDAR (SALL) method, which aggregates Integrated Gradient attributions across scenes to produce universal saliency maps for LiDAR-based 3D object detectors. Guided by these maps, we design the Explainability-aware Frustum Attack (EFA), which selectively perturbs only the most influential frustums rather than uniformly attacking entire object regions. Experiments on KITTI and nuScenes, across detectors such as PointPillars and SECOND, show that EFA reduces detection recall by more than 15 percentage points while requiring 25-50% fewer perturbed frustums than the state-of-the-art non-saliency-aware baseline. These findings reveal that modern 3D detectors concentrate discriminative evidence in a small subset of spatial regions, exposing a structural robustness vulnerability in current LiDAR perception systems. Our code is released at https://github.com/SecMindLab/Saliency_LiDAR.