GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion
Authors: Santiago Montiel-Marín, Miguel Antunes-García, Fabio Sánchez-García, Angel Llamazares, Holger Caesar, Luis M. Bergasa
Organizations: Department of Electronics. University of Alcalá, Spain. · Department of Cognitive Robotics. Delft University of Technology, The Netherlands.
Robust and accurate perception of dynamic objects and map elements is crucial for autonomous vehicles performing safe navigation in complex traffic scenarios. While vision-only methods have become the de facto standard due to their technical advances, they can benefit from effective and cost-efficient fusion with radar measurements. In this work, we advance fusion methods by repurposing Gaussian Splatting as an efficient universal view transformer that bridges the view disparity gap, mapping both image pixels and radar points into a common Bird's-Eye View (BEV) representation. Our main contribution is GaussianCaR, an end-to-end network for BEV segmentation that, unlike prior BEV fusion methods, leverages Gaussian Splatting to map raw sensor information into latent features for efficient camera-radar fusion. Our architecture combines multi-scale fusion with a transformer decoder to efficiently extract BEV features. Experimental results demonstrate that our approach achieves performance on par with, or even surpassing, the state of the art on BEV segmentation tasks (57.3%, 82.9%, and 50.1% IoU for vehicles, roads, and lane dividers) on the nuScenes dataset, while maintaining a 3.2x faster inference runtime. Code and project page are available online.
Figures & tables
Fig. 1: We propose GaussianCaR , a novel method for efficient camera-radar fusion. We envision sensor fusion as a modality → Gaussians → BEV transformation, achieving competitive accuracy with significantly fast inference times for BEV segmentation tasks.
Fig. 2: Main diagram of our proposal, GaussianCaR. Given multi-view camera images and radar point clouds , we leverage GS as a universal view transformer and formulate sensor fusion as modality → Gaussians → BEV transformation. The model predicts BEV segmentation maps for dynamic vehicles and map elements. We employ two feature encoding branches: Pixels-to-Gaussians for camera features and Points-to-Gaussians for radar point clouds. Features are splatted and fused in BEV space using a CMX-based fuser, and decoded via a DPT decoder.
Fig. 3: Gaussian modeling process. In (a), we present the process of extracting a Gaussian from a discrete probability distribution; in (b), we depict the behavior of the offset head, displacing the final Gaussian position from the original set of candidates; in (c), we illustrate the Gaussian rasterization process, projecting Gaussians from 3D space to BEV space via orthographic projection.
Fig. 4: Our Pixels-to-Gaussians extracts low-resolution feature maps using an EfficientViT backbone and a neck. A set of convolutional heads predicts Gc Gaussians. To position the Gaussians in 3D space, camera intrinsic and extrinsic matrices are used.
Fig. 5: Our proposed Points-to-Gaussians module processes radar point clouds using a lightweight PTv3, composed of E encoder and D decoder blocks. A set of MLP heads then predicts Gr Gaussians, each parameterized by geometric and semantic attributes.
Method
Code
Cam Enc
Radar Enc
IoU ( ↑ )
Camera-only
BEVFormer [ 5 ]
✓
RN-101
-
43.2
GaussianLSS [ 22 ]
✓
RN-101
-
46.1
SimpleBEV [ 6 ]
✓
RN-101
-
47.4
PointBeV [ 29 ]
✓
EN-b4
-
47.8
GaussianBeV [ 23 ]
✗
EN-b4
-
50.3
TABLE I: BEV Vehicle Segmentation on the nuScenes Validation Set
Method
Driv. Area ( ↑ )
Lane Div. ( ↑ )
Camera-only
LSS [ 3 ]
72.9
20.0
BEVFormer [ 5 ]
80.1
25.7
GaussianBeV [ 23 ]
82.6
47.4
Camera-radar
BEVGuide [ 30 ]
76.7
44.2
TABLE II: BEV Map Segmentation on the nuScenes Validation Set
Method
IoU ( ↑ )
ms ( ↓ )
FPS ( ↑ )
Baseline: GaussianLSS [ 22 ]
46.1
53.9
18.6
Image Encoding Branch
+ EffViT L2
47.3
56.6
17.8
+ Offset Head
47.7
56.9
17.6
+ Early auxiliary loss
47.8
56.9
17.6
+ Dice loss
48.0
56.9
17.6
TABLE III: Ablation Study
Fig. 6: Qualitative results on the nuScenes validation set. Each row shows, from left to right: multi-view camera images, PCA camera latent features, PCA radar latent features, and predictions. For vehicle segmentation, we report an error map where correctness is indicated by color: correct , missing , and incorrect . For map segmentation, we report classes by color: drivable area , lane and road dividers , pedestrian crossings , walkway and carpark areas .
Long-range 3D object detection is critical for safe autonomous driving at highway speeds, yet existing radar-camera fusion methods remain limited at extended ranges. BEV-based methods capture scene-level context but incur rapidly growing computation and often lose fine-grained object detail, while query-based methods are efficient but provide limited scene-level context. Temporal fusion further requires both multi-frame accumulation for sparse distant observations and object-level motion modeling for fast-moving objects. We propose Horizon3D, a sparse radar-camera fusion framework for long-range 3D object detection that combines Gaussian primitives with sparse BEV features. Horizon3D initializes Gaussian primitives at radar- and camera-estimated object keypoints using Keypoint-Guided Gaussian Initialization, refines them through Object-Centric Sparse Fusion, and splats them onto the BEV plane to fuse object-level detail with sparse radar BEV context. It further introduces Dual-Path Temporal Fusion, which aggregates temporal cues through a BEV path for scene-level accumulation and a Gaussian path for object-level motion propagation. Experiments on TruckScenes show that Horizon3D achieves state-of-the-art radar-camera 3D detection performance. On the validation set, it outperforms the previous best method by +3.0 NDS and +1.6 mAP while maintaining competitive inference speed.
Geonho Bang, Geunju Baek, Dongyoung Lee +2
Seoul National University, Seoul, Republic of Korea
In autonomous driving, camera-radar fusion offers complementary sensing and low deployment cost. Existing methods perform fusion through input mixing, feature map mixing, or query-based feature sampling. We propose a new fusion paradigm, termed heterogeneous query interaction, and present ConFusion, a camera-radar 3D object detector. ConFusion combines image queries, radar queries, and learnable world queries distributed in 3D space to improve query initialization and object coverage. To encourage cross-type interaction among heterogeneous queries, we introduce heterogeneous query mixing (QMix), which performs dedicated cross-type attention after feature sampling to consolidate complementary object evidence. We further propose interactive query swap sampling (QSwap), which improves feature sampling by allowing related queries to exchange informative feature tokens under attention and geometric constraints. Experiments on the nuScenes dataset show that ConFusion achieves state-of-the-art performance, reaching 59.1 mAP and 65.6 NDS on the validation set, and 61.6 mAP and 67.9 NDS on the test set.
Jialong Wu, Yihan Wang, Matthias Rottmann
1Osnabrück University · 3Aptiv Services Deutschland GmbH · University of Wuppertal
A realistic view of the vehicle's surroundings is generally offered by camera sensors, which is crucial for environmental perception. Affordable radar sensors, on the other hand, are becoming invaluable due to their robustness in variable weather conditions. However, because of their noisy output and reduced classification capability, they work best when combined with other sensor data. Specifically, we address the challenge of multimodal sensor fusion by aligning radar and camera data in a unified domain, prioritizing not only accuracy, but also computational efficiency. Our work leverages the raw range-Doppler (RD) spectrum from radar and front-view camera images as inputs. To enable effective fusion, we employ a variational encoder-decoder architecture that learns the transformation of front-view camera data into the Bird's-Eye View (BEV) polar domain. Concurrently, a radar encoder-decoder learns to recover the angle information from the RD data that produce Range-Azimuth (RA) features. This alignment ensures that both modalities are represented in a compatible domain, facilitating robust and efficient sensor fusion. We evaluated our fusion strategy for vehicle detection and free space segmentation against state-of-the-art methods using the RADIal dataset.