GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion
Authors: Santiago Montiel-Marín, Miguel Antunes-García, Fabio Sánchez-García, Angel Llamazares, Holger Caesar, Luis M. Bergasa
Organizations: Department of Electronics. University of Alcalá, Spain. · Department of Cognitive Robotics. Delft University of Technology, The Netherlands.
Robust and accurate perception of dynamic objects and map elements is crucial for autonomous vehicles performing safe navigation in complex traffic scenarios. While vision-only methods have become the de facto standard due to their technical advances, they can benefit from effective and cost-efficient fusion with radar measurements. In this work, we advance fusion methods by repurposing Gaussian Splatting as an efficient universal view transformer that bridges the view disparity gap, mapping both image pixels and radar points into a common Bird's-Eye View (BEV) representation. Our main contribution is GaussianCaR, an end-to-end network for BEV segmentation that, unlike prior BEV fusion methods, leverages Gaussian Splatting to map raw sensor information into latent features for efficient camera-radar fusion. Our architecture combines multi-scale fusion with a transformer decoder to efficiently extract BEV features. Experimental results demonstrate that our approach achieves performance on par with, or even surpassing, the state of the art on BEV segmentation tasks (57.3%, 82.9%, and 50.1% IoU for vehicles, roads, and lane dividers) on the nuScenes dataset, while maintaining a 3.2x faster inference runtime. Code and project page are available online.
Figures & tables
Fig. 1: We propose GaussianCaR , a novel method for efficient camera-radar fusion. We envision sensor fusion as a modality → Gaussians → BEV transformation, achieving competitive accuracy with significantly fast inference times for BEV segmentation tasks.
Fig. 2: Main diagram of our proposal, GaussianCaR. Given multi-view camera images and radar point clouds , we leverage GS as a universal view transformer and formulate sensor fusion as modality → Gaussians → BEV transformation. The model predicts BEV segmentation maps for dynamic vehicles and map elements. We employ two feature encoding branches: Pixels-to-Gaussians for camera features and Points-to-Gaussians for radar point clouds. Features are splatted and fused in BEV space using a CMX-based fuser, and decoded via a DPT decoder.
Fig. 3: Gaussian modeling process. In (a), we present the process of extracting a Gaussian from a discrete probability distribution; in (b), we depict the behavior of the offset head, displacing the final Gaussian position from the original set of candidates; in (c), we illustrate the Gaussian rasterization process, projecting Gaussians from 3D space to BEV space via orthographic projection.
Fig. 4: Our Pixels-to-Gaussians extracts low-resolution feature maps using an EfficientViT backbone and a neck. A set of convolutional heads predicts Gc Gaussians. To position the Gaussians in 3D space, camera intrinsic and extrinsic matrices are used.
Fig. 5: Our proposed Points-to-Gaussians module processes radar point clouds using a lightweight PTv3, composed of E encoder and D decoder blocks. A set of MLP heads then predicts Gr Gaussians, each parameterized by geometric and semantic attributes.
Method
Code
Cam Enc
Radar Enc
IoU ( ↑ )
Camera-only
BEVFormer [ 5 ]
✓
RN-101
-
43.2
GaussianLSS [ 22 ]
✓
RN-101
-
46.1
SimpleBEV [ 6 ]
✓
RN-101
-
47.4
PointBeV [ 29 ]
✓
EN-b4
-
47.8
GaussianBeV [ 23 ]
✗
EN-b4
-
50.3
TABLE I: BEV Vehicle Segmentation on the nuScenes Validation Set
Method
Driv. Area ( ↑ )
Lane Div. ( ↑ )
Camera-only
LSS [ 3 ]
72.9
20.0
BEVFormer [ 5 ]
80.1
25.7
GaussianBeV [ 23 ]
82.6
47.4
Camera-radar
BEVGuide [ 30 ]
76.7
44.2
TABLE II: BEV Map Segmentation on the nuScenes Validation Set
Method
IoU ( ↑ )
ms ( ↓ )
FPS ( ↑ )
Baseline: GaussianLSS [ 22 ]
46.1
53.9
18.6
Image Encoding Branch
+ EffViT L2
47.3
56.6
17.8
+ Offset Head
47.7
56.9
17.6
+ Early auxiliary loss
47.8
56.9
17.6
+ Dice loss
48.0
56.9
17.6
TABLE III: Ablation Study
Fig. 6: Qualitative results on the nuScenes validation set. Each row shows, from left to right: multi-view camera images, PCA camera latent features, PCA radar latent features, and predictions. For vehicle segmentation, we report an error map where correctness is indicated by color: correct , missing , and incorrect . For map segmentation, we report classes by color: drivable area , lane and road dividers , pedestrian crossings , walkway and carpark areas .