M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis
Authors: Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, Giuseppe Loianno
Organizations: New York University, New York, USA · Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720, USA · U.S. Army Combat Capabilities Development Command, Army Research Laboratory, Adelphi, MD 20783, USA
Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator. We show that, after camera projection, frozen LiDAR and image features exhibit substantial shared spatial structure, providing a natural cross-modal representation. M3GD conditions generation on LiDAR through this structure: it combines explicit geometry statistics with learned point-cloud descriptors into view-aligned packets on the image-latent grid, injected through a lightweight residual adapter into a multi-view flow-matching generator whose latent space, decoders, and training objective remain intact. On the GrandTour dataset, M3GD improves target-view RGB and depth synthesis over an image-only version of the same backbone. Ablations show that the gains come from pixel-aligned LiDAR content and that target-view LiDAR acts as a geometric query linking the requested view to source observations. Deployment on a ground robot demonstrates practical real-world operation, with a configurable quality--cost trade-off controlled by the number of Euler integration steps.
Figures & tables
Fig. 1: Projected RGB–LiDAR representation. Frozen Utonia per-point descriptors are projected and pooled into the same H×W image-latent grid produced by the frozen DA3 encoder, so each cell q indexes both zq and hq (first three PCA components; hatched cells have no LiDAR support, nq=0 ).
Fig. 3: Clearpath Warthog platform used for the real-robot deployment.
Feature pair
PCA ch. corr. ↑
CKA ↑
SVCCA- 20↑
Jaccard@ 10↑
Affinity corr. ↑
DA3-Base feature level vs. Utonia
DA3-Base level- 0 –Utonia
0.6412
0.6590
0.6040
0.3551
0.5737
DA3-Base level- 1 –Utonia
0.6212
0.6451
0.5911
0.3402
0.5997
DA3-Base level- 2 –Utonia
0.6347
0.5459
0.5789
0.3460
0.5372
DA3-Base level- 3 –Utonia
0.6426
0.4688
0.5796
0.3565
0.4580
Image backbone comparison vs. Utonia
TABLE I: Feature-structure similarity on GrandTour. All metrics are computed on shared 16×24 image-grid cells with valid projected LiDAR support. The study covers 599,774 successfully processed images from 48 recordings and three HDR camera streams. Best and second-best are highlighted within each comparison block.
Fig. 4: Qualitative feature-structure visualization on GrandTour. For each camera image, we show the RGB input, the DA3-Base level- 0 PCA pseudo-color map, projected Utonia features, and DINOv2 final-layer image features. For each image, DA3-Base, Utonia, and DINOv2 features are independently projected to three local PCA components on the shared valid grid cells. The Utonia and DINOv2 coordinates are aligned to DA3-Base level- 0 using orthogonal Procrustes. Black regions in the Utonia column indicate cells without valid projected LiDAR support. Projected Utonia features recover large-scale geometric partitions similar to DA3-Base level- 0 on LiDAR-supported cells, while DINOv2 provides an image-only reference.
2D Appearance
2D Geometry (Depth)
3D Reconstruction
Method
PSNR ↑
SSIM ↑
LPIPS ↓
AbsRel ↓
RMSE ↓
δ<1.25↑
MEt3R ↓
Reproj. ↓
Diffusion-based methods
Matrix3D [ 73 ]
16.90
0.374
0.502
0.282
0.570
0.587
0.184
0.071
MVGenMaster [ 74 ]
17.45
0.406
0.446
0.285
0.568
0.629
0.243
0.072
GLD [ 8 ]
17.92
0.414
0.450
0.303
0.602
0.592
0.283
0.064
GLD ∗† [ 8 ]
21.62
0.531
0.302
0.173
0.354
0.777
0.191
0.059
TABLE II: Baseline comparison on the GrandTour evaluation split ( 2,400 clips). ↑ / ↓ indicate that higher/lower is better; best and second-best per column are highlighted. ∗ fine-tuned on the GrandTour training set; † trained on the DA3 level- 0 feature.
Method
AbsRel ↓
RMSE ↓
δ<1.25↑
Native readout (GLD-default)
GLD
0.358
0.747
0.478
GLD ∗†
0.248
0.500
0.565
M3GD ∗† (Ours)
0.227
0.472
0.579
RGB readout (protocol of Table II )
GLD
0.303
0.602
0.592
TABLE III: Depth readout comparison, using identical latent predictions as Table II . The native readout scores depth decoded directly from the synthesized latent (GLD-default); the RGB readout first decodes to RGB and re-encodes through the DA3 encoder (protocol of Table II ). Both are scored against the same pseudo-ground truth. Notation follows Table II ; best within each readout.
Method
Latency (s) ↓
Memory (GiB) ↓
Matrix3D [ 73 ]
77.7±0.1
11.0
MVGenMaster [ 74 ]
10.3±0.1
10.5
GLD [ 8 ]
17.3±0.8
9.7
M3GD ∗† (Ours)
10.2±0.2
6.5
TABLE IV: Sampling cost of the diffusion methods on the GrandTour evaluation split. All methods use N=2 source views and M=2 target views. Latency is measured per target pair over the full sampling procedure, including the VAE decode to RGB and, for GLD, both stages of its native cascade; for M3GD it also includes generation of the LiDAR packet. Values are averaged over three runs ( ± standard deviation); peak reserved GPU memory is measured on a single NVIDIA RTX A5000, following the protocol of Table VII . Best per column.
Fig. 5: Qualitative RGB and depth comparison on the GrandTour evaluation split. For each scene, the two source images are shown on the left. The middle block compares RGB predictions, and is organized into three rows: feed-forward splatting baselines DepthSplat [ 16 ] , MVSplat [ 15 ] , and PixelSplat [ 14 ] ; diffusion-based baselines Matrix3D [ 73 ] and MVGenMaster [ 74 ] together with the original GLD model [ 8 ] ; and GLD ∗† , our GrandTour-fine-tuned image-only GLD baseline, followed by M3GD and the ground-truth target view. The right block shows the corresponding DA3 depth estimates for the same predicted views and the ground-truth target depth, using the same method layout. Compared with splatting and image-only generative baselines, M3GD better preserves target-view appearance and produces depth maps that more closely follow the ground-truth scene geometry in challenging robot-captured scenes.
Fig. 6: Qualitative visualization of M3GD on GrandTour evaluation samples. Each group follows the actual evaluation layout: two observed source views, two M3GD target-view predictions, and the corresponding ground-truth target views. M3GD synthesizes targets from sparse source images and target-view LiDAR geometry; target RGB is never provided. Examples are ordered from indoor scenes to semi-open structures and outdoor environments, showing that M3GD preserves dominant scene layout and reconstructs plausible target-view appearance across diverse and geometrically challenging robot-captured environments.
Training mode
LiDAR
PSNR ↑
SSIM ↑
LPIPS ↓
freeze
✗ ‡
17.39
0.380
0.468
✓ §
19.11
0.434
0.394
scratch
✗
21.53
0.520
0.311
✓
22.16
0.534
0.295
fine-tune
✗
21.56
0.525
0.303
✓
22.22
0.542
0.285
TABLE V: Effect of LiDAR conditioning across training modes (multi-step, level- 0 latent). Best per column. ‡ Off-the-shelf: frozen released GLD checkpoint, no LiDAR adapter. § As ‡ , with the LiDAR adapter trained. tgt-only scores the loss on the target-view latents only, not on the jointly denoised source views. ¶ Classifier-free LiDAR (cfg-dropout 0.1 ); at evaluation a single CFG scale of 1.5 is applied to the joint camera+LiDAR condition.
LiDAR input mode
PSNR ↑
SSIM ↑
LPIPS ↓
Full (real packet)
21.00
0.476
0.314
Geom-only
20.90
0.471
0.319
Zero
20.37
0.455
0.336
Learnable param
20.37
0.455
0.336
No LiDAR (floor)
20.31
0.455
0.336
TABLE VI: Effect of LiDAR input mode on the hard subset (maximum camera baseline between views >1 m, n=1,407 ; Section V-B ). Each row is a separately trained model under the default recipe of Section VI-D1 ; only the LiDAR input changes. Highlighting follows Table II .
Fig. 7: Qualitative examples highlighting the effect of target-view LiDAR conditioning. Each row shows two source images, the image-only GLD ∗† baseline, the raw target-view LiDAR projection, our downsampled LiDAR representation, the geometry-only (geom-only) variant, M3GD, and the ground-truth target image. These examples contain large source-target viewpoint changes where RGB-only conditioning is ambiguous and GLD ∗† often fails to infer the target-facing surface. The LiDAR-conditioned geom-only model recovers the dominant scene layout more accurately, while M3GD further improves appearance fidelity and local details such as boxes, road surfaces, walls, and vehicles by combining LiDAR geometry with Utonia features.
Fig. 8: Qualitative results for two real-robot queries Q1 (rows 1 – 2 ) and Q2 (rows 3 – 4 ), using K=10 Euler steps, N=2 source views, and M=2 target views. Each query forms a two-row block: in the first two columns, the upper row shows the two source RGB images and the lower row shows their corresponding projected LiDAR point clouds. The remaining columns show the ground-truth RGB, M3GD synthesized RGB, our RGB readout depth, GLD synthesized RGB, and GLD RGB readout depth (defined in Section VI-B ) for target T1 in the upper row and target T2 in the lower row.
Configuration
M3GD Quality
M3GD Computational Cost
Steps K
Targets M
PSNR (dB) ↑
SSIM ↑
LPIPS ↓
Latency (s) ↓
Lat./Target (s) ↓
Mem. (GiB) ↓
1
2
14.13
0.279
0.889
3.13±0.44
1.56±0.22
4.22
5
2
15.86
0.281
0.480
3.63±0.45
1.81±0.22
5.70
5
3
15.65
0.280
0.489
3.77±0.45
1.26±0.15
5.96
5
4
15.54
0.276
0.494
3.91±0.44
0.98±0.11
6.11
10
2
15.84
0.270
0.476
4.26±0.44
2.13±0.22
5.80
TABLE VII: On-robot quality, latency, and memory scaling on the Warthog platform with two fixed source views ( N=2 ). Quality is averaged over 30 queries; latency and peak GPU reserved memory are measured over four synchronized post-warm-up queries.