Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a framework for geometry-enhanced novel view synthesis via a disentangled 3D prior. Specifically, GenNVS models foreground objects and the background with 3D Gaussian Splatting and aligns them through a coarse-to-fine geometric optimization process to form a unified 3D scene. This scene conditions a video diffusion model through the proposed Dual-Stream Masking mechanism, which guides synthesis by jointly exploiting rendered validity masks and geometry-aware warping. Experimental results show that GenNVS performs favorably against recent methods in both visual quality and geometric accuracy, while naturally supporting flexible scene editing.
Figures & tables
Figure 1: Geometric prior from entangled scene modeling [ 32 ] and our disentangled formulation.
Figure 2: Overview of GenNVS. Given a single input image, our framework first generates separate 3DGS for foreground objects and background. These are aligned into a unified 3D scene through coarse-to-fine geometric optimization. The aligned prior then conditions a video diffusion model via a Dual-Stream Masking mechanism, yielding high-fidelity and geometrically consistent novel view sequences.
Figure 3: Training pipeline with Dual-Stream Masking. A coarse 3DGS reconstructed from the video is rendered into a warped image sequence and a validity mask. During training, the mask latent is concatenated with both the noisy video latent and the warped latent along the channel dimension, and the two streams are fused along the temporal dimension for denoising.
Figure 4: Qualitative comparisons on the DL3DV and Mip-NeRF 360 datasets. We evaluate novel view sequences generated by our method against VMem [ 16 ] , GenWarp [ 25 ] , and HW-Mirror [ 20 ] . GenNVS produces more geometrically consistent and temporally stable results. Red circles highlight geometric inconsistencies, and blue circles indicate unseen holes.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Rdist↓
Tdist↓
HW-Mirror [ 20 ]
17.18
0.41
0.28
20.81
1.607
0.467
GenWarp [ 25 ]
14.38
0.18
0.47
28.68
1.811
0.464
VMem [ 16 ]
16.32
0.30
0.34
19.17
1.888
0.460
RecamMaster [ 1 ]
17.60
0.44
0.33
18.82
1.805
0.465
ViewCrafter [ 32 ]
18.72
0.50
0.26
16.38
1.320
0.472
CameraCtrl [ 10 ]
16.74
0.43
0.32
22.10
1.680
0.465
Table 1: Quantitative comparison on DL3DV dataset. Single-view novel view synthesis results on DL3DV dataset with diverse environments and complex camera trajectories. Our method demonstrates favorable geometric consistency and image quality.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Rdist↓
Tdist↓
HW-Mirror [ 20 ]
15.90
0.20
0.35
16.25
1.098
0.420
GenWarp [ 25 ]
14.59
0.33
0.47
29.33
1.184
0.629
VMem [ 16 ]
19.04
0.47
0.30
16.81
2.150
0.481
RecamMaster [ 1 ]
17.86
0.48
0.34
15.16
1.881
0.432
ViewCrafter [ 32 ]
16.75
0.48
0.39
19.23
1.067
0.595
CameraCtrl [ 10 ]
16.74
0.43
0.30
22.10
1.680
0.465
Table 2: Quantitative comparison on Mip-NeRF 360 dataset. Evaluation on Mip-NeRF 360 dataset further validates the robustness of our approach in handling complex indoor and outdoor scenes. GenNVS obtains favorable results against baselines in both perceptual quality and geometric accuracy.
Figure 5: Ablation study on alignment stages. Qualitative comparisons demonstrate the necessity of both the coarse and fine alignment stages.
PSNR ↑
LPIPS ↓
Time/obj (s)
w/o disent.
15.90
0.35
w/o coarse align.
18.39
0.22
177.45
w/o fine align.
19.65
0.18
1.32
GenNVS
21.75
0.16
22.88
Table 3: Ablation studies on Mip-NeRF 360 dataset. We evaluate the contribution of each component. The full model achieves the best quality while maintaining a trade-off in efficiency.
Disent. gen.
Coarse (/obj)
Fine (/obj)
Diff.
Total
Time (s)
20.15
1.28
21.60
60.10
103.13
Table 4: Analysis of inference time. The table details the processing time per stage, including disentangled generation, coarse and fine alignment (per object), diffusion-based rendering, and the total end-to-end latency.
Figure 6: User study results. We evaluate our method along with VMem [ 16 ] , GenWarp [ 25 ] , HW-Mirror [ 20 ] on O.A., O.C., F.A., and F.C.. Our method performs favorably against baselines, receiving over 40% of votes as the preferred approach in each category.
Figure 7: Scene editing results. Replacing the foreground "lego" with a "tank" preserves visual and geometric consistency across synthesized novel views.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Comparison of 3D prior from a single image. Top row: initial point cloud from a monocular model. Bottom row: refined results after disentangled foreground-background generation and coarse-to-fine alignment.
Metric
Default
Lrgb↑30%
Ldepth↑30%
Lmask↑30%
Lhole↑30%
Loverflow↑30%
PSNR
24.03
24.18
23.82
23.65
21.47
23.95
SSIM
0.83
0.84
0.81
0.80
0.74
0.82
LPIPS
0.18
0.19
0.16
0.21
0.28
0.18
Appendix
Table 5: Sensitivity analysis of loss weights in the prior generation stage. Best values are highlighted in bold.
Figure 10: Additional results on the DL3DV [ 17 ] dataset.
Figure 11: Additional comparisons with CameraCtrl [ 10 ] , MotionCtrl [ 31 ] , and SEVA [ 34 ] . Our method shows stronger geometric consistency.
Figure 12: Generalization to Diverse Artistic Styles. We evaluate our framework on out-of-distribution artistic inputs, including hand-drawn sketches, game scenes, and anime imagery. Results show that our method synthesizes high-quality, stylized novel views while maintaining strict multi-view 3D consistency, highlighting that our disentangled prior and conditioning mechanism provide strong geometric guidance without sacrificing adaptation to varied visual styles.
Figure 13: Results under boundary viewpoint changes.