Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a framework for geometry-enhanced novel view synthesis via a disentangled 3D prior. Specifically, GenNVS models foreground objects and the background with 3D Gaussian Splatting and aligns them through a coarse-to-fine geometric optimization process to form a unified 3D scene. This scene conditions a video diffusion model through the proposed Dual-Stream Masking mechanism, which guides synthesis by jointly exploiting rendered validity masks and geometry-aware warping. Experimental results show that GenNVS performs favorably against recent methods in both visual quality and geometric accuracy, while naturally supporting flexible scene editing.
Figures & tables
Figure 1: Geometric prior from entangled scene modeling [ 32 ] and our disentangled formulation.
Figure 2: Overview of GenNVS. Given a single input image, our framework first generates separate 3DGS for foreground objects and background. These are aligned into a unified 3D scene through coarse-to-fine geometric optimization. The aligned prior then conditions a video diffusion model via a Dual-Stream Masking mechanism, yielding high-fidelity and geometrically consistent novel view sequences.
Figure 3: Training pipeline with Dual-Stream Masking. A coarse 3DGS reconstructed from the video is rendered into a warped image sequence and a validity mask. During training, the mask latent is concatenated with both the noisy video latent and the warped latent along the channel dimension, and the two streams are fused along the temporal dimension for denoising.
Figure 4: Qualitative comparisons on the DL3DV and Mip-NeRF 360 datasets. We evaluate novel view sequences generated by our method against VMem [ 16 ] , GenWarp [ 25 ] , and HW-Mirror [ 20 ] . GenNVS produces more geometrically consistent and temporally stable results. Red circles highlight geometric inconsistencies, and blue circles indicate unseen holes.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Rdist↓
Tdist↓
HW-Mirror [ 20 ]
17.18
0.41
0.28
20.81
1.607
0.467
GenWarp [ 25 ]
14.38
0.18
0.47
28.68
1.811
0.464
VMem [ 16 ]
16.32
0.30
0.34
19.17
1.888
0.460
RecamMaster [ 1 ]
17.60
0.44
0.33
18.82
1.805
0.465
ViewCrafter [ 32 ]
18.72
0.50
0.26
16.38
1.320
0.472
CameraCtrl [ 10 ]
16.74
0.43
0.32
22.10
1.680
0.465
Table 1: Quantitative comparison on DL3DV dataset. Single-view novel view synthesis results on DL3DV dataset with diverse environments and complex camera trajectories. Our method demonstrates favorable geometric consistency and image quality.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Rdist↓
Tdist↓
HW-Mirror [ 20 ]
15.90
0.20
0.35
16.25
1.098
0.420
GenWarp [ 25 ]
14.59
0.33
0.47
29.33
1.184
0.629
VMem [ 16 ]
19.04
0.47
0.30
16.81
2.150
0.481
RecamMaster [ 1 ]
17.86
0.48
0.34
15.16
1.881
0.432
ViewCrafter [ 32 ]
16.75
0.48
0.39
19.23
1.067
0.595
CameraCtrl [ 10 ]
16.74
0.43
0.30
22.10
1.680
0.465
Table 2: Quantitative comparison on Mip-NeRF 360 dataset. Evaluation on Mip-NeRF 360 dataset further validates the robustness of our approach in handling complex indoor and outdoor scenes. GenNVS obtains favorable results against baselines in both perceptual quality and geometric accuracy.
Figure 5: Ablation study on alignment stages. Qualitative comparisons demonstrate the necessity of both the coarse and fine alignment stages.
PSNR ↑
LPIPS ↓
Time/obj (s)
w/o disent.
15.90
0.35
w/o coarse align.
18.39
0.22
177.45
w/o fine align.
19.65
0.18
1.32
GenNVS
21.75
0.16
22.88
Table 3: Ablation studies on Mip-NeRF 360 dataset. We evaluate the contribution of each component. The full model achieves the best quality while maintaining a trade-off in efficiency.
Disent. gen.
Coarse (/obj)
Fine (/obj)
Diff.
Total
Time (s)
20.15
1.28
21.60
60.10
103.13
Table 4: Analysis of inference time. The table details the processing time per stage, including disentangled generation, coarse and fine alignment (per object), diffusion-based rendering, and the total end-to-end latency.
Figure 6: User study results. We evaluate our method along with VMem [ 16 ] , GenWarp [ 25 ] , HW-Mirror [ 20 ] on O.A., O.C., F.A., and F.C.. Our method performs favorably against baselines, receiving over 40% of votes as the preferred approach in each category.
Figure 7: Scene editing results. Replacing the foreground "lego" with a "tank" preserves visual and geometric consistency across synthesized novel views.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Comparison of 3D prior from a single image. Top row: initial point cloud from a monocular model. Bottom row: refined results after disentangled foreground-background generation and coarse-to-fine alignment.
Metric
Default
Lrgb↑30%
Ldepth↑30%
Lmask↑30%
Lhole↑30%
Loverflow↑30%
PSNR
24.03
24.18
23.82
23.65
21.47
23.95
SSIM
0.83
0.84
0.81
0.80
0.74
0.82
LPIPS
0.18
0.19
0.16
0.21
0.28
0.18
Appendix
Table 5: Sensitivity analysis of loss weights in the prior generation stage. Best values are highlighted in bold.
Figure 10: Additional results on the DL3DV [ 17 ] dataset.
Figure 11: Additional comparisons with CameraCtrl [ 10 ] , MotionCtrl [ 31 ] , and SEVA [ 34 ] . Our method shows stronger geometric consistency.
Figure 12: Generalization to Diverse Artistic Styles. We evaluate our framework on out-of-distribution artistic inputs, including hand-drawn sketches, game scenes, and anime imagery. Results show that our method synthesizes high-quality, stylized novel views while maintaining strict multi-view 3D consistency, highlighting that our disentangled prior and conditioning mechanism provide strong geometric guidance without sacrificing adaptation to varied visual styles.
Figure 13: Results under boundary viewpoint changes.
Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, they often suffer from geometric distortions and limited camera controllability. To overcome these challenges, we introduce GeoNVS, a geometry-grounded novel-view synthesizer that enhances both geometric fidelity and camera controllability through explicit 3D geometric guidance. Our key innovation is the Gaussian Splat Feature Adapter (GS-Adapter), which lifts input-view diffusion features into 3D Gaussian representations, renders geometry-constrained novel-view features, and adaptively fuses them with diffusion features to correct geometrically inconsistent representations. Unlike prior methods that inject geometry at the input level, GS-Adapter operates in feature space, avoiding view-dependent color noise that degrades structural consistency. Its plug-and-play design enables zero-shot compatibility with diverse feed-forward geometry models without additional training, and can be adapted to other video diffusion backbones. Experiments across 9 scenes and 18 settings demonstrate state-of-the-art performance, achieving 11.3% and 14.9% improvements over SEVA and CameraCtrl, with up to 2x reduction in translation error and 7x in Chamfer Distance.
Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas video generative models offer rich appearance priors but accumulate inconsistencies during sequential view generation. We propose GeoVerse, a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model. Specifically, GeoVerse extracts multilevel features from Wan2.2 VACE and injects them into the geometric latent diffusion model via a ControlNet-style adapter, incorporating video-learned appearance priors to enhance structural completion. To enforce cross-view coherence, a global spatial memory continuously aggregates observed and synthesized content, reprojecting target-aligned guidance to anchor subsequent predictions to a shared scene representation. Extensive experiments across diverse datasets demonstrate improved visual quality and geometric consistency, with a 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360 compared to GLD.
Kerui Ren, Tao Lu, Linning Xu +5
Shanghai Jiao Tong University · Shanghai Artificial Intelligence Laboratory · The Chinese University of Hong Kong +3
Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this paper, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization. Our code is publicly available at https://github.com/KlingAIResearch/MetaView.
Yufei Cai, Xuesong Niu, Hao Lu +3
Nanyang Technological University · Kolors Team, Kuaishou Technology · The Hong Kong University of Science and Technology (Guangzhou)