GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
Authors: Kerui Ren, Tao Lu, Linning Xu, Changjian Jiang, Mu Huang, Chunhua Shen, Mulin Yu, Bo Dai
Organizations: Shanghai Jiao Tong University · Shanghai Artificial Intelligence Laboratory · The Chinese University of Hong Kong · The University of Hong Kong · Fudan University · Zhejiang University
Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas video generative models offer rich appearance priors but accumulate inconsistencies during sequential view generation. We propose GeoVerse, a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model. Specifically, GeoVerse extracts multilevel features from Wan2.2 VACE and injects them into the geometric latent diffusion model via a ControlNet-style adapter, incorporating video-learned appearance priors to enhance structural completion. To enforce cross-view coherence, a global spatial memory continuously aggregates observed and synthesized content, reprojecting target-aligned guidance to anchor subsequent predictions to a shared scene representation. Extensive experiments across diverse datasets demonstrate improved visual quality and geometric consistency, with a 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360 compared to GLD.
Figures & tables
Figure 1: GeoVerse enables long-sequence novel view synthesis by iteratively integrating generated observations into a persistent spatial memory, which in turn guides subsequent view synthesis. Project page: https://geoverse-nvs.github.io/ .
Figure 2: Overview of GeoVerse. Context observations initialize a global spatial memory that provides target-aligned RGB-D hints. Guided by the hints and Wan2.2 VACE features injected through a ControlNet-style adapter, geometric latent diffusion synthesizes target views in four denoising updates. The decoded RGB and geometry update the memory for subsequent view expansion.
Method
2D Metrics
3D Metrics
Efficiency
PSNR ↑
SSIM ↑
LPIPS ↓
ATE ↓
RPE r↓
RPE t↓
Reproj. ↓
MEt3R ↓
Time (s) ↓
DL3DV
ViewCrafter
15.98
0.516
0.494
0.367
10.214
0.728
0.738
0.313
286.75
NeoVerse
12.14
0.353
0.622
0.164
7.934
0.382
0.638
0.278
249.19
GEN3C
17.06
0.561
0.435
0.114
5.788
0.238
0.692
0.301
343.10
MVGenMaster
17.28
0.573
0.384
0.098
5.235
0.205
0.669
0.282
37.64
Matrix3D
13.47
0.416
0.494
0.153
6.127
0.323
0.715
0.288
53.39
Table 1: Quantitative comparison across multiple datasets in terms of visual quality, geometric accuracy, and average inference time. Bold and underline denote the best and second-best results.
Figure 3: Qualitative comparisons across multiple datasets. Top: comparisons of results from a single inference pass. Bottom: comparisons of long-sequence generation over three rounds.
Steps
L1 sweep (Cascade L0 fixed to 1)
Cascade L0 sweep (L1 fixed to 3)
Time (s) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
ATE ↓
MEt3R ↓
Time (s) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
ATE ↓
MEt3R ↓
49
80.04
17.72
0.535
0.376
0.037
0.290
52.55
19.48
0.597
0.342
0.025
0.260
24
41.33
17.87
0.541
0.370
0.035
0.286
29.88
19.50
0.598
0.342
0.027
0.261
12
23.02
18.15
0.550
0.362
0.033
0.282
19.29
19.53
0.599
0.342
0.029
0.261
6
13.84
18.61
0.567
0.349
0.032
0.272
13.75
19.60
0.600
0.342
0.031
0.271
3
9.24
19.61
0.598
0.339
0.028
0.261
11.02
19.66
0.601
0.341
0.029
0.261
Table 2: Quantitative comparisons on DL3DV ( Ling et al., 2024 ) with different denoising budgets. Three L1 updates and one Cascade L0 update provide the best overall trade-off between synthesis quality and inference speed among the compared configurations.
Figure 4: Qualitative comparisons on DL3DV ( Ling et al., 2024 ) with different denoising budgets. One L1 update produces blurrier fine details than three L1 updates, while one Cascade L0 update preserves detail comparable to 49 Cascade L0 updates.
Variant
2D Metrics
3D Metrics
Scaled time
PSNR ↑
SSIM ↑
LPIPS ↓
ATE ↓
RPE r↓
RPE t↓
Reproj. ↓
MEt3R ↓
Reference (s) ↓
GeoVerse
19.65
0.781
0.398
0.009
0.223
0.016
0.589
0.088
33.11
w/o Wan2.2 prior
16.83
0.696
0.463
0.022
0.388
0.034
0.591
0.092
29.19
w/o spatial memory
17.87
0.727
0.434
0.014
0.268
0.023
0.591
0.087
33.48
MoT video-KV
19.24
0.769
0.403
0.009
0.238
0.016
0.587
0.089
36.29
w/o RGB loss
18.58
0.756
0.415
0.014
0.360
0.023
0.581
0.091
33.21
Table 3: Ablation study on ScanNetV2 ( Dai et al., 2017 ) . We evaluate the contributions of key components, training losses, and alternative architectures for video-prior injection.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Data type
#Scenes
#Images
Aria Digital Twin ( Pan et al., 2023 )
Real / digital twin
188
232,115
ARKitScenes ( Baruch et al., 2021 )
Real / RGB-D
644
125,134
DL3DV ( Ling et al., 2024 )
Real / multi-view video
10,475
3,594,809
MapFree ( Arnold et al., 2022 )
Real / multi-view video
460
515,113
ScanNet++ v2 ( Yeshwanth et al., 2023 )
Real / RGB + scans
954
1,032,198
Waymo ( Sun et al., 2020 )
Real / driving
3,990
790,405
Appendix
Table 4: Training datasets used by GeoVerse. We summarize the data types, scene counts, and image counts of the real and synthetic multi-view datasets in our training mixture.
Figure 5: Additional qualitative comparisons on RE10K and DL3DV. Two target views are shown for each of three scenes, alongside context images and ground truth.
Figure 6: Additional long-sequence comparisons on ScanNetv2. Each row contains two context images followed by target frames 1, 3, and 5 from each of three rounds.
Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a framework for geometry-enhanced novel view synthesis via a disentangled 3D prior. Specifically, GenNVS models foreground objects and the background with 3D Gaussian Splatting and aligns them through a coarse-to-fine geometric optimization process to form a unified 3D scene. This scene conditions a video diffusion model through the proposed Dual-Stream Masking mechanism, which guides synthesis by jointly exploiting rendered validity masks and geometry-aware warping. Experimental results show that GenNVS performs favorably against recent methods in both visual quality and geometric accuracy, while naturally supporting flexible scene editing.
Novel view synthesis requires strong 3D geometric consistency and the ability to generate visually coherent images across diverse viewpoints. While recent camera-controlled video diffusion models show promising results, they often suffer from geometric distortions and limited camera controllability. To overcome these challenges, we introduce GeoNVS, a geometry-grounded novel-view synthesizer that enhances both geometric fidelity and camera controllability through explicit 3D geometric guidance. Our key innovation is the Gaussian Splat Feature Adapter (GS-Adapter), which lifts input-view diffusion features into 3D Gaussian representations, renders geometry-constrained novel-view features, and adaptively fuses them with diffusion features to correct geometrically inconsistent representations. Unlike prior methods that inject geometry at the input level, GS-Adapter operates in feature space, avoiding view-dependent color noise that degrades structural consistency. Its plug-and-play design enables zero-shot compatibility with diverse feed-forward geometry models without additional training, and can be adapted to other video diffusion backbones. Experiments across 9 scenes and 18 settings demonstrate state-of-the-art performance, achieving 11.3% and 14.9% improvements over SEVA and CameraCtrl, with up to 2x reduction in translation error and 7x in Chamfer Distance.
Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this paper, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization. Our code is publicly available at https://github.com/KlingAIResearch/MetaView.
Yufei Cai, Xuesong Niu, Hao Lu +3
Nanyang Technological University · Kolors Team, Kuaishou Technology · The Hong Kong University of Science and Technology (Guangzhou)