GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
Authors: Kerui Ren, Tao Lu, Linning Xu, Changjian Jiang, Mu Huang, Chunhua Shen, Mulin Yu, Bo Dai
Organizations: Shanghai Jiao Tong University · Shanghai Artificial Intelligence Laboratory · The Chinese University of Hong Kong · The University of Hong Kong · Fudan University · Zhejiang University
Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas video generative models offer rich appearance priors but accumulate inconsistencies during sequential view generation. We propose GeoVerse, a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model. Specifically, GeoVerse extracts multilevel features from Wan2.2 VACE and injects them into the geometric latent diffusion model via a ControlNet-style adapter, incorporating video-learned appearance priors to enhance structural completion. To enforce cross-view coherence, a global spatial memory continuously aggregates observed and synthesized content, reprojecting target-aligned guidance to anchor subsequent predictions to a shared scene representation. Extensive experiments across diverse datasets demonstrate improved visual quality and geometric consistency, with a 2.23 dB higher PSNR on DL3DV and 32.4% lower ATE on Mip-NeRF360 compared to GLD.
Figures & tables
Figure 1: GeoVerse enables long-sequence novel view synthesis by iteratively integrating generated observations into a persistent spatial memory, which in turn guides subsequent view synthesis. Project page: https://geoverse-nvs.github.io/ .
Figure 2: Overview of GeoVerse. Context observations initialize a global spatial memory that provides target-aligned RGB-D hints. Guided by the hints and Wan2.2 VACE features injected through a ControlNet-style adapter, geometric latent diffusion synthesizes target views in four denoising updates. The decoded RGB and geometry update the memory for subsequent view expansion.
Method
2D Metrics
3D Metrics
Efficiency
PSNR ↑
SSIM ↑
LPIPS ↓
ATE ↓
RPE r↓
RPE t↓
Reproj. ↓
MEt3R ↓
Time (s) ↓
DL3DV
ViewCrafter
15.98
0.516
0.494
0.367
10.214
0.728
0.738
0.313
286.75
NeoVerse
12.14
0.353
0.622
0.164
7.934
0.382
0.638
0.278
249.19
GEN3C
17.06
0.561
0.435
0.114
5.788
0.238
0.692
0.301
343.10
MVGenMaster
17.28
0.573
0.384
0.098
5.235
0.205
0.669
0.282
37.64
Matrix3D
13.47
0.416
0.494
0.153
6.127
0.323
0.715
0.288
53.39
Table 1: Quantitative comparison across multiple datasets in terms of visual quality, geometric accuracy, and average inference time. Bold and underline denote the best and second-best results.
Figure 3: Qualitative comparisons across multiple datasets. Top: comparisons of results from a single inference pass. Bottom: comparisons of long-sequence generation over three rounds.
Steps
L1 sweep (Cascade L0 fixed to 1)
Cascade L0 sweep (L1 fixed to 3)
Time (s) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
ATE ↓
MEt3R ↓
Time (s) ↓
PSNR ↑
SSIM ↑
LPIPS ↓
ATE ↓
MEt3R ↓
49
80.04
17.72
0.535
0.376
0.037
0.290
52.55
19.48
0.597
0.342
0.025
0.260
24
41.33
17.87
0.541
0.370
0.035
0.286
29.88
19.50
0.598
0.342
0.027
0.261
12
23.02
18.15
0.550
0.362
0.033
0.282
19.29
19.53
0.599
0.342
0.029
0.261
6
13.84
18.61
0.567
0.349
0.032
0.272
13.75
19.60
0.600
0.342
0.031
0.271
3
9.24
19.61
0.598
0.339
0.028
0.261
11.02
19.66
0.601
0.341
0.029
0.261
Table 2: Quantitative comparisons on DL3DV ( Ling et al., 2024 ) with different denoising budgets. Three L1 updates and one Cascade L0 update provide the best overall trade-off between synthesis quality and inference speed among the compared configurations.
Figure 4: Qualitative comparisons on DL3DV ( Ling et al., 2024 ) with different denoising budgets. One L1 update produces blurrier fine details than three L1 updates, while one Cascade L0 update preserves detail comparable to 49 Cascade L0 updates.
Variant
2D Metrics
3D Metrics
Scaled time
PSNR ↑
SSIM ↑
LPIPS ↓
ATE ↓
RPE r↓
RPE t↓
Reproj. ↓
MEt3R ↓
Reference (s) ↓
GeoVerse
19.65
0.781
0.398
0.009
0.223
0.016
0.589
0.088
33.11
w/o Wan2.2 prior
16.83
0.696
0.463
0.022
0.388
0.034
0.591
0.092
29.19
w/o spatial memory
17.87
0.727
0.434
0.014
0.268
0.023
0.591
0.087
33.48
MoT video-KV
19.24
0.769
0.403
0.009
0.238
0.016
0.587
0.089
36.29
w/o RGB loss
18.58
0.756
0.415
0.014
0.360
0.023
0.581
0.091
33.21
Table 3: Ablation study on ScanNetV2 ( Dai et al., 2017 ) . We evaluate the contributions of key components, training losses, and alternative architectures for video-prior injection.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Data type
#Scenes
#Images
Aria Digital Twin ( Pan et al., 2023 )
Real / digital twin
188
232,115
ARKitScenes ( Baruch et al., 2021 )
Real / RGB-D
644
125,134
DL3DV ( Ling et al., 2024 )
Real / multi-view video
10,475
3,594,809
MapFree ( Arnold et al., 2022 )
Real / multi-view video
460
515,113
ScanNet++ v2 ( Yeshwanth et al., 2023 )
Real / RGB + scans
954
1,032,198
Waymo ( Sun et al., 2020 )
Real / driving
3,990
790,405
Appendix
Table 4: Training datasets used by GeoVerse. We summarize the data types, scene counts, and image counts of the real and synthetic multi-view datasets in our training mixture.
Figure 5: Additional qualitative comparisons on RE10K and DL3DV. Two target views are shown for each of three scenes, alongside context images and ground truth.
Figure 6: Additional long-sequence comparisons on ScanNetv2. Each row contains two context images followed by target frames 1, 3, and 5 from each of three rounds.