Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training to a narrow set of purpose-built datasets. We propose in-the-wild monocular pretraining: a frozen depth estimator lifts each source image into 3D and reprojects under sampled poses to yield pseudo-target views; masked losses restrict supervision to valid regions and an adversarial objective covers disoccluded areas. Scaled to 30 million uncurated images, this produces OVIE, requiring only a source image and target pose at inference. Prior work trains without multi-view data but needs a depth estimator at inference, or drops this dependency but requires multi-view training pairs; OVIE is the first to require neither. Without multi-view supervision, OVIE rivals in-domain baselines on RealEstate10K and surpasses all on DL3DV, producing the most multi-view-consistent trajectories of any geometry-free method; brief multi-view fine-tuning outperforms all geometry-free methods on their training domain. At 116 FPS, it is over 600x faster than the fastest baseline. Code and pretrained models are at https://github.com/kyutai-labs/ovie; video results are on the project page, https://kyutai.org/blog/2026-04-14-ovie/.
Figures & tables
Figure 1 : OVIE generates novel views from a single image in real time across diverse domains given a source image (gray) and target poses (colored), regardless of content or style.
Figure 2 : Monocular pretraining method overview. Top: A frozen depth estimator lifts web-sourced images I0 into point clouds P ; sampled transformations T0→1∈SE(3) and reprojection yield pseudo-target views I1∗ . Bottom: Feed-forward model fθ takes I0 conditioned on T0→1 and predicts I^1 . Training uses a masked reconstruction loss Lrecon and a perceptual loss Lperc to compare I1∗ and I^1 , and an adversarial loss Ladv where discriminator model Dϕ distinguishes I0 from I^1 .
Source
Pseudo-Target
Prediction
Figure 3 : Pseudo-pairs and predictions.
Figure 4 : Qualitative comparison with baseline geometry-free methods. Pretrained on single images only, OVIE produces sharp, geometrically consistent novel views.
Figure 5 : Metric scale awareness. The same translation applied to a close-up scene (left) produces a larger apparent displacement than on a room-scale scene (right), consistent with metric parallax.
RealEstate10K [ 91 ]
DL3DV [ 38 ]
Method
OOD
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
MEt3R ↓
OOD
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
MEt3R ↓
GeoGPT [ 60 ]
✗
15.3
0.480
0.446
18.0
0.184
✓
13.1
0.339
0.560
35.9
0.300
ViewCrafter ∗ [ 84 ]
✗
16.6
0.559
0.340
9.98
0.193
✗
15.4
0.419
0.425
10.50
0.261
PhotoNVS [ 83 ]
✗
18.9
0.601
0.314
10.6
0.139
✓
13.8
0.349
0.525
37.6
0.301
VIVID [ 20 ]
✗
20.5
0.661
0.241
4.26
0.076
✓
14.5
0.362
0.471
18.0
0.169
OVIE (ours)
✓
18.8
0.602
0.279
6.74
0.035
✓
14.8
0.369
0.464
13.6
0.078
Table 1 : In-distribution and out-of-domain robustness on RealEstate10K and DL3DV. ↑ / ↓ : higher/lower is better. Bold : best; underline : second best. OOD: not trained on the benchmark. MEt3R [ 1 ] scores multi-view 3D consistency between consecutive generated views. ∗ Trained in-domain on DL3DV: scores greyed out and excluded from rankings for fairness.
Figure 6 : Quality vs. Inference speed tradeoff on RealEstate10K. Bubble radius indicates parameter count. OVIE-ft achieves competitive quality at drastically higher FPS than baselines.
RealEstate10K [ 91 ]
DL3DV [ 38 ]
P-DINO
LPIPS
GAN
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Loss component ablation
✓
✓
✓
18.9
0.596
0.284
7.12
15.0
0.373
0.468
14.3
✓
✓
19.0 +0.1
0.599 +.003
0.288 +.004
8.34 +1.22
15.1 +0.1
0.377 +.004
0.472 +.004
15.7 +1.4
✓
✓
18.7 − 0.2
0.592 − .004
0.297 +.013
8.43 +1.31
14.9 − 0.1
0.371 − .002
0.478 +.010
15.3 +1.0
✓
18.7 − 0.2
0.584 − .012
0.367 +.083
18.7 +11.6
14.9 − 0.1
0.368 − .005
0.540 +.072
27.0 +12.7
Table 2 : Loss ablation studies on RealEstate10K and DL3DV . Each group varies one design axis while keeping all others at the default configuration ( bold ).
Figure 7 : Scaling with dataset size. PSNR and FID improve as data volume increases.
RealEstate10K [ 91 ]
DL3DV [ 38 ]
Dataset
Domain
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Mix
Mixed
18.8
0.595
0.284
7.08
15.0
0.372
0.467
14.2
OSV5M [ 2 ]
Street View
18.2 − 0.6
0.566 − .029
0.318 +.034
8.59 +1.51
15.0
0.369 − .003
0.481 +.014
18.9 +4.7
ImageNet21K [ 59 ]
Objects
18.8
0.593 − .002
0.289 +.005
7.32 +0.24
14.9 − 0.1
0.370 − .003
0.472 +.005
15.9 +1.7
Places [ 89 ]
Scenes
18.8
0.596 +.001
0.286 +.002
7.41 +0.33
14.8 − 0.2
0.367 − .006
0.477 +.010
20.3 +6.1
OpenImages [ 33 ]
General
18.8
0.593 − .002
0.287 +.003
7.20 +0.12
14.9 − 0.1
0.370 − .003
0.469 +.002
14.8 +0.6
Table 3 : Comparison of data coverage on model performance on RealEstate10K and DL3DV. Bold : best; underline : second best. Differences are relative to the Mix baseline.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Scaling with dataset size – SSIM and LPIPS. Complementary to Figure 7 in the main paper, SSIM and LPIPS on RealEstate10K [ 91 ] follow the same monotonic improvement as data volume increases.
Figure 9 : Quality vs. Inference tradeoff on RealEstate10K – SSIM and FID. Complementary to Figure 6 in the main paper. Bubble radius indicates parameter count. The same trend holds: OVIE is faster and better-performing than baseline methods.
Method
Input
Native output
Scored at
OVIE
Single 256×256 image
256×256
256×256
SHARP [ 43 ]
Single 256×256 image (center-crop, below SHARP’s native 1536 )
3DGS render, arbitrary resolution
256×256
FlexWorld [ 12 ]
Single image at 16:9; video diffusion stage at 480×720
3DGS render, 288×512
256×256 center-crop
ViewCrafter [ 84 ]
Single image and its DUSt3R [ 76 ] reconstruction
576×1024
256×256 center-crop
VIVID [ 20 ]
Single image
256×256
256×256
PhotoNVS [ 83 ]
Single image
256×256
256×256
Appendix
Table 4 : Input and output resolution of each evaluated method. All metrics are computed at 256×256 .
Method
RealEstate10K [ 91 ]
DL3DV [ 38 ]
Geometry-free
GeoGPT [ 60 ]
0.184
0.300
PhotoNVS [ 83 ]
0.139
0.301
VIVID [ 20 ]
0.076
0.169
Geometry-based references
ViewCrafter [ 84 ]
0.193
0.261
Appendix
Table 5 : Multi-view 3D consistency on RealEstate10K and DL3DV. MEt3R [ 1 ] , lower is better, over 750 scenes per benchmark. Bold : best; underline : second best. Geometry-based methods estimate depth or a point cloud at inference; they are greyed out and excluded from the ranking for fairness. ∗ Trained in-domain on DL3DV: its DL3DV score is greyed out and excluded from the ranking, as in Table 1 .
Figure 10 : Generated trajectories on RealEstate10K. Five evaluation sequences, each shown at six points along the camera path. Every view is produced independently from the first frame of the sequence and its target pose, with no access to any intermediate ground-truth frame and no recurrence between views. Content stays stable along each path, including in regions disoccluded by the motion, which is the qualitative counterpart to the MEt3R scores in Section B.5 . These are the released model’s evaluation outputs, so they are the same generations the reported metrics are computed on.
Motion offset (frames)
0
3
15
45
90
Round-trip PSNR ↑
22.2
22.9
22.2
20.0
18.1
Co-visible fraction
0.99
0.94
0.80
0.58
0.41
Appendix
Table 6 : Round-trip consistency on RealEstate10K. A pose change followed by its reverse, scored on the pixels visible in both passes. The co-visible fraction is the share of the source still visible after the motion.
Figure 11 : Camera displacement distributions for the training mix and for both benchmarks, over all pairs. Boxes span the first and third quartiles around the median; whiskers mark the 10th and 90th percentiles. The training mix spans a wider range of both rotation and translation than either benchmark, so evaluation stays inside the training envelope.
Figure 12 : Quality as a function of co-visibility on RealEstate10K, where 1.0 denotes full source-target overlap and lower values mean larger motion and more disocclusion. 12(a) Restricting the measurement to co-visible pixels separates reconstruction of observed content from hallucination of disoccluded content: OVIE beats the warp it is trained to imitate below 0.9 and the warp wins in the near-copy regime, which holds most pairs, so pooled over all pairs the two essentially tie (19.5 against 19.7 dB). 12(b) Binning by co-visibility compares methods at matched motion: OVIE-ft leads in every bin and widens its margin as motion grows, while PhotoNVS collapses at the largest motions. Pooled over all pairs the methods reach 21.87 (OVIE-ft), 20.51 (VIVID), 18.83 (OVIE), 18.78 (PhotoNVS) and 15.25 dB (GeoGPT).
RealEstate10K [ 91 ]
DL3DV [ 38 ]
Depth estimator
Params
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
MoGe-2 [ 75 ] (metric, with mask)
331M
19.05
0.605
0.276
15.00
0.367
0.476
Depth Anything 3 [ 37 ] (relative, no mask)
135M
18.77 − 0.28
0.592 − .013
0.287 +.011
14.58 − 0.42
0.353 − .014
0.497 +.021
Appendix
Table 7 : Depth-estimator ablation. Models trained from scratch on RealEstate10K under an identical recipe, changing only the frozen depth network. Differences are relative to MoGe-2.
RealEstate10K [ 91 ]
DL3DV [ 38 ]
Training sampler
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
High overlap (median co-visibility 0.59)
19.0
0.600
0.280
6.7
14.9
0.366
0.466
13.7
Released sampler (0.43)
18.8
0.595
0.284
7.1
15.0
0.372
0.467
14.2
Low overlap (0.21)
18.4
0.572
0.310
8.5
14.8
0.366
0.489
19.8
Appendix
Table 8 : Training-pose overlap ablation. Three samplers differing only in the co-visibility distribution of the sampled training poses, at ablation scale. The released-sampler row is the Mix run of Table 3 . Bold : best; underline : second best.
Figure 13 : Co-visible PSNR of the training-pose overlap variants on RealEstate10K, binned by co-visibility, against the geometric warp teacher. The low-overlap sampler is the most robust at large viewpoint change and leads below co-visibility 0.4, while remaining softer near copies; the high-overlap sampler behaves the other way round. The two curves cross, which is the trade-off the released sampler sits between.
Figure 14 : Qualitative loss ablation on three RealEstate10K test scenes. All four models train for 250K steps under an identical recipe and differ only in their loss weights, matching rows 1, 4, 5 and 6 of Table 2 . Each row uses one source frame and the benchmark’s own relative pose to the target frame 24 frames later, held fixed across the four models, so Target is the true novel view. Removing the perceptual terms costs texture fidelity; removing the adversarial term leaves streaks and soft patches along the disoccluded edge; removing all three blurs the whole frame.
RealEstate10K [ 91 ]
DL3DV [ 38 ]
Model
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
MEt3R ↓
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
MEt3R ↓
OVIE, eval@256
18.8
0.602
0.279
6.74
0.035
14.8
0.369
0.464
13.6
0.078
OVIE-512, eval@256
19.1
0.611
0.283
7.62
0.034
15.2
0.389
0.464
14.8
0.077
OVIE-512, eval@512
18.9
0.658
0.344
–
–
15.1
0.475
0.509
–
–
Appendix
Table 9 : OVIE at 512×512 , zero-shot on both benchmarks. eval@256 is the like-for-like comparison, with every output brought to 256×256 before scoring; eval@512 scores OVIE-512 at its native resolution and is given for reference only. Bold : better of the two models at eval@256 .
Figure 15 : OVIE-512 novel views at 512×512 . Source images are held-out OpenImages [ 33 ] validation photographs at native 512×512 , so no upscaling is involved. Target poses are drawn from the training sampler of Section C.4 , and the pseudo-target column shows the MoGe-2 point cloud reprojected under that pose, exactly as during training. The generated views recover detail the reprojection cannot provide, such as fur, foliage, and crowd texture, and fill the regions the motion disoccludes. Co-visibility for these rows is 0.67, 0.47 and 0.43.
Figure 16 : OVIE-512 novel views at 512×512 , continued. Same protocol as Figure 15 . These rows show specular metal, a dense repeated structure, and layered flaky texture. Co-visibility is 0.52, 0.43 and 0.42.
Method
Params
PSNR ↑
SSIM ↑
LPIPS ↓
FID ↓
Time to first view
Zero-shot on RealEstate10K
SHARP [ 43 ]
702M
18.51
0.589
0.299
10.18
∼1 s
FlexWorld [ 12 ]
11.3B
18.96
0.625
0.290
19.32
∼225 s
OVIE (ours)
143M
18.8
0.602
0.279
6.74
<10 ms
In-domain on RealEstate10K
DiffusionGS [ 6 ]
–
21.34
0.731
0.217
8.99
∼0.8 s
Appendix
Table 10 : Comparison to single-image methods that build an explicit 3D representation , on RealEstate10K under our protocol. Time to first view includes any per-scene 3D construction; subsequent views are amortized for the explicit-3D methods and cost one forward pass for OVIE. The DiffusionGS row uses the public re-implementation, since the original model is unreleased and its authors report that it differs from the paper. Bold : best per regime.
Range
OVIE-ft (ours)
VIVID [ 20 ] (published)
Source-copy
Mid-range (gap 30–60)
17.24
17.36
13.12
Long-range (gap 60–120)
15.30
15.21
11.91
Appendix
Table 11 : RealEstate10K under VIVID’s original protocol (PSNR, higher is better). Source-copy returns the source image unchanged and bounds the difficulty of each range.
Parameter
Symbol
Value
Sampling probabilities
Identity
–
0.15
Pure translation
–
0.10
Pure rotation
–
0.10
Combined rotation & translation
–
0.35
Normal-derived
–
0.05
Appendix
Table 12 : Camera sampling hyperparameters and model settings used during training.
Hyperparameter
Value
Hyperparameter
Value
Generator Architecture
Losses
Resolution
256×256
Reconstruction loss
L2 (MSE)
Base channels
128
LPIPS weight λLPIPS
1.0
Channel multipliers
[1,2,4]
P-DINO model
DINOv3-ViT-B/16
Downsampling factor
8×
P-DINO weight λDINO
0.5
ViT bottleneck
Adversarial weight λadv
0.75
Appendix
Table 13 : Architecture, optimization, and loss hyperparameters.
Figure 17 : Qualitative results on out-of-distribution images. Each pair shows the input source image followed by the generated novel view. Source views are, from left to right and top to bottom: Gas by Edward Hopper, Untitled by Ralambo, A Sunday on La Grande Jatte by Georges Seurat, Nighthawks by Edward Hopper, Portrait of an Artist (Pool with Two Figures) by David Hockney, and The Sea of Ice by Caspar David Friedrich.
Figure 18 : Comparison of source inputs, training pseudo-targets, and generated views. During training, OVIE is supervised on pseudo-targets (middle) created by depth-lifting the source image (left) to a sampled pose. At inference, it generates novel views (right) from a source image and target pose. Here, the generated views are rendered at the same poses as their corresponding pseudo-targets.
Figure 19 : Comparison of source inputs, training pseudo-targets, and generated views. During training, OVIE is supervised on pseudo-targets (middle) created by depth-lifting the source image (left) to a sampled pose. At inference, it generates novel views (right) from a source image and target pose. Here, the generated views are rendered at the same poses as their corresponding pseudo-targets.
Figure 20 : Comparison of source inputs, training pseudo-targets, and generated views. During training, OVIE is supervised on pseudo-targets (middle) created by depth-lifting the source image (left) to a sampled pose. At inference, it generates novel views (right) from a source image and target pose. Here, the generated views are rendered at the same poses as their corresponding pseudo-targets.
Figure 21 : Comparison of source inputs, training pseudo-targets, and generated views. During training, OVIE is supervised on pseudo-targets (middle) created by depth-lifting the source image (left) to a sampled pose. At inference, it generates novel views (right) from a source image and target pose. Here, the generated views are rendered at the same poses as their corresponding pseudo-targets.
Figure 22 : Qualitative comparison with state-of-the-art methods. Given a source image and target camera pose, each method synthesizes a novel view. Despite training on no multi-view data, OVIE generates novel views that match or exceed the quality of concurrent methods.
Figure 23 : Qualitative comparison with state-of-the-art methods. Given a source image and target camera pose, each method synthesizes a novel view. Despite training on no multi-view data, OVIE generates novel views that match or exceed the quality of concurrent methods.
Figure 24 : Qualitative comparison with state-of-the-art methods. Given a source image and target camera pose, each method synthesizes a novel view. Despite training on no multi-view data, OVIE generates novel views that match or exceed the quality of concurrent methods.
Figure 25 : Qualitative comparison with InfiniteNature-Zero [ 36 ] . Given a source image (left), we show the ground-truth target view, the view generated by OVIE (ours), and the view generated by InfiniteNature-Zero. OVIE’s generated views are sharper and more realistic than those of InfiniteNature-Zero.
We study the challenging problem of novel view video synthesis from single images or monocular videos. Existing methods, which operate under the assumption that pre-trained video models lack native novel view synthesis capability and enforce view alignment via camera conditioning, task-specific fine-tuning, or stepwise hard denoising guidance, often suffer from artifacts and compromised global scene consistency. In this paper, we introduce NeoMap, a novel training-free framework designed to locate high-fidelity, view-consistent novel view solutions from general pre-trained video models. The key to our approach is the core insight that promising novel view solutions are inherently encoded within the natural video data manifold learned by pre-trained models, and the core challenge is simply to locate this optimal solution. We solve this via our core mechanism: convergent manifold alternating projection iterations that optimize the initial noise. Extensive experiments demonstrate that NeoMap significantly outperforms all existing methods across 3 standard novel view synthesis benchmarks, including the challenging Tanks-and-Temples, LLFF and DAVIS datasets, achieving state-of-the-art generation fidelity and top-tier view consistency.
Jinxi Li, Tianyi Zhang, Yafei Yang +4
Shenzhen Research Institute, The Hong Kong Polytechnic University · vLAR Group, The Hong Kong Polytechnic University
Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this paper, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization. Our code is publicly available at https://github.com/KlingAIResearch/MetaView.
Yufei Cai, Xuesong Niu, Hao Lu +3
Nanyang Technological University · Kolors Team, Kuaishou Technology · The Hong Kong University of Science and Technology (Guangzhou)
The abundance of casually captured monocular videos and images on social media provides a valuable source for immersive content creation, where generating novel views from such sparse observations can greatly enhance user experiences. However, producing photorealistic and geometrically consistent views with precise camera control remains challenging when input coverage is extremely limited. Reconstruction-based approaches such as NeRF and 3D Gaussian Splatting (3DGS) deteriorate severely under sparse inputs and fail to explicitly handle occlusions. Generative methods ease data requirements but still struggle with large-baseline view synthesis due to inaccurate or implicit geometric guidance. To overcome these limitations, we introduce UniWorld-View, a unified framework for controllable large-baseline novel view synthesis from monocular inputs. UniWorld-View integrates explicit 3D guidance with generative diffusion modeling to enable precise camera control and geometrically consistent view generation. The geometric guidance is obtained through an occlusion-aware point cloud rendering strategy that resolves visibility ambiguities and provides accurate priors for diffusion-based synthesis. By coupling this rendering strategy with powerful video diffusion backbones, UniWorld-View achieves high-fidelity novel view generation even under extreme camera motions and wide-baseline changes, and can further provide multi-view videos for downstream dynamic 3DGS reconstruction. Experiments on the WorldScore benchmark and zero-shot NVS benchmarks demonstrate the effectiveness of UniWorld-View in controllability, geometric consistency, and visual fidelity.