3D foundation models enable efficient novel view synthesis by carrying a Gaussian head on the representation they already use for reconstruction. However, the views they render fall short of the geometry they recover, because that geometry is estimated under a metric objective and never scored on how it renders. Recent methods alleviate this by updating the backbone weights, but they thereby discard the metric predictions the model was built for and must be repeated for every new backbone. To this end, we propose RDGSplat, a framework that decodes a second geometry dedicated to rendering from a frozen 3D foundation model, leaving its metric predictions intact. In particular, we devise Render-Dedicated Geometry Decoding, which duplicates the pretrained decoders and optimizes the duplicates under photometric supervision alone. Then, a Target-Pose Conditioned Adapter is introduced to reformulate the representation those decoders read, conditioned on the target camera pose rather than the target image. Extensive experiments show that RDGSplat improves novel view synthesis across three feed-forward backbones on four benchmarks, with every pretrained weight frozen. On RE10K, it raises WM2.0 from 20.918 to 24.266,dB while training 205.5,M added parameters against a frozen 1.4,B backbone, and the depth and pose the same model predicts are unchanged.
Figures & tables
Figure 2 : The two geometries of a single scene. Both are decoded from the same frozen backbone and the same context views. The metric geometry thickens the far wall into an opaque shell, and the views rendered from it lose the corridor behind it. The render-dedicated geometry resolves those surfaces into thin layers, through which the corridor is rendered.
Figure 3 : Overall framework. RDGSplat evaluates a frozen 3D foundation model twice and produces a metric geometry and a render-dedicated geometry. The token bars mark which views each evaluation carries: zA spans the context and target views, zB the context views alone. The unmodified evaluation feeds the pretrained heads, which predict the depth, camera poses and pointmaps forming the metric geometry. TCA takes Plücker rays, high-frequency content and the projected state as its condition, and adapts the second evaluation within the residual stream. RDGD decodes a render depth and per-primitive attributes from zB , which the rasterizer renders into the target.
Figure 4 : Render-Dedicated Geometry Decoding. (a) Duplicated depth and Gaussian decoders read zB and predict a render depth and per-primitive attributes, while the residual pose head reads both streams and corrects the context and target poses; the primitives are then unprojected and rasterized. (b) The residual pose head projects both streams, attends from the joint stream to the adapted context stream, and updates a frozen anchor pose.
Figure 5 : Target-Pose Conditioned Adapter. Plücker rays, high-frequency content and the projected state form the condition C , of which the per-frame patch encoder receives only the per-view prefix Clocal . Each adapter block is initialized from the layer it adapts, and enters the residual through a learned scale.
Method
Init.
Self-sup.
RE10K
ACID
DL3DV
DTU
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
pixelSplat [ 1 ]
random
–
23.859
0.808
0.184
25.477
0.770
0.207
21.370
0.713
0.250
15.067
0.539
0.341
MVSplat [ 3 ]
random
–
24.012
0.812
0.175
25.525
0.773
0.199
20.344
0.673
0.258
14.542
0.537
0.324
NoPoSplat [ 62 ]
random
–
21.382
0.709
0.266
23.193
0.667
0.274
18.235
0.518
0.364
15.113
0.455
0.450
NAS3R † [ 12 ]
random
✓
23.594
0.766
0.188
25.030
0.735
0.209
17.646
0.466
0.375
15.229
0.524
0.317
NAS3R + RDGSplat
random
✓
24.906
0.791
0.170
26.343
0.763
0.185
19.077
0.505
0.342
15.461
0.532
0.307
Table 1 : Two-view novel view synthesis. The best results among self-supervised methods under each initialization are highlighted. † indicates variants reproduced from officially released codebase and weights.
Method
Video depth (KITTI)
Two-view pose (RE10K)
Abs Rel ↓
δ<1.25↑
@10∘↑
@20∘↑
DUSt3R [ 52 ]
0.144
81.3
54.1
70.2
MASt3R [ 24 ]
0.183
74.5
49.4
67.1
VGGT [ 48 ]
0.061
97.0
47.4
65.8
WM2.0 + RDGSplat
0.060
97.4
71.6
82.9
Table 2 : Metric geometry. Video depth follows the CUT3R protocol with per-sequence scale alignment [ 51 ] , and two-view pose the AUC protocol of NAS3R [ 12 ] . The best and second-best results are highlighted.
RDGD
TCA
RE10K
ACID
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
1
–
–
20.918
0.747
0.210
21.752
0.679
0.266
2
✓
–
23.485
0.801
0.175
24.145
0.740
0.216
3
–
✓
22.613
0.780
0.194
23.615
0.727
0.227
4
✓
✓
23.861
0.808
0.171
24.387
0.748
0.210
Table 3 : Ablation studies of each component. All rows are trained for 30k iterations, and Row 1 denotes WM2.0 with neither module attached. The best results are highlighted.
Method
Pretrained weights updated
RDGD
TCA
Trainable
Share
WM2.0 [ 45 ]
0
–
–
–
–
NAS3R [ 12 ]
all
–
–
1.4 B
100%
WM2.0 + RDGSplat
0
67.3 M
138.2 M
205.5 M
14.7%
NAS3R + RDGSplat
0
67.1 M
138.2 M
205.3 M
14.7%
Table 4 : Training efficiency. A dash marks a quantity that does not apply to a frozen baseline.
Views
Method
PSNR ↑
SSIM ↑
LPIPS ↓
2
WM2.0
20.918
0.747
0.210
WM2.0 + RDGSplat
23.582
0.800
0.176
3
WM2.0
22.592
0.792
0.179
WM2.0 + RDGSplat
24.996
0.838
0.146
5
WM2.0
23.796
0.822
0.157
WM2.0 + RDGSplat
26.303
0.871
0.121
Table 5 : Performance on novel view synthesis across different numbers of context views. The best results are highlighted.
Figure 6 : Qualitative comparison on RE10K. Scenes sampled at random from the test split. The leftmost column shows the two context views, and two held-out target views follow.
Component
Value
Injection sites ∣S∣
4 per stream, 8 in total
Site placement
blocks 4, 11, 17, 23
Token width d
1024
Condition width
153
MLPl
153→4096→1024
Fourier bands
4
Table 1 : Architecture detail. The same four injection depths are used in the patch encoder and in the cross-view blocks. The three decoder sizes sum to the 67.3 M of Table 4 of the main paper.
Figure 1 : The two geometries across scenes on RE10K. Both are decoded from the same frozen backbone and the same context views, and the views rendered from each are shown beneath it. The metric geometry thickens the far surfaces into opaque shells, and the views rendered from it lose what lies behind them. The render-dedicated geometry resolves those surfaces into thin layers, through which the same regions are rendered.
Figure 2 : Additional novel view synthesis on RE10K. The leftmost column shows the two context views, and two held-out target views follow under each method.
gs
gd
Res. pose
RE10K
PSNR ↑
SSIM ↑
LPIPS ↓
1
–
–
–
20.918
0.747
0.210
2
✓
–
–
22.456
0.776
0.193
3
✓
✓
–
23.338
0.796
0.176
4
✓
✓
✓
23.485
0.801
0.175
Table 2 : Decomposition of RDGD. All rows are trained for 30k iterations, and Row 1 denotes WM2.0 with nothing attached. The best results are highlighted.
Inherited
RE10K
PSNR ↑
SSIM ↑
LPIPS ↓
1
–
17.449
0.545
0.407
2
✓
23.485
0.801
0.175
Table 3 : Initialization of the duplicated decoders. All rows are trained for 30k iterations, and Row 1 starts the three decoders from random weights. The best results are highlighted.
Blocks
Cond.
RE10K
PSNR ↑
SSIM ↑
LPIPS ↓
1
–
–
20.918
0.747
0.210
2
✓
–
21.520
0.758
0.205
3
✓
✓
22.613
0.780
0.194
Table 4 : The condition in TCA. All rows are trained for 30k iterations, and Row 1 denotes WM2.0 with nothing attached. The best results are highlighted.
Geometry
Median error
AUC
Rot. ↓
Trans. ↓
@10 ∘ ↑
@20 ∘ ↑
Metric geometry
0.622
1.589
71.6
82.9
Render-dedicated geometry
0.622
1.609
71.5
82.9
Table 5 : Two-view pose, the two branches. Median angular error in degrees, and AUC of the larger of the two errors, under the protocol of Table 2 of the main paper.
Views
Method
Latency (ms) ↓
Peak memory (GB) ↓
2
WM2.0 [ 45 ]
262.3
7.74
WM2.0 + RDGSplat
419.8
7.78
3
WM2.0
286.9
7.85
WM2.0 + RDGSplat
464.0
7.86
5
WM2.0
351.4
8.11
WM2.0 + RDGSplat
579.1
8.11
Table 6 : Inference cost. One forward pass per row on a single NVIDIA A6000, with latency the median over eight scenes and three repetitions.
We present ViewSplat, a view-adaptive 3D Gaussian splatting network for novel view synthesis from unposed images. While recent feed-forward 3D Gaussian splatting has significantly accelerated 3D scene reconstruction by bypassing per-scene optimization, a fundamental fidelity gap remains. We attribute this gap to the limited capacity of single-step feed-forward networks to regress static Gaussian primitives that satisfy all viewpoints. To address this limitation, we shift the paradigm from static primitive regression to view-adaptive splatting. Instead of a rigid Gaussian representation, our pipeline learns a view-adaptive latent representation. Specifically, ViewSplat initially predicts base Gaussian primitives alongside the weights of scene-conditioned View MLPs. During rendering, these MLPs take target-view coordinates as input and predict view-dependent residual updates for each Gaussian attribute (i.e., 3D position, scale, rotation, opacity, and color). This mechanism, which we term view-adaptive splatting, allows each primitive to rectify initial estimation errors, effectively capturing high-fidelity appearances. Extensive experiments demonstrate that ViewSplat achieves state-of-the-art fidelity while maintaining fast inference and real-time rendering; our large backbone variant runs at 15 FPS during inference and 90 FPS during rendering. Our project page is available at https://cvlab-uos.github.io/ViewSplat.
Moonyeon Jeong, Seunggi Min, Suhyeon Lee +1
University of Seoul · Korea Electronics Technology Institute
3D Gaussians have become a powerful scene representation for real-time splatting and high-quality novel-view synthesis. This has motivated generalizable splatting -- methods that adapt feed-forward geometry prediction networks to produce per-pixel Gaussians from a set of images. However, most generalizable splatting pipelines are supervised primarily through a view-synthesis loss to predict Gaussian orientation, anisotropic scale, opacity, and appearance in addition to their locations. We show that this learning objective is under-constrained. Models trained with view synthesis alone produce splats whose orientations and scales have no geometric connotation. The result is that, while producing decent view-synthesis performance, nearly all generalizable splatting methods produce geometrically inaccurate and misaligned Gaussians. We introduce G3Splat, a geometry-consistent generalizable splatting framework that addresses these degeneracies through differentiable geometric priors on the predicted 3D Gaussians, making the learning problem well-posed. These priors encourage the per-pixel splats to remain on their viewing rays and to orient themselves in accordance with local surfaces. Our priors are architecture-agnostic and can be incorporated into any previously studied geometric backbone for generalizable splatting, as well as different scene representations. We test G3Splat with both DUSt3R-style and VGGT-style backbones to predict pixel-aligned full-rank 3DGS as well as surfel-like 2DGS. Trained on RE10K, G3Splat produces Gaussian splats with significantly higher geometric fidelity than baselines, providing state-of-the-art novel-view depth, mesh reconstruction, and relative pose estimation performance while preserving novel-view synthesis quality, as evaluated on datasets such as ACID and ScanNet. Code and pretrained models are released on our project page.
Mehdi Hosseinzadeh, Shin-Fang Chng, Yi Xu +3
Australian Institute for Machine Learning · Goertek Alpha Labs · MBZUAI
While feed-forward 3D Gaussian Splatting (3DGS) enables efficient 3D reconstruction, achieving high-fidelity rendering remains challenging. Existing pixel-aligned approaches suffer from spatial inflexibility and massive structural redundancy, whereas query-based methods lack 3D priors and entangle geometry with appearance, yielding blurry, pose-dependent results. To overcome these deficiencies, we propose \textbf{QuerySplat}, a feed-forward 3DGS framework driven by geometric priors and explicit appearance decoupling. Specifically, we design a dual-branch query-based decoder: the geometry branch leverages a pretrained Vision Geometric Model for spatial understanding, which intrinsically endows QuerySplat with pose-free modeling capabilities, while the appearance branch recovers high-frequency details through a dedicated pathway separated from geometric attribute regression. Extensive experiments demonstrate that QuerySplat mitigates the blurry rendering issues of early query-based models and consistently outperforms pixel-aligned approaches in rendering fidelity. On the challenging DL3DV benchmark, it achieves state-of-the-art novel view synthesis performance, with average PSNR gains of 2.30 dB and 1.04 dB over the best pose-free and pose-required baselines, respectively. Project Page: https://inspatio.github.io/querysplat.
Yinglong Li, Donghui Shen, Xiaoyu Zhang +5
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University · InSpatio Research · State Key Lab of CAD&CG, Zhejiang University