3D foundation models enable efficient novel view synthesis by carrying a Gaussian head on the representation they already use for reconstruction. However, the views they render fall short of the geometry they recover, because that geometry is estimated under a metric objective and never scored on how it renders. Recent methods alleviate this by updating the backbone weights, but they thereby discard the metric predictions the model was built for and must be repeated for every new backbone. To this end, we propose RDGSplat, a framework that decodes a second geometry dedicated to rendering from a frozen 3D foundation model, leaving its metric predictions intact. In particular, we devise Render-Dedicated Geometry Decoding, which duplicates the pretrained decoders and optimizes the duplicates under photometric supervision alone. Then, a Target-Pose Conditioned Adapter is introduced to reformulate the representation those decoders read, conditioned on the target camera pose rather than the target image. Extensive experiments show that RDGSplat improves novel view synthesis across three feed-forward backbones on four benchmarks, with every pretrained weight frozen. On RE10K, it raises WM2.0 from 20.918 to 24.266,dB while training 205.5,M added parameters against a frozen 1.4,B backbone, and the depth and pose the same model predicts are unchanged.
Figures & tables
Figure 2 : The two geometries of a single scene. Both are decoded from the same frozen backbone and the same context views. The metric geometry thickens the far wall into an opaque shell, and the views rendered from it lose the corridor behind it. The render-dedicated geometry resolves those surfaces into thin layers, through which the corridor is rendered.
Figure 3 : Overall framework. RDGSplat evaluates a frozen 3D foundation model twice and produces a metric geometry and a render-dedicated geometry. The token bars mark which views each evaluation carries: zA spans the context and target views, zB the context views alone. The unmodified evaluation feeds the pretrained heads, which predict the depth, camera poses and pointmaps forming the metric geometry. TCA takes Plücker rays, high-frequency content and the projected state as its condition, and adapts the second evaluation within the residual stream. RDGD decodes a render depth and per-primitive attributes from zB , which the rasterizer renders into the target.
Figure 4 : Render-Dedicated Geometry Decoding. (a) Duplicated depth and Gaussian decoders read zB and predict a render depth and per-primitive attributes, while the residual pose head reads both streams and corrects the context and target poses; the primitives are then unprojected and rasterized. (b) The residual pose head projects both streams, attends from the joint stream to the adapted context stream, and updates a frozen anchor pose.
Figure 5 : Target-Pose Conditioned Adapter. Plücker rays, high-frequency content and the projected state form the condition C , of which the per-frame patch encoder receives only the per-view prefix Clocal . Each adapter block is initialized from the layer it adapts, and enters the residual through a learned scale.
Method
Init.
Self-sup.
RE10K
ACID
DL3DV
DTU
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
pixelSplat [ 1 ]
random
–
23.859
0.808
0.184
25.477
0.770
0.207
21.370
0.713
0.250
15.067
0.539
0.341
MVSplat [ 3 ]
random
–
24.012
0.812
0.175
25.525
0.773
0.199
20.344
0.673
0.258
14.542
0.537
0.324
NoPoSplat [ 62 ]
random
–
21.382
0.709
0.266
23.193
0.667
0.274
18.235
0.518
0.364
15.113
0.455
0.450
NAS3R † [ 12 ]
random
✓
23.594
0.766
0.188
25.030
0.735
0.209
17.646
0.466
0.375
15.229
0.524
0.317
NAS3R + RDGSplat
random
✓
24.906
0.791
0.170
26.343
0.763
0.185
19.077
0.505
0.342
15.461
0.532
0.307
Table 1 : Two-view novel view synthesis. The best results among self-supervised methods under each initialization are highlighted. † indicates variants reproduced from officially released codebase and weights.
Method
Video depth (KITTI)
Two-view pose (RE10K)
Abs Rel ↓
δ<1.25↑
@10∘↑
@20∘↑
DUSt3R [ 52 ]
0.144
81.3
54.1
70.2
MASt3R [ 24 ]
0.183
74.5
49.4
67.1
VGGT [ 48 ]
0.061
97.0
47.4
65.8
WM2.0 + RDGSplat
0.060
97.4
71.6
82.9
Table 2 : Metric geometry. Video depth follows the CUT3R protocol with per-sequence scale alignment [ 51 ] , and two-view pose the AUC protocol of NAS3R [ 12 ] . The best and second-best results are highlighted.
RDGD
TCA
RE10K
ACID
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
1
–
–
20.918
0.747
0.210
21.752
0.679
0.266
2
✓
–
23.485
0.801
0.175
24.145
0.740
0.216
3
–
✓
22.613
0.780
0.194
23.615
0.727
0.227
4
✓
✓
23.861
0.808
0.171
24.387
0.748
0.210
Table 3 : Ablation studies of each component. All rows are trained for 30k iterations, and Row 1 denotes WM2.0 with neither module attached. The best results are highlighted.
Method
Pretrained weights updated
RDGD
TCA
Trainable
Share
WM2.0 [ 45 ]
0
–
–
–
–
NAS3R [ 12 ]
all
–
–
1.4 B
100%
WM2.0 + RDGSplat
0
67.3 M
138.2 M
205.5 M
14.7%
NAS3R + RDGSplat
0
67.1 M
138.2 M
205.3 M
14.7%
Table 4 : Training efficiency. A dash marks a quantity that does not apply to a frozen baseline.
Views
Method
PSNR ↑
SSIM ↑
LPIPS ↓
2
WM2.0
20.918
0.747
0.210
WM2.0 + RDGSplat
23.582
0.800
0.176
3
WM2.0
22.592
0.792
0.179
WM2.0 + RDGSplat
24.996
0.838
0.146
5
WM2.0
23.796
0.822
0.157
WM2.0 + RDGSplat
26.303
0.871
0.121
Table 5 : Performance on novel view synthesis across different numbers of context views. The best results are highlighted.
Figure 6 : Qualitative comparison on RE10K. Scenes sampled at random from the test split. The leftmost column shows the two context views, and two held-out target views follow.
Component
Value
Injection sites ∣S∣
4 per stream, 8 in total
Site placement
blocks 4, 11, 17, 23
Token width d
1024
Condition width
153
MLPl
153→4096→1024
Fourier bands
4
Table 1 : Architecture detail. The same four injection depths are used in the patch encoder and in the cross-view blocks. The three decoder sizes sum to the 67.3 M of Table 4 of the main paper.
Figure 1 : The two geometries across scenes on RE10K. Both are decoded from the same frozen backbone and the same context views, and the views rendered from each are shown beneath it. The metric geometry thickens the far surfaces into opaque shells, and the views rendered from it lose what lies behind them. The render-dedicated geometry resolves those surfaces into thin layers, through which the same regions are rendered.
Figure 2 : Additional novel view synthesis on RE10K. The leftmost column shows the two context views, and two held-out target views follow under each method.
gs
gd
Res. pose
RE10K
PSNR ↑
SSIM ↑
LPIPS ↓
1
–
–
–
20.918
0.747
0.210
2
✓
–
–
22.456
0.776
0.193
3
✓
✓
–
23.338
0.796
0.176
4
✓
✓
✓
23.485
0.801
0.175
Table 2 : Decomposition of RDGD. All rows are trained for 30k iterations, and Row 1 denotes WM2.0 with nothing attached. The best results are highlighted.
Inherited
RE10K
PSNR ↑
SSIM ↑
LPIPS ↓
1
–
17.449
0.545
0.407
2
✓
23.485
0.801
0.175
Table 3 : Initialization of the duplicated decoders. All rows are trained for 30k iterations, and Row 1 starts the three decoders from random weights. The best results are highlighted.
Blocks
Cond.
RE10K
PSNR ↑
SSIM ↑
LPIPS ↓
1
–
–
20.918
0.747
0.210
2
✓
–
21.520
0.758
0.205
3
✓
✓
22.613
0.780
0.194
Table 4 : The condition in TCA. All rows are trained for 30k iterations, and Row 1 denotes WM2.0 with nothing attached. The best results are highlighted.
Geometry
Median error
AUC
Rot. ↓
Trans. ↓
@10 ∘ ↑
@20 ∘ ↑
Metric geometry
0.622
1.589
71.6
82.9
Render-dedicated geometry
0.622
1.609
71.5
82.9
Table 5 : Two-view pose, the two branches. Median angular error in degrees, and AUC of the larger of the two errors, under the protocol of Table 2 of the main paper.
Views
Method
Latency (ms) ↓
Peak memory (GB) ↓
2
WM2.0 [ 45 ]
262.3
7.74
WM2.0 + RDGSplat
419.8
7.78
3
WM2.0
286.9
7.85
WM2.0 + RDGSplat
464.0
7.86
5
WM2.0
351.4
8.11
WM2.0 + RDGSplat
579.1
8.11
Table 6 : Inference cost. One forward pass per row on a single NVIDIA A6000, with latency the median over eight scenes and three repetitions.