Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes. Its quality, however, degrades sharply as the number of input images drops. The bottleneck is upstream of the reconstruction heads: from a few unposed views, the internal representation they read carries no evidence for unobserved regions, leaving holes, floaters, and blur. The common remedy supplies that evidence as pixels, synthesizing extra views with an image or video generator and re-encoding them, which is costly and not 3D-consistent by construction. We instead densify the evidence itself. We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given. DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads. The same module design can be integrated into different pretrained predictors while keeping each backbone and its reconstruction heads frozen. Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes. Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.
Figures & tables
Figure 1: DensiTok augments geometry tokens for unobserved views while keeping the pretrained encoder and reconstruction heads frozen. With only three input views, it enables AnySplat to exceed its 16-view baseline in PSNR on RealEstate10K.
Figure 2: Geometry token compression and completion. Stage 1 trains a VAE with Level-Adaptive Modulation (LAM-VAE) to compress and reconstruct geometry tokens. level-adaptive fusion and the VAE encoder together form the LAM-VAE encoder, with a level-wise residual into the posterior mean. Stage 2 freezes the VAE and trains a camera-conditioned DiT to complete latents for unobserved views. The geometry encoder remains frozen throughout, while the reconstruction heads remain frozen or undergo minimal appearance adaptation.
Methods
RealEstate10K ( Zhou et al., 2018 )
DL3DV ( Ling et al., 2024 )
Runtime
Appearance
Camera
Depth
Appearance
Camera
Depth
PSNR ↑
SSIM ↑
LPIPS ↓
AUC@ 3∘↑
AUC@ 30∘↑
δ1.25↑
AbsRel ↓
PSNR ↑
SSIM ↑
LPIPS ↓
AUC@ 3∘↑
AUC@ 30∘↑
δ1.25↑
AbsRel ↓
(s) ↓
Dense-view ( N=16 )
AnySplat ( Jiang et al., 2025 )
24.04
0.792
0.189
20.4
77.5
0.835
0.151
22.30
0.683
0.246
24.1
85.4
0.803
0.158
0.434
WorldMirror ( Liu et al., 2025 )
25.11
0.856
0.161
41.8
83.3
0.880
0.110
24.47
0.819
0.173
69.0
94.0
0.825
0.141
0.495
Sparse-view ( n=2 )
Table 1: Quantitative results on RealEstate10K and DL3DV with n∈{2,3,5,7} sparse input views and dense-view references ( N=16 ).
Figure 3: Qualitative novel-view RGB and depth comparisons for AnySplat with and without DensiTok on RE10K and DL3DV, using n=2 input views (left) and n=3 input views (right).
RE10K
DL3DV
LAM-VAE encoder
PSNR
SSIM
LPIPS
AUC@ 3∘
AUC@ 30∘
δ1.25
AbsRel
PSNR
SSIM
LPIPS
AUC@ 3∘
AUC@ 30∘
δ1.25
AbsRel
Uniform fusion
19.76
0.673
0.266
18.3
76.3
0.808
0.184
19.53
0.611
0.314
21.7
85.2
0.760
0.199
MLP fusion
21.19
0.725
0.243
20.8
77.4
0.831
0.164
20.53
0.644
0.291
23.8
86.1
0.781
0.181
+ level-adaptive fusion
22.11
0.756
0.228
25.0
78.1
0.836
0.159
21.89
0.684
0.267
28.8
87.6
0.787
0.177
+ level-wise residual
21.86
0.747
0.234
24.2
78.6
0.841
0.156
21.56
0.675
0.274
27.5
88.2
0.794
0.173
+ both
22.56
0.766
0.223
26.7
78.9
0.846
0.153
22.40
0.699
0.255
30.4
88.9
0.801
0.170
Table 2: Ablations of Level-Adaptive Modulation with AnySplat and n=2 observed views.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Architecture of a DiT block for geometry latent augmentation. Time and camera embeddings produce per-token AdaLN shift, scale, and gate parameters for the attention and MLP branches. Frame and global blocks differ only in their attention scope.
Setting
VAE encoder
VAE decoder
DiT
Hidden channels
512
512
768
Frame/global block pairs
6
6
12
Total transformer blocks
12
12
24
Attention heads
8
8
12
Appendix
Table 4: Shared VAE and DiT architecture for all three backbones. Each pair contains a frame block followed by a global block.
Setting
VAE
DiT
One-step finetuning
Optimizer
AdamW
AdamW
AdamW
Maximum learning rate
2×10−4
2×10−4
4×10−5
Minimum learning rate
10−6
2×10−6
2×10−6
Learning-rate schedule
Cosine decay
Cosine decay
Cosine decay
Optimizer betas
(0.9,0.95)
(0.9,0.95)
(0.9,0.95)
Weight decay
0.01
0.01
0.01
Appendix
Table 5: Training configurations for the three stages of the main AnySplat experiments.
Methods
RealEstate10K ( Zhou et al., 2018 )
DL3DV ( Ling et al., 2024 )
Runtime
Appearance
Camera
Depth
Appearance
Camera
Depth
PSNR ↑
SSIM ↑
LPIPS ↓
AUC@ 3∘↑
AUC@ 30∘↑
δ1.25↑
AbsRel ↓
PSNR ↑
SSIM ↑
LPIPS ↓
AUC@ 3∘↑
AUC@ 30∘↑
δ1.25↑
AbsRel ↓
(s) ↓
Dense-view ( N=16 )
DA3 ( Lin et al., 2025 )
21.28
0.712
0.249
43.8
86.2
0.538
0.346
22.73
0.717
0.222
87.8
97.2
0.556
0.322
0.514
Sparse-view ( n=2 )
DA3
19.69
0.667
0.300
55.8
89.6
0.369
0.471
20.86
0.646
0.287
88.3
98.8
0.402
0.452
0.112
Appendix
Table 6: Quantitative results for DA3 on RealEstate10K and DL3DV with n∈{2,3,5,7} sparse input views and dense-view references ( N=16 ). Camera AUC is reported in percent. Runtime is reported in seconds in a shared column.
Figure 5: Qualitative novel-view RGB and depth comparisons on RE10K and DL3DV with n=2 input views (left) and n=3 input views (right), each with and without DensiTok. Red boxes mark the enlarged RGB and depth crops. GT denotes the ground-truth target image.
Figure 6: Qualitative novel-view RGB and depth comparisons on RE10K and DL3DV with n=5 input views (left) and n=7 input views (right), each with and without DensiTok. Red boxes mark the enlarged RGB and depth crops. GT denotes the ground-truth target image.
Figure 7: Visualization of Gaussian reconstructions and camera poses for AnySplat with and without DensiTok. The top row shows the original AnySplat backbone and the bottom row shows DensiTok. Each scene is paired with a camera visualization, where red denotes ground truth and blue denotes predictions. DensiTok produces clearer scene structure and closer agreement with ground-truth camera positions and viewing orientations.
RE10K
DL3DV
Input cameras
Appearance
Depth
Appearance
Depth
Method
Intr.
Extr.
PSNR
SSIM
LPIPS
δ1.25
AbsRel
PSNR
SSIM
LPIPS
δ1.25
AbsRel
Runtime
↑
↑
↓
↑
↓
↑
↑
↓
↑
↓
(s) ↓
Sparse-view ( n=2 )
pixelSplat
∘
∘
26.125
0.864
0.131
0.608
0.359
23.358
0.721
0.239
0.489
0.467
0.088
MVSplat
∘
∘
26.319
0.865
0.127
0.769
0.220
22.894
0.704
0.235
0.687
0.275
0.039
Appendix
Table 7: Comparison with sparse-view feed-forward Gaussian Splatting methods on RE10K and DL3DV at 256×256 with n∈{2,3,5,7} . NoPoSplat results are included for n∈{2,3} . Intr. and Extr. indicate whether the evaluated configuration requires externally supplied input camera intrinsics and extrinsics for reconstruction. Orange and yellow mark the best and second-best values for each metric at the same n , respectively. Rankings are determined before rounding, with ties sharing the same color.
Figure 8: Qualitative novel-view RGB comparisons with sparse-view feed-forward Gaussian Splatting models. DensiTok uses the AnySplat backbone and requires neither input camera intrinsics nor extrinsics. Methods are grouped by the input camera information they require. The leftmost column shows two input views for each scene, and GT denotes the ground-truth target image. Red boxes mark regions enlarged below each rendering.
RE10K
DL3DV
Method
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Runtime
↑
↑
↓
↑
↑
↓
(s) ↓
Sparse-view ( n=2 )
NVComposer
14.897
0.528
0.508
12.078
0.359
0.610
48.502
Matrix3D
20.934
0.700
0.284
18.149
0.523
0.376
25.592
GLD
23.608
0.768
0.219
21.440
0.621
0.270
76.824
Appendix
Table 8: Comparison with diffusion-based novel-view synthesis methods on RE10K and DL3DV at 512×512 with n∈{2,3,5,7} . Runtime is reported in seconds in a shared column. Orange and yellow mark the best and second-best values for each metric at the same n , respectively. Rankings are determined before rounding, with ties sharing the same color.
Figure 9: Qualitative novel-view RGB comparisons of DensiTok with GLD, Matrix3D, and NVComposer using n=2 input views. The leftmost column shows the input views for each scene, and GT denotes the ground-truth target image. Red boxes mark regions enlarged below each rendering to compare the preservation of scene content and local geometry.
n
Methods
Appearance
Camera
Depth
PSNR ↑
SSIM ↑
LPIPS ↓
AUC@ 3∘↑
AUC@ 30∘↑
δ1.25↑
AbsRel ↓
ETH3D ( Schops et al., 2017 )
2
AnySplat
15.99
0.462
0.489
27.3
89.4
0.767
0.148
+ DensiTok
17.01
0.531
0.466
33.3
90.3
0.771
0.135
3
AnySplat
16.79
0.487
0.461
29.3
88.0
0.831
0.119
+ DensiTok
17.92
0.554
0.447
31.3
89.2
0.837
0.116
Appendix
Table 9: Generalization to ETH3D, ScanNet, and Mip-NeRF 360 with the AnySplat backbone and n∈{2,3,5,7} observed views. Camera AUC is reported in percent. Depth is evaluated against official ground truth on ETH3D and ScanNet and MoGe-2 predictions on Mip-NeRF 360.
RE10K
DL3DV
Efficiency
Method
NFE
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Runtime
↓
↑
↑
↓
↑
↑
↓
(s) ↓
Deterministic regression
1
20.86
0.712
0.246
19.97
0.631
0.288
0.258
One-step training from scratch
1
21.67
0.746
0.240
21.28
0.670
0.285
0.258
FM pretraining only
1
21.56
0.743
0.243
21.09
0.665
0.289
0.258
8
21.84
0.755
0.236
21.59
0.676
0.279
0.990
Appendix
Table 10: Effect of one-step finetuning with AnySplat, dz=4 , n=2 , and N=16 . FM denotes flow matching and NFE denotes the number of DiT forward evaluations. Runtime is reported in a shared column.
RE10K
DL3DV
Appearance
Camera
Depth
Appearance
Camera
Depth
n
dz
PSNR ↑
SSIM ↑
LPIPS ↓
AUC@ 3∘↑
AUC@ 30∘↑
δ1.25↑
AbsRel ↓
PSNR ↑
SSIM ↑
LPIPS ↓
AUC@ 3∘↑
AUC@ 30∘↑
δ1.25↑
AbsRel ↓
2
4
22.56
0.766
0.223
26.7
78.9
0.846
0.153
22.40
0.699
0.255
30.4
88.9
0.801
0.170
16
22.64
0.770
0.215
27.5
79.0
0.854
0.151
22.24
0.700
0.247
39.6
90.8
0.804
0.164
3
4
24.28
0.813
0.184
25.8
79.7
0.853
0.146
24.28
0.755
0.205
34.4
89.6
0.809
0.158
16
24.57
0.822
0.171
27.1
79.5
0.862
0.145
24.29
0.760
0.194
38.5
90.5
0.812
0.155
Appendix
Table 11: Effect of latent channel count with AnySplat, N=16 , and one-step inference.
RE10K
DL3DV
Appearance
Camera
Depth
Appearance
Camera
Depth
n
Gaussian head
PSNR ↑
SSIM ↑
LPIPS ↓
AUC@ 3∘↑
AUC@ 30∘↑
δ1.25↑
AbsRel ↓
PSNR ↑
SSIM ↑
LPIPS ↓
AUC@ 3∘↑
AUC@ 30∘↑
δ1.25↑
AbsRel ↓
2
Fully frozen
22.14
0.754
0.233
26.4
78.8
0.833
0.161
22.01
0.685
0.267
30.6
88.8
0.790
0.179
Image embedding only
22.56
0.766
0.223
26.7
78.9
0.846
0.153
22.40
0.699
0.255
30.4
88.9
0.801
0.170
Full finetuning
22.88
0.775
0.216
26.8
78.8
0.854
0.148
22.69
0.708
0.248
30.2
88.8
0.807
0.166
3
Fully frozen
24.07
0.806
0.190
25.6
78.8
0.846
0.150
24.05
0.747
0.212
34.1
88.8
0.803
0.163
Appendix
Table 12: Effect of Gaussian head finetuning with AnySplat, dz=4 , N=16 , and one-step inference. Image embedding only denotes finetuning himg while keeping the remaining pretrained Gaussian head weights frozen.