DensiTok: Making Feed-Forward 3D Gaussian Splatting See More Views Than It Is Given
Organizations: Yonsei University · NAVER AI Lab · Korea Institute of Science and Technology (KIST)
Abstract
Feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene in a single forward pass, replacing per-scene optimization with a network trained across many scenes. Its quality, however, degrades sharply as the number of input images drops. The bottleneck is upstream of the reconstruction heads: from a few unposed views, the internal representation they read carries no evidence for unobserved regions, leaving holes, floaters, and blur. The common remedy supplies that evidence as pixels, synthesizing extra views with an image or video generator and re-encoding them, which is costly and not 3D-consistent by construction. We instead densify the evidence itself. We present DensiTok, a plug-in module for pretrained feed-forward 3DGS models that densifies their internal geometry tokens directly, making a frozen backbone behave as though it had observed many more views than it was given. DensiTok compresses those tokens into a compact latent space, completes the latents of the unobserved viewpoints in a single flow-matching step conditioned on camera geometry, and decodes them back into tokens that the original reconstruction heads. The same module design can be integrated into different pretrained predictors while keeping each backbone and its reconstruction heads frozen. Completion in a low-dimensional latent space requires no image synthesis or additional encoder passes. Across three pretrained backbones and two benchmarks, DensiTok consistently improves sparse-view reconstruction and recovers much of the gap to dense-view reconstruction.
Figures & tables
| Methods | RealEstate10K ( Zhou et al., 2018 ) | DL3DV ( Ling et al., 2024 ) | Runtime | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Appearance | Camera | Depth | Appearance | Camera | Depth | ||||||||||
| PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | (s) | |||
| Dense-view ( ) | |||||||||||||||
| AnySplat ( Jiang et al., 2025 ) | 24.04 | 0.792 | 0.189 | 20.4 | 77.5 | 0.835 | 0.151 | 22.30 | 0.683 | 0.246 | 24.1 | 85.4 | 0.803 | 0.158 | 0.434 |
| WorldMirror ( Liu et al., 2025 ) | 25.11 | 0.856 | 0.161 | 41.8 | 83.3 | 0.880 | 0.110 | 24.47 | 0.819 | 0.173 | 69.0 | 94.0 | 0.825 | 0.141 | 0.495 |
| Sparse-view ( ) | |||||||||||||||
| RE10K | DL3DV | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| LAM-VAE encoder | PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | ||
| Uniform fusion | 19.76 | 0.673 | 0.266 | 18.3 | 76.3 | 0.808 | 0.184 | 19.53 | 0.611 | 0.314 | 21.7 | 85.2 | 0.760 | 0.199 |
| MLP fusion | 21.19 | 0.725 | 0.243 | 20.8 | 77.4 | 0.831 | 0.164 | 20.53 | 0.644 | 0.291 | 23.8 | 86.1 | 0.781 | 0.181 |
| + level-adaptive fusion | 22.11 | 0.756 | 0.228 | 25.0 | 78.1 | 0.836 | 0.159 | 21.89 | 0.684 | 0.267 | 28.8 | 87.6 | 0.787 | 0.177 |
| + level-wise residual | 21.86 | 0.747 | 0.234 | 24.2 | 78.6 | 0.841 | 0.156 | 21.56 | 0.675 | 0.274 | 27.5 | 88.2 | 0.794 | 0.173 |
| + both | 22.56 | 0.766 | 0.223 | 26.7 | 78.9 | 0.846 | 0.153 | 22.40 | 0.699 | 0.255 | 30.4 | 88.9 | 0.801 | 0.170 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | VAE encoder | VAE decoder | DiT |
|---|---|---|---|
| Hidden channels | 512 | 512 | 768 |
| Frame/global block pairs | 6 | 6 | 12 |
| Total transformer blocks | 12 | 12 | 24 |
| Attention heads | 8 | 8 | 12 |
| Setting | VAE | DiT | One-step finetuning |
|---|---|---|---|
| Optimizer | AdamW | AdamW | AdamW |
| Maximum learning rate | |||
| Minimum learning rate | |||
| Learning-rate schedule | Cosine decay | Cosine decay | Cosine decay |
| Optimizer betas | |||
| Weight decay |
| Methods | RealEstate10K ( Zhou et al., 2018 ) | DL3DV ( Ling et al., 2024 ) | Runtime | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Appearance | Camera | Depth | Appearance | Camera | Depth | ||||||||||
| PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | (s) | |||
| Dense-view ( ) | |||||||||||||||
| DA3 ( Lin et al., 2025 ) | 21.28 | 0.712 | 0.249 | 43.8 | 86.2 | 0.538 | 0.346 | 22.73 | 0.717 | 0.222 | 87.8 | 97.2 | 0.556 | 0.322 | 0.514 |
| Sparse-view ( ) | |||||||||||||||
| DA3 | 19.69 | 0.667 | 0.300 | 55.8 | 89.6 | 0.369 | 0.471 | 20.86 | 0.646 | 0.287 | 88.3 | 98.8 | 0.402 | 0.452 | 0.112 |
| RE10K | DL3DV | ||||||||||||
| Input cameras | Appearance | Depth | Appearance | Depth | |||||||||
| Method | Intr. | Extr. | PSNR | SSIM | LPIPS | AbsRel | PSNR | SSIM | LPIPS | AbsRel | Runtime | ||
| (s) | |||||||||||||
| Sparse-view ( ) | |||||||||||||
| pixelSplat | 26.125 | 0.864 | 0.131 | 0.608 | 0.359 | 23.358 | 0.721 | 0.239 | 0.489 | 0.467 | 0.088 | ||
| MVSplat | 26.319 | 0.865 | 0.127 | 0.769 | 0.220 | 22.894 | 0.704 | 0.235 | 0.687 | 0.275 | 0.039 | ||
| RE10K | DL3DV | ||||||
| Method | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | Runtime |
| (s) | |||||||
| Sparse-view ( ) | |||||||
| NVComposer | 14.897 | 0.528 | 0.508 | 12.078 | 0.359 | 0.610 | 48.502 |
| Matrix3D | 20.934 | 0.700 | 0.284 | 18.149 | 0.523 | 0.376 | 25.592 |
| GLD | 23.608 | 0.768 | 0.219 | 21.440 | 0.621 | 0.270 | 76.824 |
| Methods | Appearance | Camera | Depth | |||||
|---|---|---|---|---|---|---|---|---|
| PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | |||
| ETH3D ( Schops et al., 2017 ) | ||||||||
| AnySplat | 15.99 | 0.462 | 0.489 | 27.3 | 89.4 | 0.767 | 0.148 | |
| + DensiTok | 17.01 | 0.531 | 0.466 | 33.3 | 90.3 | 0.771 | 0.135 | |
| AnySplat | 16.79 | 0.487 | 0.461 | 29.3 | 88.0 | 0.831 | 0.119 | |
| + DensiTok | 17.92 | 0.554 | 0.447 | 31.3 | 89.2 | 0.837 | 0.116 | |
| RE10K | DL3DV | Efficiency | ||||||
| Method | NFE | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | Runtime |
| (s) | ||||||||
| Deterministic regression | 1 | 20.86 | 0.712 | 0.246 | 19.97 | 0.631 | 0.288 | 0.258 |
| One-step training from scratch | 1 | 21.67 | 0.746 | 0.240 | 21.28 | 0.670 | 0.285 | 0.258 |
| FM pretraining only | 1 | 21.56 | 0.743 | 0.243 | 21.09 | 0.665 | 0.289 | 0.258 |
| 8 | 21.84 | 0.755 | 0.236 | 21.59 | 0.676 | 0.279 | 0.990 | |
| RE10K | DL3DV | ||||||||||||||
| Appearance | Camera | Depth | Appearance | Camera | Depth | ||||||||||
| PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | ||||
| 22.56 | 0.766 | 0.223 | 26.7 | 78.9 | 0.846 | 0.153 | 22.40 | 0.699 | 0.255 | 30.4 | 88.9 | 0.801 | 0.170 | ||
| 22.64 | 0.770 | 0.215 | 27.5 | 79.0 | 0.854 | 0.151 | 22.24 | 0.700 | 0.247 | 39.6 | 90.8 | 0.804 | 0.164 | ||
| 24.28 | 0.813 | 0.184 | 25.8 | 79.7 | 0.853 | 0.146 | 24.28 | 0.755 | 0.205 | 34.4 | 89.6 | 0.809 | 0.158 | ||
| 24.57 | 0.822 | 0.171 | 27.1 | 79.5 | 0.862 | 0.145 | 24.29 | 0.760 | 0.194 | 38.5 | 90.5 | 0.812 | 0.155 | ||
| RE10K | DL3DV | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Appearance | Camera | Depth | Appearance | Camera | Depth | ||||||||||
| Gaussian head | PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | PSNR | SSIM | LPIPS | AUC@ | AUC@ | AbsRel | |||
| Fully frozen | 22.14 | 0.754 | 0.233 | 26.4 | 78.8 | 0.833 | 0.161 | 22.01 | 0.685 | 0.267 | 30.6 | 88.8 | 0.790 | 0.179 | |
| Image embedding only | 22.56 | 0.766 | 0.223 | 26.7 | 78.9 | 0.846 | 0.153 | 22.40 | 0.699 | 0.255 | 30.4 | 88.9 | 0.801 | 0.170 | |
| Full finetuning | 22.88 | 0.775 | 0.216 | 26.8 | 78.8 | 0.854 | 0.148 | 22.69 | 0.708 | 0.248 | 30.2 | 88.8 | 0.807 | 0.166 | |
| Fully frozen | 24.07 | 0.806 | 0.190 | 25.6 | 78.8 | 0.846 | 0.150 | 24.05 | 0.747 | 0.212 | 34.1 | 88.8 | 0.803 | 0.163 | |