CATSplat: Context-Aware Transformer with Spatial Guidance for Generalizable 3D Gaussian Splatting from A Single-View Image
Organizations: Korea University · Google · Purdue University
Abstract
Recently, generalizable feed-forward methods based on 3D Gaussian Splatting have gained significant attention for their potential to reconstruct 3D scenes using finite resources. These approaches create a 3D radiance field, parameterized by per-pixel 3D Gaussian primitives, from just a few images in a single forward pass. However, unlike multi-view methods that benefit from cross-view correspondences, 3D scene reconstruction with a single-view image remains an underexplored area. In this work, we introduce CATSplat, a novel generalizable transformer-based framework designed to break through the inherent constraints in monocular settings. First, we propose leveraging textual guidance from a visual-language model to complement insufficient information from a single image. By incorporating scene-specific contextual details from text embeddings through cross-attention, we pave the way for context-aware 3D scene reconstruction beyond relying solely on visual cues. Moreover, we advocate utilizing spatial guidance from 3D point features toward comprehensive geometric understanding under single-view settings. With 3D priors, image features can capture rich structural insights for predicting 3D Gaussians without multi-view techniques. Extensive experiments on large-scale datasets demonstrate the state-of-the-art performance of CATSplat in single-view 3D scene reconstruction with high-quality novel view synthesis.
Figures & tables
| (frames) | (frames) | (frames) | |||||||||
| Method | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | ||
| MPI [ 54 ] | 27.10 | 0.870 | – | 24.40 | 0.812 | – | 23.52 | 0.785 | – | ||
| BTS [ 60 ] | – | – | – | – | – | – | 24.00 | 0.755 | 0.194 | ||
| Splatter Image [ 51 ] | 28.15 | 0.894 | 0.110 | 25.34 | 0.842 | 0.144 | 24.15 | 0.810 | 0.177 | ||
| MINE [ 26 ] | 28.45 | 0.897 | 0.111 | 25.89 | 0.850 | 0.150 | 24.75 | 0.820 | 0.179 | ||
| Flash3D [ 50 ] | 28.46 | 0.899 | 0.100 | 25.94 | 0.857 | 0.133 | 24.93 | 0.833 | 0.160 | ||
| RE10K Interpolation | RE10K Extrapolation | ||||||||
| Input | Method | Framework | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS | |
| Two-View | pixelNeRF [ 65 ] | NeRF | 20.51 | 0.592 | 0.550 | 20.05 | 0.575 | 0.567 | |
| Du et al . [ 14 ] | NeRF | 24.78 | 0.820 | 0.213 | 21.83 | 0.790 | 0.242 | ||
| pixelSplat [ 8 ] | 3DGS | 26.09 | 0.864 | 0.136 | 21.84 | 0.777 | 0.216 | ||
| latentSplat [ 58 ] | 3DGS | 23.93 | 0.812 | 0.164 | 22.62 | 0.777 | 0.196 | ||
| MVSplat [ 10 ] | 3DGS | 26.39 | 0.869 | 0.128 | 23.04 | 0.813 | 0.185 | ||
| Cross Dataset | Method | PSNR | SSIM | LPIPS |
| RE10K NYUv2 | Flash3D [ 50 ] | 25.09 | 0.775 | 0.182 |
| CATSplat (Ours) | 25.57 | 0.781 | 0.157 | |
| RE10K ACID | Flash3D [ 50 ] | 24.28 | 0.730 | 0.263 |
| CATSplat (Ours) | 24.73 | 0.739 | 0.250 | |
| RE10K KITTI | Flash3D [ 50 ] | 21.96 | 0.826 | 0.132 |
| CATSplat (Ours) | 22.43 | 0.833 | 0.122 |
| (frames) | (frames) | |||||
| Method | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS |
| Baseline | 26.04 | 0.857 | 0.132 | 25.02 | 0.834 | 0.159 |
| w/ Contextual | 26.40 | 0.864 | 0.127 | 25.40 | 0.838 | 0.153 |
| w/ Spatial | 26.38 | 0.864 | 0.127 | 25.42 | 0.837 | 0.153 |
| CATSplat | 26.44 | 0.866 | 0.125 | 25.45 | 0.841 | 0.151 |
| (frames) | (frames) | |||||
| Method | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS |
| w/ ConvNeXt-B | 26.15 | 0.856 | 0.132 | 25.09 | 0.832 | 0.158 |
| w/ ConvNeXt-L | 26.17 | 0.857 | 0.132 | 25.12 | 0.833 | 0.157 |
| w/ DINOv2-B | 26.17 | 0.858 | 0.131 | 25.11 | 0.833 | 0.157 |
| w/ DINOv2-g | 26.19 | 0.859 | 0.131 | 25.17 | 0.834 | 0.156 |
| w/ Contextual | 26.40 | 0.864 | 0.127 | 25.40 | 0.838 | 0.153 |
| (frames) | (frames) | |||||
| Method | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS |
| Baseline | 26.04 | 0.857 | 0.132 | 25.02 | 0.834 | 0.159 |
| w/ Scene Type | 26.14 | 0.859 | 0.130 | 25.13 | 0.835 | 0.158 |
| w/ Object List | 26.23 | 0.862 | 0.128 | 25.25 | 0.836 | 0.155 |
| w/ Extended | 26.31 | 0.862 | 0.128 | 25.29 | 0.837 | 0.154 |
| w/ Single Sent. | 26.40 | 0.864 | 0.127 | 25.40 | 0.838 | 0.153 |
| (frames) | (frames) | |||||
| Method | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS |
| Baseline | 26.04 | 0.857 | 0.132 | 25.02 | 0.834 | 0.159 |
| w/o Depth Conc. | 25.91 | 0.855 | 0.134 | 24.82 | 0.827 | 0.165 |
| w/ Point Conc. | 26.06 | 0.857 | 0.132 | 25.04 | 0.834 | 0.158 |
| w/ Depth Feat. | 26.18 | 0.859 | 0.130 | 25.16 | 0.835 | 0.157 |
| w/ Point Feat. | 26.38 | 0.864 | 0.127 | 25.42 | 0.837 | 0.153 |
| RE10K [ 72 ] | ACID [ 29 ] | ||||
| Method | Preference ( ) | Likert | Preference ( ) | Likert | |
| Flash3D [ 50 ] | 11.58 1.09 | 4.56 0.30 | 8.59 0.63 | 4.14 0.21 | |
| CATSplat (Ours) | 88.42 1.09 | 6.04 0.22 | 91.41 0.63 | 5.27 0.18 | |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Training (hrs) | Inference (secs) | (frames) | ||
| PSNR | SSIM | LPIPS | |||
| Flash3D [ 50 ] | 23.2 | 0.327 | 24.93 | 0.833 | 0.160 |
| CATSplat (Ours) | 29.1 | 0.393 | 25.45 | 0.841 | 0.151 |
| Method | (frames) | ||||
| Baseline | Contextual | Spatial | PSNR | SSIM | LPIPS |
| - | - | 25.11 | 0.775 | 0.178 | |
| - | 25.51 | 0.779 | 0.163 | ||
| - | 25.48 | 0.778 | 0.165 | ||
| 25.57 | 0.781 | 0.157 | |||
| Method | (frames) | ||||
| Baseline | Contextual | Spatial | PSNR | SSIM | LPIPS |
| - | - | 24.26 | 0.732 | 0.261 | |
| - | 24.57 | 0.735 | 0.253 | ||
| - | 24.62 | 0.737 | 0.254 | ||
| 24.73 | 0.739 | 0.250 | |||
| (frames) | (frames) | |||||
| Method | PSNR | SSIM | LPIPS | PSNR | SSIM | LPIPS |
| OpenFlamingo | 26.08 | 0.858 | 0.131 | 25.06 | 0.832 | 0.158 |
| BLIP2 T5 | 26.29 | 0.860 | 0.129 | 25.27 | 0.833 | 0.156 |
| LLaVA 7B | 26.19 | 0.861 | 0.129 | 25.23 | 0.834 | 0.156 |
| LLaVA 13B | 26.40 | 0.864 | 0.127 | 25.40 | 0.838 | 0.153 |