Recently, generalizable feed-forward methods based on 3D Gaussian Splatting have gained significant attention for their potential to reconstruct 3D scenes using finite resources. These approaches create a 3D radiance field, parameterized by per-pixel 3D Gaussian primitives, from just a few images in a single forward pass. However, unlike multi-view methods that benefit from cross-view correspondences, 3D scene reconstruction with a single-view image remains an underexplored area. In this work, we introduce CATSplat, a novel generalizable transformer-based framework designed to break through the inherent constraints in monocular settings. First, we propose leveraging textual guidance from a visual-language model to complement insufficient information from a single image. By incorporating scene-specific contextual details from text embeddings through cross-attention, we pave the way for context-aware 3D scene reconstruction beyond relying solely on visual cues. Moreover, we advocate utilizing spatial guidance from 3D point features toward comprehensive geometric understanding under single-view settings. With 3D priors, image features can capture rich structural insights for predicting 3D Gaussians without multi-view techniques. Extensive experiments on large-scale datasets demonstrate the state-of-the-art performance of CATSplat in single-view 3D scene reconstruction with high-quality novel view synthesis.
Figures & tables
Figure 1 : Overview of the generalizable 3D scene reconstruction pipeline. The feed-forward network creates a 3D radiance field using 3D Gaussians, all within an end-to-end differentiable system.
Figure 2 : We introduce CATSplat , a C ontext- A ware T ransformer with S patial Guidance for Generalizable 3D Gaussian Splatting from a single image. (a) Our two main priors, and (b) Examples of text descriptions (from the VLM) representing an input image.
Figure 3 : Overview of CATSplat framework. CATSplat takes an image I and predicts 3D Gaussian primitives {(μj,αj,Σj,cj)}jJ to construct a scene-representative 3D radiance field in a single forward pass. Our primary goal is to go beyond the finite knowledge inherent in single-view image features leveraging our two innovative priors. Through cross-attention layers, we enhance image features FiI to be highly informative by incorporating valuable insights: contextual cues from text features FiC , and spatial cues from 3D point features FiS .
Figure 4 : Detailed transformer pipeline. In the i -th layer, we first operate cross-attention between FiI and FiC , then proceed cross-attention with FiS . We also use a ratio γ to preserve visual information from FiI while incorporating extra cues from FiC and FiS .
n=5 (frames)
n=10 (frames)
n=Random (frames)
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
MPI [ 54 ]
27.10
0.870
–
24.40
0.812
–
23.52
0.785
–
BTS [ 60 ]
–
–
–
–
–
–
24.00
0.755
0.194
Splatter Image [ 51 ]
28.15
0.894
0.110
25.34
0.842
0.144
24.15
0.810
0.177
MINE [ 26 ]
28.45
0.897
0.111
25.89
0.850
0.150
24.75
0.820
0.179
Flash3D [ 50 ]
28.46
0.899
0.100
25.94
0.857
0.133
24.93
0.833
0.160
Table 1: Comparisons of Novel View Synthesis (NVS) performance with state-of-the-art single-view 3D reconstruction approaches on the RealEstate10K [ 72 ] dataset. Following the standard protocol from [ 26 , 50 ] , we evaluate NVS metrics on unseen target frames located n frames away from the input source frame. Also, we randomly sample an extra target frame within 30 frames apart from the source frame.
RE10K Interpolation
RE10K Extrapolation
Input
Method
Framework
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Two-View
pixelNeRF [ 65 ]
NeRF
20.51
0.592
0.550
20.05
0.575
0.567
Du et al . [ 14 ]
NeRF
24.78
0.820
0.213
21.83
0.790
0.242
pixelSplat [ 8 ]
3DGS
26.09
0.864
0.136
21.84
0.777
0.216
latentSplat [ 58 ]
3DGS
23.93
0.812
0.164
22.62
0.777
0.196
MVSplat [ 10 ]
3DGS
26.39
0.869
0.128
23.04
0.813
0.185
Table 2: Comparisons of NVS performance with state-of-the-art few-view 3D reconstruction approaches on the RealEstate10K [ 72 ] . Although we mainly focus on comparing with the leading single-view method, Flash3D [ 50 ] , we also provide scores of two-view methods for additional references. Following Flash3D, we use interpolation and extrapolation protocols from previous works, [ 8 ] and [ 58 ] , respectively.
Cross Dataset
Method
PSNR ↑
SSIM ↑
LPIPS ↓
RE10K → NYUv2
Flash3D [ 50 ]
25.09
0.775
0.182
CATSplat (Ours)
25.57
0.781
0.157
RE10K → ACID
Flash3D [ 50 ]
24.28
0.730
0.263
CATSplat (Ours)
24.73
0.739
0.250
RE10K → KITTI
Flash3D [ 50 ]
21.96
0.826
0.132
CATSplat (Ours)
22.43
0.833
0.122
Table 3 : Comparisons of cross-dataset generalization with the state-of-the-art single-view 3DGS method, Flash3D [ 50 ] , on various real-world datasets: NYUv2 [ 47 ] , ACID [ 29 ] , and KITTI [ 17 ] .
Figure 5 : Ablation study to see the effect of iteratively incorporating our novel priors on the RE10K [ 72 ] ( n = Random ). For clear ablations, we keep the number of entire transformer layers consistent across the experiments and adjust only the number of cross-attentions (CA).
n=10 (frames)
n=Random (frames)
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline
26.04
0.857
0.132
25.02
0.834
0.159
w/ Contextual
26.40
0.864
0.127
25.40
0.838
0.153
w/ Spatial
26.38
0.864
0.127
25.42
0.837
0.153
CATSplat
26.44
0.866
0.125
25.45
0.841
0.151
Table 4 : Ablation study to explore the effect of our two intelligent priors (Contextual and Spatial) across three different settings, as in Tab. 1 , on the RE10K [ 72 ] dataset. Here, the “Baseline” indicates our basic transformer architecture without any proposed priors.
Figure 6 : Qualitative comparisons of NVS performance between Flash3D [ 50 ] and ours with Ground Truth on the novel view frames from RealEstate10K [ 72 ] and ACID [ 29 ] (cross-dataset). We provide more visual results and details of user study in the supplementary material.
n=10 (frames)
n=Random (frames)
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
w/ ConvNeXt-B
26.15
0.856
0.132
25.09
0.832
0.158
w/ ConvNeXt-L
26.17
0.857
0.132
25.12
0.833
0.157
w/ DINOv2-B
26.17
0.858
0.131
25.11
0.833
0.157
w/ DINOv2-g
26.19
0.859
0.131
25.17
0.834
0.156
w/ Contextual
26.40
0.864
0.127
25.40
0.838
0.153
Table 5 : Ablation study to see the effect of our novel priors versus visual priors from large-scale pre-trained image encoders, such as DINOv2 [ 39 ] and ConvNeXt V2 [ 61 ] , on the RE10K [ 72 ] dataset.
n=10 (frames)
n=Random (frames)
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline
26.04
0.857
0.132
25.02
0.834
0.159
w/ Scene Type
26.14
0.859
0.130
25.13
0.835
0.158
w/ Object List
26.23
0.862
0.128
25.25
0.836
0.155
w/ Extended
26.31
0.862
0.128
25.29
0.837
0.154
w/ Single Sent.
26.40
0.864
0.127
25.40
0.838
0.153
Table 6 : Ablation study to see the impact of different text description formats on generalizable tasks. The “Baseline” is as in Tab. 4 .
n=10 (frames)
n=Random (frames)
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline
26.04
0.857
0.132
25.02
0.834
0.159
w/o Depth Conc.
25.91
0.855
0.134
24.82
0.827
0.165
w/ Point Conc.
26.06
0.857
0.132
25.04
0.834
0.158
w/ Depth Feat.
26.18
0.859
0.130
25.16
0.835
0.157
w/ Point Feat.
26.38
0.864
0.127
25.42
0.837
0.153
Table 7 : Ablation study to explore strategies for enriching geometric knowledge from a single image. The “Baseline” is as in Tab. 4 .
RE10K [ 72 ]
ACID [ 29 ]
Method
Preference ( % )
Likert ↑
Preference ( % )
Likert ↑
Flash3D [ 50 ]
11.58 ± 1.09
4.56 ± 0.30
8.59 ± 0.63
4.14 ± 0.21
CATSplat (Ours)
88.42 ± 1.09
6.04 ± 0.22
91.41 ± 0.63
5.27 ± 0.18
Table 8 : User study comparisons. We report mean preference percentage and a 7-point Likert scale with a 95% confidence interval.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1 : Examples of two types of user study questions. The first type of question (above) asks about preference between ours and Flash3D [ 50 ] , and the second (below) requires participants to rate the visual quality of the rendered image compared to the target.
Figure 2 : Detailed architecture of 3D point feature extraction from a monocular input image I . Our point cloud encoder takes back-projected points P and produces point features FS based on the PointNet [ 42 ] structure. Here, T-Net indicates an affine transform network.
Figure 3 : Examples of input images with their corresponding estimated depth maps and back-projected 3D point clouds. For better visualization, we also present 3D point clouds with RGB colors.
Method
Training (hrs)
Inference (secs)
n=Random (frames)
PSNR ↑
SSIM ↑
LPIPS ↓
Flash3D [ 50 ]
23.2
0.327
24.93
0.833
0.160
CATSplat (Ours)
29.1
0.393
25.45
0.841
0.151
Appendix
Table 1 : Comparison of computational costs with the state-of-the-art single-view 3DGS method, Flash3D [ 50 ] , on the RE10K [ 72 ] .
Method
n=Random (frames)
Baseline
Contextual
Spatial
PSNR ↑
SSIM ↑
LPIPS ↓
✓
-
-
25.11
0.775
0.178
✓
✓
-
25.51
0.779
0.163
✓
-
✓
25.48
0.778
0.165
✓
✓
✓
25.57
0.781
0.157
Appendix
Table 2 : Ablation study to investigate the effect of our two intelligent priors on the NYUv2 [ 47 ] dataset in cross-dataset settings.
Method
n=Random (frames)
Baseline
Contextual
Spatial
PSNR ↑
SSIM ↑
LPIPS ↓
✓
-
-
24.26
0.732
0.261
✓
✓
-
24.57
0.735
0.253
✓
-
✓
24.62
0.737
0.254
✓
✓
✓
24.73
0.739
0.250
Appendix
Table 3 : Ablation study to investigate the effect of our two intelligent priors on the ACID [ 72 ] dataset in cross-dataset settings.
Figure 4 : Examples of four different formats of text descriptions from the VLM [ 31 ] , as described in Tab.5 in the main paper.
Figure 5 : Ablation study to see the effect of iteratively incorporating our novel priors on the RE10K [ 72 ] ( n = Random ). For clear ablations, we keep the number of entire transformer layers consistent across the experiments and adjust only the number of cross-attentions (CA).
n=10 (frames)
n=Random (frames)
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
OpenFlamingo
26.08
0.858
0.131
25.06
0.832
0.158
BLIP2 T5
26.29
0.860
0.129
25.27
0.833
0.156
LLaVA 7B
26.19
0.861
0.129
25.23
0.834
0.156
LLaVA 13B
26.40
0.864
0.127
25.40
0.838
0.153
Appendix
Table 4 : Ablation study to see the impact of text embeddings from different VLMs: OpenFlamingo [ 5 ] , BLIP2 [ 27 ] , and LLaVA [ 31 ] .
Figure 6 : Qualitative comparisons with various novel view synthesis approaches, including MASt3R [ 25 ] , Splatt3R [ 49 ] , MoGe [ 56 ] , and ZeroNVS [ 45 ] . All methods are fairly evaluated on unseen scenes in cross-dataset settings across ACID [ 29 ] and NYUv2 [ 47 ] datasets.
Figure 7 : Qualitative comparisons of 3D reconstruction between Flash3D [ 50 ] and ours with Ground Truth. We visualize zoom-in views of 3D Gaussian distributions and depth maps from them.
Figure 8 : Failure cases of CATSplat. When invisible areas in the input become visible in the target, ours might be less productive.
Figure 9 : Qualitative comparisons between Flash3D [ 50 ] and Ours with Input Image and Ground Truth on the RealEstate10K [ 72 ] dataset.
Figure 10 : Qualitative comparisons between Flash3D [ 50 ] and Ours with Input Image and Ground Truth on the RealEstate10K [ 72 ] dataset.
Figure 11 : Qualitative comparisons between Flash3D [ 50 ] and Ours with Input Image and Ground Truth on the ACID [ 29 ] dataset.
Figure 12 : Qualitative comparisons between Flash3D [ 50 ] and Ours with Input Image and Ground Truth on the KITTI [ 17 ] dataset.
Feed-forward 3D Gaussian Splatting (3DGS) enables efficient and generalizable 3D reconstruction, but current feed-forward 3DGS methods for scene understanding remain largely category-oriented. In contrast, instance-aware 3DGS methods typically rely on per-scene optimization and often decouple reconstruction from instance and semantic learning, limiting reciprocal interactions among them. We present InstanceSplat, a unified feed-forward 3DGS framework for generalizable 3D reconstruction and instance-aware scene understanding from pose-free multi-view images. In a single forward pass, InstanceSplat constructs an instance-aware Gaussian representation that jointly encodes appearance, geometry, instance identity, and language-aligned semantics. Shared 3D Gaussians ground instance identities across views, producing renderable and cross-view-consistent instance features. To allow reconstruction and scene understanding to benefit from each other, we further design an instance-centric learning strategy that connects reconstruction, instance learning, and semantic learning through shared instance structure. Specifically, instance cues guide reconstruction, language-aligned semantics strengthen the discrimination of confusing same-category instances, and instance regions aggregate semantic evidence into coherent object-level predictions. Experiments on novel-view synthesis, instance segmentation, and open-vocabulary semantic understanding under varying input-view settings and on an unseen dataset demonstrate state-of-the-art performance, practical efficiency, and strong generalization.
Minchao Jiang, Xiaoxuan Ma, Shunyu Jia +3
1Shanghai Jiao Tong University · 2Eastern Institute of Technology, Ningbo · 3Carnegie Mellon University +2
Reconstructing 3D scenes from sparse, unposed images remains challenging under real-world conditions with varying illumination and transient occlusions. Existing methods rely on scene-specific optimization using appearance embeddings or dynamic masks, which requires extensive per-scene training and fails under sparse views. Moreover, evaluations on limited scenes raise questions about generalization. We present GenWildSplat, a feed-forward framework for sparse-view outdoor reconstruction that requires no per-scene optimization. Given unposed internet images, GenWildSplat predicts depth, camera parameters, and 3D Gaussians in a canonical space using learned geometric priors. An appearance adapter modulates appearance for target lighting conditions, while semantic segmentation handles transient objects. Through curriculum learning on synthetic and real data, GenWildSplat generalizes across diverse illumination and occlusion patterns. Evaluations on PhotoTourism and MegaScenes benchmark demonstrate state-of-the-art feed-forward rendering quality, achieving real-time inference without test-time optimization
Vinayak Gupta, Chih-Hao Lin, Shenlong Wang +2
University of Maryland, College Park · University of Illinois Urbana-Champaign · Johns Hopkins University
We introduce S2C-3D, a novel sparse-view 3D reconstruction framework for high-fidelity and complete scene reconstruction from as few as six to eight images. Our framework features three components: a specialized diffusion model for scene-specific image restoration, a training-free view-consistency conditioned sampling process in the diffusion model for refined Gaussian optimization, and a camera trajectory planning scheme to ensure comprehensive scene coverage. The specialized diffusion model is developed by finetuning a pretrained architecture on the input views and their corresponding degraded counterparts. The adaptation to the scene distribution allows the model to repair Gaussian renderings while effectively eliminating domain gaps. Meanwhile, the trajectory planning scheme optimizes scene coverage by connecting each newly sampled camera to its two nearest neighbors. By iteratively constructing paths and retaining only those that significantly enhance visibility, the scheme establishes a trajectory that covers the entire scene. To address multi-view conflicts, the view-consistency conditioned sampling process quantifies the consistency between neighboring repaired images. This information is injected as a condition into the sampling process of the frozen diffusion model, facilitating the generation of view-consistent images without additional training. Consequently, our approach produces high-fidelity 3D Gaussians that are robust to artifacts. Experimental results demonstrate that S2C-3D outperforms state-of-the-art methods, constructing high-quality scenes that are free from missing regions, blurring, or other artifacts with very sparse inputs. The source code and data are available at https://gapszju.github.io/S2C-3D.
Yiyang Shen, Yin Yang, Kun Zhou +1
State Key Lab of CAD&CG, Zhejiang University, China · University of Utah, USA · State Key Lab of CAD&CG, Zhejiang University, China and Hangzhou Research Institute of Holographic and AI Technology, China +1