Feed-forward 3D Gaussian Splatting (3DGS) enables reconstruction without per- scene optimisation, but practical stereo-camera applications require nearby-view extrapolation beyond the input views. Stereo depth anchors visible surfaces, yet rendering newly exposed regions also requires learned appearance and additional scene capacity. We introduce StereoGaussians, which predicts a metric 3DGS representation from a single calibrated stereo pair. It reuses intermediate repre- sentations from frozen pretrained stereo networks to predict Gaussian attributes, while calibrated disparity anchors the geometry. A second Gaussian layer and an expanded image canvas provide capacity for disoccluded and outside-field-of- view content. For training, we construct SceneSplat-Stereo from quality-filtered 3DGS teachers, pairing stereo inputs with nearby target views across 803 training scenes. Experiments on unseen real and photorealistic stereo benchmarks demon- strate improvements over strong view-synthesis baselines, while ablation studies support our main design choices.
Figures & tables
Figure 1: SceneSplat-Stereo synthetic data generation. (a) Camera sampling. A quality-filtered real-scene 3DGS teacher renders a calibrated stereo pair and ten nearby targets around a reference pose from the source trajectory. (b) Generated sample. Stereo inputs and target images are accompanied by intrinsics, relative target poses, and opacity-derived valid masks. Camera placement is schematic and illustrates the relative arrangement of input and target views.
Data source
Scenes / images
Stereo input
Target-view sup.
Static scenes
Real appearance
RealEstate10K ( Zhou et al., 2018 )
80K / 10M
✓
✓
✓
DL3DV-10K ( Ling and others, 2024 )
10.5K / 51.2M
✓
✓
✓
Middlebury v3 ( Scharstein et al., 2014 )
30 / 60
✓
✓
✓
Flickr1024 ( Wang et al., 2019b )
1K / 2K
✓
✓
✓
InStereo2K ( Bao et al., 2020 )
– / 4.1K
✓
✓
✓
KITTI ( Geiger et al., 2012 )
789 / 1.6K
✓
✓
Table 1: Dataset comparison for feed-forward stereo-to-3DGS prediction. Scale reports scenes (or sequences) / input images, counting both stereo views or individual monocular frames; separately rendered targets are excluded. See Appendix B for details on dataset-specific counting conventions and the supervision available from each source.
Figure 2: StereoGaussians overview. (a) Stereo geometry and representation. Frozen stereo features and calibrated depth condition Gaussian prediction. (b) Extrapolation-aware representation. Two layers and an expanded canvas support disoccluded and outside-FoV content. (c) Extrapolation and supervision. Target-view rendering trains the decoder and attribute predictor. Scene and Gaussian illustrations are schematic.
StereoNVS
ReplicaGS
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Single-view feed-forward 3DGS
SHARP-mono
18.845
0.6140
0.2604
23.388
0.7877
0.2103
SHARP-stereo
18.767
0.6030
0.2788
23.363
0.7808
0.2164
Sparse-view feed-forward 3DGS
MVSplat
17.458
0.7248
0.2087
14.705
0.5342
0.4152
Table 2: Zero-shot comparison using the default 16-pixel padding. Our model is trained on SceneSplat-Stereo. Baselines retain their respective pretraining. Bold indicates the best score in each column for the corresponding dataset and evaluation metric.
Figure 3: Qualitative novel-view comparison on StereoNVS (top two rows) and ReplicaGS (bottom two rows). Each row shows the input stereo pair, predictions from representative methods SHARP, DepthSplat, and LVSM, StereoGaussians (Ours), and ground truth (GT). Prediction panels report PSNR in dB.
StereoNVS
ReplicaGS
Training dataset / model
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
StereoNVS
22.707
0.7585
0.2024
23.929
0.8914
0.1935
IRS
19.656
0.6546
0.2412
24.614
0.8421
0.1873
SceneSplat-Stereo
21.261
0.7549
0.1653
30.266
0.9447
0.0818
DepthSplat (feature init.)
16.063
0.4861
0.4018
21.764
0.7026
0.2350
DepthSplat (RE10K FT)
21.152
0.7313
0.2057
28.996
0.9206
0.0911
Table 3: Training-data comparison for StereoGaussians (top) and DepthSplat on SceneSplat-Stereo (bottom). Feature init.: DA-V2 + UniMatch; FT: RE10K fine-tuning. StereoNVS-to-StereoNVS evaluation is in-domain. Bold denotes the best out-of-domain result in each column.
StereoNVS
ReplicaGS
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Shared default
21.261
0.7549
0.1653
30.266
0.9447
0.0818
(A) Input conditioning
Without descriptor
20.752
0.7417
0.1871
28.284
0.9362
0.1045
Without RGB-D conditioning
21.269
0.7594
0.1717
29.953
0.9391
0.0915
(B) Frozen stereo encoder
Table 4: Ablations of input conditioning, frozen stereo encoders, Gaussian prediction, and FoV expansion under a shared training and zero-shot evaluation protocol. Each block varies the default in § 5.1 ; variant definitions are given in § 5.4 – 5.5 .
Figure 4: Qualitative ablations. Each row shows the input stereo pair, the full model, variants without the feature descriptor, back layer, composer, or FoV padding, and ground truth (GT). Red outlines highlight missing content and distortions near occlusion boundaries and in newly exposed image-boundary regions; corresponding full-model and GT regions provide visual references. Prediction panels report PSNR in dB.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Subset
Split
Scenes
Stereo pairs
Target views
Teacher PSNR
ScanNetGS
Train / Val
491 / 30
175K / 150
1.75M / 600
31.69
ScanNetPPGS-V2
Train / Val
312 / 20
149K / 100
1.49M / 400
33.04
Total
Train / Val
803 / 50
324K / 250
3.24M / 1K
–
ReplicaGS
Test
8
800
8K
41.21
Appendix
Table 5: Dataset composition. Slash-separated counts denote train / val; training and validation are scene-disjoint. ReplicaGS is test-only. Pair and target counts describe the stored training pool or the selected validation/test views; training draws four of the ten stored targets per pair. Counts with K/M are rounded. Teacher PSNR averages over the quality-filtered source pool (train and validation together), or the eight ReplicaGS test scenes.
Teacher PSNR
B (m)
ρ (m)
Train scenes
Val scenes
(30,35) dB
0.05
0.10
473 / 257
28 / 14
[35,40) dB
0.10
0.20
18 / 48
2 / 6
[40,∞) dB
0.10
0.30
0 / 7
0 / 0
Appendix
Table 6: Teacher-quality tiers, stereo baselines, target radii, and scene counts. Scene counts are ordered as ScanNetGS / ScanNetPPGS-V2.
StereoNVS
ReplicaGS
Resolution
Pad
Batch
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
2562 (default)
16
32(1×32)
21.261
0.7549
0.1653
30.266
0.9447
0.0818
2562 (larger FoV)
64
32(1×32)
23.189
0.7689
0.1612
30.165
0.9443
0.0823
5122
32
24(2×12)
20.414
0.7090
0.2274
29.676
0.9306
0.1061
10242
64
32(8×4)
19.861
0.7183
0.2935
29.615
0.9324
0.1350
Appendix
Table 7: Higher-resolution reference results using the same model architecture. Each predictor is trained from scratch for 30K steps and evaluated at its training resolution, with a frozen pretrained stereo encoder. Batch denotes effective batch size (GPUs × per-GPU batch). Padding scales with resolution except for the additional 2562 / pad-64 reference.
Figure 5: Additional baseline comparisons on StereoNVS (1/2). Columns show the left and right inputs, SHARP, DepthSplat, LVSM, StereoGaussians (Ours), and ground truth (GT). Prediction panels report PSNR in dB. In these lobby, plant, outdoor-seating, and sofa examples, our model better preserves furniture contours and the separation between foreground objects and their surroundings.
Figure 6: Additional baseline comparisons on StereoNVS (2/2). Column order and score annotations follow Figure 5 . These examples include wooden desks, closely spaced chairs and tables, a cleaning cart, and small decorative objects. Our predictions retain clearer object boundaries and more of the small-scale structure in the cart and decorations.
Figure 7: Additional baseline comparisons on ReplicaGS (1/2). Columns show the stereo inputs, SHARP, DepthSplat, LVSM, StereoGaussians (Ours), and GT, with per-example PSNR in dB. The display panels, cushions, floor patterns, and sofa boundaries are more faithfully reproduced by our model in these examples. Several baseline predictions exhibit displaced structures, blurred textures, or missing content near image boundaries.
Figure 8: Additional baseline comparisons on ReplicaGS (2/2). Column order and score annotations follow Figure 7 . Examples cover tabletops, wall partitions, a furnished living room, framed pictures, and a dining area. Our model better maintains straight architectural edges and furniture placement, while reducing the distortions and missing border regions visible in several baseline renderings.
Figure 9: Qualitative ablations on StereoNVS (1/2). Columns show the stereo inputs, full model, four ablated variants, and GT. Prediction panels report PSNR in dB. Removing the descriptor or back layer degrades table and chair boundaries in the outdoor examples. Removing padding leaves uncovered border regions, particularly in the sofa views.
Figure 10: Qualitative ablations on StereoNVS (2/2). The layout follows Figure 9 . The desk and chair examples expose missing regions and distorted boundaries in the ablated predictions. The cleaning cart and small decorations show changes in fine structure and appearance. The full model reduces these artifacts and better preserves the depicted objects’ structure.
Figure 11: Qualitative ablations on ReplicaGS (1/2). Columns follow Figure 9 , with per-example PSNR in dB. Removing the descriptor reduces fidelity around cushions and floor patterns. Removing the back layer can introduce missing content, including the upper region of the bright display scene. Without padding, black bands appear at newly exposed image boundaries in several views.
Figure 12: Qualitative ablations on ReplicaGS (2/2). The layout follows Figure 9 . The office-chair and dining-table views reveal gaps around foreground objects when the descriptor or back layer is removed. The no-padding variant additionally leaves large uncovered regions at the target image borders. The full model improves coverage while retaining clearer furniture and wall boundaries.
Recent advances in 3D Gaussian Splatting (3DGS) have enabled high-quality, render-ready scene representations for novel-view synthesis. However, most existing 3DGS pipelines rely on multi-view observations (or non-causal access to future frames) to achieve sufficient coverage, which is often unavailable in on-device robotics and AR settings where sensing is restricted to a single stereo rig. Recovering a high-quality 3DGS scene from one stereo observation, therefore, remains challenging due to occlusions, limited field of view, and missing geometry. We present StereoSplat+, a diffusion-enhanced feed-forward framework that enables causal reconstruction from a single stereo pair. Our method builds on two key components. First, we propose StereoSplat, an input-invariant feed-forward 3D Gaussian estimator that takes a variable number of posed stereo pairs as input and predicts high-quality 3D Gaussians. StereoSplat fuses complementary geometry cues via a cost-volume branch and a triplane-based 3D volume branch and leverages continuous pose encoding to generalize across view counts and camera configurations. Second, since multiple posed stereo pairs are typically unavailable at inference time, we introduce a diffusion-enhanced one-shot progressive inference scheme called StereoSplat+: starting from one stereo pair, we render novel stereo views from the predicted 3DGS, refine them with a one-step diffusion enhancer, and feed them back as additional inputs to update the 3DGS. Experiments on the KITTI-360 dataset show that StereoSplat+ improves novel-view rendering quality and geometry accuracy, especially in occluded regions and under strong view extrapolation, outperforming recent feed-forward 3DGS baselines.
Zihua Liu, Masatoshi Okutomi
Department of systems and Control Engineering, Institute of Science Tokyo, Japan.
3D Gaussian Splatting (3DGS) has achieved remarkable success in real-time novel view synthesis, yet it suffers from severe overfitting under sparse-view settings due to insufficient geometric constraints. While recent methods introduce monocular depth priors to mitigate this, they inherently struggle with scale ambiguity and cross-view inconsistency, leading to defective geometry. In this paper, we propose StereoGS, a novel sparse-view 3DGS framework that integrates stereo priors to establish reliable binocular consistency. Unlike scale-agnostic monocular constraints, StereoGS introduces a Stereo Depth Regularization by constructing virtual stereo pairs during optimization and leveraging a foundation stereo model to enforce absolute scale and binocular-consistent structures. To further suppress overfitting and eliminate redundant primitives, we design a Gradient-Aware Opacity Decay strategy that dynamically penalizes Gaussians based on their relative opacity gradient magnitudes. Combined with a Consistency-Aware Dense Initialization using zero-shot multi-view depth estimation, StereoGS effectively anchors primitives to accurate scene surfaces. Extensive experiments on LLFF, DTU, Mip-NeRF360, and Blender datasets demonstrate that StereoGS achieves state-of-the-art performance in sparse-view settings without incurring any additional inference overhead. Project Page: https://stringerywh00.github.io/StereoGS_project_page/
In this work, we revisit several key design choices of modern Transformer-based approaches for feed-forward 3D Gaussian Splatting (3DGS) prediction. We argue that the common practice of regressing Gaussian means as depths along camera rays is suboptimal, and instead propose to directly regress 3D mean coordinates using only a self-supervised rendering loss. This formulation allows us to move from the standard encoder-only design to an encoder-decoder architecture with learnable Gaussian tokens, thereby unbinding the number of predicted primitives from input image resolution and number of views. Our resulting method, TokenGS, demonstrates improved robustness to pose noise and multiview inconsistencies, while naturally supporting efficient test-time optimization in token space without degrading learned priors. TokenGS achieves state-of-the-art feed-forward reconstruction performance on both static and dynamic scenes, producing more regularized geometry and more balanced 3DGS distribution, while seamlessly recovering emergent scene attributes such as static-dynamic decomposition and scene flow.
Jiawei Ren, Michal Jan Tyszkiewicz, Jiahui Huang +1