Remote sensing novel view synthesis under sparse observations remains challenging due to insufficient geometric constraints and limited cross-view supervision. Existing Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) methods are prone to overfitting and face challenges of depth ambiguities, missing cross-view information, and insufficient constraints in under-observed regions. To address these challenges, we propose DIBR-GS, a neural Gaussian Splatting framework that exploits Depth Image-Based Rendering (DIBR) to generate pseudo views for cross-view consistency supervision. Specifically, reliable geometric initialization is constructed by aligning monocular depth priors with sparse SfM reconstruction, and cross-view appearance priors are incorporated into neural Gaussian representations to enhance appearance modeling under sparse observations. Furthermore, we introduce a progressive DIBR-based pseudo-view supervision strategy to provide additional geometric and appearance constraints, enabling more complete reconstruction of weakly observed regions. In addition, a height-constrained anchor growth strategy is designed to suppress unreasonable Gaussian expansion. Experiments demonstrate that the proposed method achieves superior performance over existing approaches when training with only 3 input views. Compared with the previous best-performing method, it improves PSNR by 6.83 dB, with relative gains of 14% in SSIM and 60% in LPIPS, while maintaining competitive computational efficiency. Our code is available at https://github.com/kanehub/DIBR-GS
Figures & tables
Fig. 1: Visual and quantitative comparison on the LEVIR-NVS dataset with 3 input views. The proposed method produces more complete structures and finer details compared with existing approaches, especially in weakly observed regions. The radar chart summarizes the performance comparison in terms of reconstruction quality and rendering efficiency. For visualization, we report the negative LPIPS, AVGE, and 1+log(FPS) values to enable unified comparison.
Fig. 2: Overview of the proposed DIBR-GS framework. The framework consists of a DIBR branch and a neural Gaussian branch. In the DIBR branch, monocular depth priors are aligned with sparse SfM reconstruction through affine transformation and multi-view fusion to construct a reliable geometric initialization. Cross-view appearance features are extracted following the image-based rendering paradigm, and progressive DIBR is further employed to generate pseudo views that provide additional supervision. In the neural Gaussian branch, Gaussian anchors are initialized from the dense point cloud, while the anchor features and cross-view appearance features are adaptively integrated to predict Gaussian attributes. The rendered training views and pseudo views are jointly optimized, enabling more complete and consistent reconstruction under sparse-view remote sensing scenes.
Fig. 3: Illustration of anchor growth. (a) Anchor distribution after growth. The original anchors contain numerous floating artifacts and deviate significantly from the initial point cloud, while the Gaussians become more compact and precise with the constraint. (b) The proposed height-constrained anchor growth strategy, where only neural Gaussians whose heights fall within the predefined bounds are considered as candidate new anchors.
Fig. 4: Visual comparison on LEVIR-NVS Dataset.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
AVGE ↓
FPS ↑
RegNeRF [ 20 ]
19.83
0.695
0.389
0.131
0.19
FreeNeRF [ 21 ]
19.04
0.524
0.373
0.148
0.19
MPNeRF [ 4 ]
21.72
0.800
0.190
0.083
0.18
TriDF [ 52 ]
24.07
0.820
0.213
0.071
0.20
3DGS [ 16 ]
18.29
0.593
0.313
0.144
280
FSGS [ 27 ]
21.18
0.772
0.230
0.094
343
TABLE I: Quantitative comparison of rendering quality and efficiency between different methods. The best, second-best, and third-best entries are marked in , , and , respectively.
Fig. 5: Visualization of rendering quality when removing different model components.
Setting
PSNR ↑
SSIM ↑
LPIPS ↓
AVGE ↓
w/o Dense Init.
29.18
0.917
0.097
3.232
w/o IBR Feature
30.56
0.933
0.082
2.652
w/o Height Map
30.68
0.935
0.080
2.593
w/o Pseudo Views
21.06
0.782
0.186
8.795
Ours
30.90
0.938
0.076
2.488
TABLE III: Ablation study of different components. The AVGE values are multiplied by 1e2.
Fig. 6: Visualization of depth map and depth error with different initialization strategies.
Map
Source
Refine
MAE ↓
Abs Rel ↓
AVGE ↓
Neg.
SfM
–
3.72
0.0322
3.853
Inv.
SfM
–
3.16
0.0273
3.522
Inv.
Fused
–
3.31
0.0285
3.485
Inv.
SfM
Reproject
2.18
0.0187
3.406
Inv.
SfM
No ovlp.
1.91
0.0165
3.302
TABLE IV: Quantitative comparison under different mapping, rescaling, and refinement strategies. The AVGE values are multiplied by 1e2.
Fig. 7: Performance under different size of voxel grid down-sampling.
Fig. 8: Performance under different size of anchor voxel.
Size
Pts(k)
GS(k)
PSNR ↑
SSIM ↑
LPIPS ↓
T(min) ↓
-
574.4
525.5
28.43
0.954
0.052
43.9
0.2
463.4
421.1
28.22
0.952
0.052
38.5
0.3
282.8
263.2
28.04
0.950
0.056
28.7
0.5
130.3
139.0
27.72
0.943
0.064
19.9
1.0
42.2
92.7
27.36
0.935
0.074
17.8
∞
5.5
69.9
25.59
0.906
0.105
17.3
TABLE V: Performance comparison for different sizes of voxel grid downsampling.
Voxel Size
Pts(k)
GS(k)
PSNR ↑
SSIM ↑
LPIPS ↓
0.01
42.2
150.4
28.32
0.951
0.057
0.03
42.2
92.7
27.36
0.935
0.074
0.05
42.2
64.6
26.54
0.921
0.089
0.10
42.2
54.3
25.95
0.909
0.100
0.54
38.1
43.7
25.50
0.896
0.115
TABLE VI: Performance comparison for different sizes of anchor voxel.
Cf
Cr
PSNR ↑
SSIM ↑
LPIPS ↓
32
32
30.90
0.938
0.076
32
64
30.72
0.935
0.080
32
128
30.71
0.935
0.080
32
256
30.84
0.936
0.078
16
64
30.39
0.930
0.088
32
64
30.72
0.935
0.080
TABLE VII: Effect of different anchor feature and IBR feature channels.
Setting
PSNR ↑
SSIM ↑
LPIPS ↓
w/o MLP Emod
30.82
0.937
0.077
w/o ft. P
30.79
0.936
0.079
w/o ResNet Feature
30.85
0.937
0.078
w/o ViT Feature
30.81
0.936
0.078
Ours
30.90
0.938
0.076
TABLE VIII: Ablation study on different network settings.
Fig. 9: Illustration and analysis of the proposed pseudo view synthesis. (a) Pipeline of the pseudo view synthesis through progressive DIBR.(b) Example of the progressive synthesis process.
3D Gaussian Splatting (3DGS) has recently enabled real-time rendering of unbounded 3D scenes for novel view synthesis. However, this technique requires dense training views to accurately reconstruct 3D geometry. A limited number of input views will significantly degrade reconstruction quality, resulting in artifacts such as "floaters" and "background collapse" at unseen viewpoints. In this work, we introduce SparseGS, an efficient training pipeline designed to address the limitations of 3DGS in scenarios with sparse training views. SparseGS incorporates depth priors, novel depth rendering techniques, and a pruning heuristic to mitigate floater artifacts, alongside an Unseen Viewpoint Regularization module to alleviate background collapses. Our extensive evaluations on the Mip-NeRF360, LLFF, and DTU datasets demonstrate that SparseGS achieves high-quality reconstruction in both unbounded and forward-facing scenarios, with as few as 12 and 3 input images, respectively, while maintaining fast training and real-time rendering capabilities.
Generating high-quality novel views at real-time frame rates remains a central challenge in 3D vision, particularly in sparse-view scenarios. Neural radiance fields have demonstrated robust reconstruction from limited observations, but their reliance on volumetric rendering leads to high computational cost and slow inference. In contrast, Gaussian Splatting methods achieve real-time rendering through rasterization, but their optimization is highly sensitive to the quality of the initial geometry. This sensitivity becomes especially problematic in sparse-view settings, where limited observations often lead to incomplete or noisy point-cloud reconstructions. In this work, we present AugSplat, a simple framework for improving Gaussian Splatting in sparse-view regimes using radiance-field-based view augmentation. We first train a radiance field on the sparse input views and use it to synthesize additional images from nearby novel viewpoints, increasing the effective view-space coverage available for supervision. These synthetic views are then used as auxiliary supervision during Gaussian Splatting optimization. We study two variants: Staged AugSplat, which uses synthetic views for an initial optimization phase before switching to real images, and Dual AugSplat, which jointly trains on real and synthetic views with a decaying synthetic loss weight. Experiments on sparse-view mip-NeRF 360 scenes show that AugSplat improves reconstruction quality over standard Gaussian Splatting. Staged AugSplat achieves the strongest average performance, while Dual AugSplat provides a closely performing formulation that keeps real-image supervision active throughout training, and both variants preserve real-time rendering at inference.
Lorenzo Lazzaroni, Riccardo Bollati, Daniel Barath +2
3D Gaussian Splatting (3DGS) has achieved remarkable success in real-time novel view synthesis, yet it suffers from severe overfitting under sparse-view settings due to insufficient geometric constraints. While recent methods introduce monocular depth priors to mitigate this, they inherently struggle with scale ambiguity and cross-view inconsistency, leading to defective geometry. In this paper, we propose StereoGS, a novel sparse-view 3DGS framework that integrates stereo priors to establish reliable binocular consistency. Unlike scale-agnostic monocular constraints, StereoGS introduces a Stereo Depth Regularization by constructing virtual stereo pairs during optimization and leveraging a foundation stereo model to enforce absolute scale and binocular-consistent structures. To further suppress overfitting and eliminate redundant primitives, we design a Gradient-Aware Opacity Decay strategy that dynamically penalizes Gaussians based on their relative opacity gradient magnitudes. Combined with a Consistency-Aware Dense Initialization using zero-shot multi-view depth estimation, StereoGS effectively anchors primitives to accurate scene surfaces. Extensive experiments on LLFF, DTU, Mip-NeRF360, and Blender datasets demonstrate that StereoGS achieves state-of-the-art performance in sparse-view settings without incurring any additional inference overhead. Project Page: https://stringerywh00.github.io/StereoGS_project_page/