Surface reconstruction under sparse-view settings remains challenging due to limited geometric cues. Volume rendering methods based on signed distance functions often produce over-smoothed surfaces, while 3D Gaussian Splatting (3DGS), though time-efficient, suffers from incomplete geometry due to the lack of reliable depth supervision and the limitation of being optimized only from given input views. In this paper, we present Sparse-GS2Mesh, a stereo-aware framework for surface reconstruction from sparse views. While 3DGS and stereo matching have been leveraged for surface reconstruction under dense view settings, we extend them to operate effectively under sparse view conditions by first initializing 3DGS using epipolar depth priors to mitigate the 3DGS overfitting problem, followed by our three key components: (I) adaptive baseline selection, (II) fine-tuning with a stereo matching network, and (III) 2D/3D co-regularized fine-tuning. Given a warmed-up 3DGS initialized with epipolar depth, the adaptive baseline selection automatically determines a baseline to synthesize for each sparse view. We then fine-tune 3DGS by backpropagating depth-refining gradients from the stereo matching network, effectively specializing the 3DGS for stereo matching. The 2D/3D co-regularization further helps obtain stable reconstruction, addressing weak geometric cues in close stereo views. Sparse-GS2Mesh achieves a 15% improvement over state-of-the-art methods in little-overlap settings and comparable results in large-overlap settings. Codes will be publicly available.
Figures & tables
Figure 2: Overview of Sparse-GS2Mesh. Stereo renderings from warmed-up 3DGS are refined via backpropagated gradients from the stereo matching network (Sec. 3.3 ). Subsequently, 2D/3D co-regularization with stereo image and depth supervision stabilizes 3DGS geometry (Sec. 3.4 ), specializing it for stereo matching.
Figure 3: DTU surface reconstruction results under different sparse view settings. Our method achieves more complete reconstruction with finer details.
Figure 4: BlendedMVS reconstruction results. Our method demonstrates a more complete and detailed reconstruction.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Downsample
GPU Memory
Training Time
CD ↓
50%
21GB
31 mins
1.140
66%
34GB
33 mins
1.117
100%
73GB
37 mins
1.069
Appendix
Table 5: Computational report across downsampling ratios.
Figure 5: Generated Initial Pointcloud. The point cloud generated from epipolar depth is the most dense and preserves the finest geometric details
Init. Strategy
Dense SFM [ 23 ]
MVSNet [ 1 ]
Epipolar Depth [ 45 ]
CD ↓
1.341
1.123
1.069
Appendix
Table 6: Reconstruction quality of our method under different initialization strategies.
Figure 6: Visual comparison of the final reconstructed meshes under different initialization strategies.
Figure 7: Effect of fine-tuning on rendering quality. After the fine-tuning, artifacts are removed, resulting in more accurate stereo renderings.
Fine-Tuning
PSNR ↑
LPIPS ↓
Before
26.27
0.058
After
26.57
0.054
Appendix
Table 7: Rendering quality comparison before and after proposed fine-tuning on stereo view. The image quality improves after fine-tuning, showing more accurate stereo renderings.