3D Gaussian Splatting (3DGS) is a state-of-the-art technique for 3D scene rendering, offering high efficiency and excellent visual quality. However, because 3DGS relies on an initial sparse point set from Structure-from-Motion (SfM) and view-dependent properties, it can suffer from geometric inaccuracies and visual artifacts, particularly in complex scenes. To address these challenges, we propose an improved 3DGS approach that regularizes the optimization process by integrating geometric priors, including surface normals and dense depth information. Surface normal regularization improves geometric consistency by aligning Gaussian covariance with local surface structures, while dense depth priors combined with an initial points from SfM enhance per-pixel depth estimation, increasing accuracy and reducing ambiguities. These enhancements enable robust handling of diverse and complex real-world scenarios, minimizing visual distortions and improving reconstruction quality across various environments. To validate our method, we evaluate it on challenging datasets, including street-view scenes and highly reflective environments, while testing it across multiple SfM pipelines. Our results demonstrate compatibility across diverse environments and highlight the robustness of our approach. Experimental findings further show that our method enhances geometric accuracy and visual quality, establishing a reliable solution for real-time 3D scene rendering in complex environments.
Figures & tables
Figure 1: Hardware for Data Collection
COLMAP
SP-SG
LoFTR
Datasets
Method
MRE ↓
PSNR ↑
SSIM ↑
LPIPS ↓
MRE ↓
PSNR ↑
SSIM ↑
LPIPS ↓
MRE ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Train
3DGS
0.75
21.10
0.802
0.218
1.40
21.10
0.750
0.282
0.79
20.97
0.767
0.274
Ours
21.97
0.799
0.252
21.32
0.749
0.287
21.11
0.759
0.291
Horse
3DGS
0.71
24.18
0.889
0.239
1.31
21.01
0.802
0.239
0.80
23.39
0.870
0.174
Ours
25.50
0.903
0.153
21.12
0.801
0.246
24.80
0.881
0.165
Parking lots
3DGS
Invalid
-
-
-
1.37
28.08
0.828
0.424
0.64
29.33
0.842
0.410
Table 1: Quantitative comparison of 3DGS, Our Method, and SfM-Rendering Correlation - A lower Mean Reprojection Error (MRE) generally corresponds to a higher PSNR and SSIM, while its correlation with LPIPS is relatively weak. Among the SfM pipelines, COLMAP and LoFTR achieve relatively low MRE, whereas SP-SG exhibits a higher MRE. For the parking lot and street-view datasets, SfM with COLMAP failed to reconstruct the scenes, resulting in invalid outputs. The proposed method demonstrates improved PSNR and SSIM performance across the COLMAP, SP-SG, and LoFTR approaches.
Figure 2: Qualitative comparison of 3DGS and our method - Our method demonstrates improved rendering quality over 3DGS by leveraging surface normals and dense depth as prior information. This prior information enhance Gaussian alignment and depth representation. Normal & Depth information for parking lots and street-view, flat surfaces, is represented appropriately.
3D Gaussian Splatting (3DGS) has emerged as a powerful technique for generating photorealistic renderings of a scene in real-time. However, the volumetric nature of 3DGS limits its ability to accurately capture surface geometry. To address this, 2D Gaussian Splatting (2DGS) was proposed to enable view-consistent and geometrically accurate surface reconstruction from multi-view images. However, 2DGS can be sensitive to the initialization of the Gaussian primitives. Reliance on Structure-from-Motion (SfM) initializations, which can produce poor estimates on challenging image sets, may lead to subpar results. In this work, we enhance 2DGS by incorporating monocular depth and normal priors to improve both geometric accuracy and robustness. We propose a depth-guided initialization strategy for Gaussians and introduce a clustering-based technique for pruning degenerate Gaussians. We evaluate our method on the DTU dataset, where it achieves state-of-the-art results in mesh reconstruction while preserving high-quality novel view synthesis.
Prajwal Gupta C. R., Divyam Sheth, Jinjoo Ha +2
TU Darmstadt · ELIZA · Max Planck Institute for Intelligent Systems
High-fidelity surface reconstruction from multi-view images is a core problem in 3D computer vision. While neural implicit surfaces like SDFs offer smooth geometry, they are often bottlenecked by the computational intensity of volume rendering. Conversely, 3D Gaussian Splatting (3DGS) provides rapid training but lacks geometry continuity, often leading to fragmented surfaces. This paper presents a novel framework that integrates Signed Distance Fields directly into the splatting pipeline. By leveraging the continuous nature of SDFs to regularize Gaussian primitives, our method effectively fills geometric holes and suppresses noise inherent in sparse point clouds. Unlike hybrid approaches that rely on heavy volumetric sampling, our approach utilizes the efficiency of splatting to achieve faster convergence. Extensive evaluations demonstrate that our method produces high-quality surfaces with significantly fewer primitives, offering a more compact and efficient representation for both indoor and outdoor environments.
Baixin Xu, Jiangbei Hu, Jiaze Li +1
College of Computing and Data Science, Nanyang Technological University · School of Software, Dalian University of Technology
3D Gaussian splatting (3DGS) has emerged as a widely-used tool for novel view synthesis, offering real-time rendering in a sparse representation. However, the method's reliance on structure-from-motion initialization and photometric optimization can lead to suboptimal geometric reconstruction, particularly for objects with high specularity. In this work, we investigate the integration of geometric priors, in the form of predicted normal and depth maps, into the 3DGS framework to improve the reconstruction quality. We analyze the effect of incorporating these priors into GS-based methods and our evaluation reveals that multi-view predictions, as they are done by the recent visual geometry grounded transformer (VGGT), outperform single-view alternatives. A major factor is the existence of a confidence map for the estimations, which comes as a by-product of multi-view models and which can significantly improve the effectiveness of priors by weighting each prediction appropriately. Extensive experiments on standard benchmarks show consistent improvement in reconstruction quality and significant gains in complex scenes including specular objects.
Hongyu Zhou, Zorah Lähner
University of Bonn · Lamarr Institute for Machine Learning and Artificial Intelligence