HierGF: Hierarchical Gaussian Fields via Geometry-perception Message Passing for Sparse-view 3D Reconstruction
Authors: Bi'an Du, Zhimin Zhang, Daizong Liu, Baoquan Chen, Wei Hu
Organizations: Wangxuan Institute of Computer Technology, Peking University, No. 128, Zhongguancun North Street, Beijing, China · Institute for Math & AI, Wuhan University, Wuhan, 430072, China · School of Intelligence Science and Technology, Peking University, Beijing 100871, China
Sparse view 3D reconstruction is an important and common scenario in multimedia applications, such as augmented reality/virtual reality (AR/VR) content creation, cultural heritage digitization, and certain robotic applications, where only a limited number of randomly captured views may be available. However, sparse views contain only limited 3D information, posing two major challenges:1) too few images are available for matching, making it difficult to build multi-view consistency; 2) insufficient view coverage leads to a lack of information in under-sampled regions, resulting in missing parts of object structure. Existing methods mostly still rely on limited reprojection errors and regularization terms, which are prone to overfitting to a single view and inconsistent appearances across views. In geometrically under-sampled regions, they often rely on heuristic density control, lacking reliable guidance and often resulting in blurring and structural holes.To address these issues, this paper proposes Hierarchical Gaussian Fields (HierGF), which revisits sparse-view reconstruction from a hierarchical geometry-perception perspective and converts limited observations into reliable self-generated supervision beyond fixed priors and heuristic density control. In particular, we transform coarse 3D geometric information and additional 2D generative priors into structured pseudo-supervision through a two-stage geometry-perception backbone network, thereby enhancing multi-view consistency with very few input views. In addition, we introduce a learnable confidence network to guide gradients toward cross-view consistent content, and a geometrically consistent densification module to improve the reconstruction of multi-view alignment and under-sampled regions.
Figures & tables
Fig. 1: Motivation of HierGF. (a) Sparse inputs in multimedia scenes provide weak multi-view constraints and leave large regions under-sampled. (b) Existing sparse-view methods rely on fixed regularizers or heuristic generative priors, causing holes and inconsistent textures. (c–d) Our hierarchical pipeline with 2D-assisted perceptual supervision turns few views into reliable pseudo-supervision and produces sharper, more consistent reconstructions than recent Gaussian-based baselines.
Method
Representation path
Diffusion use
Reliability / sparse-region handling
Main distinction from HierGF
3DGS-Enhancer [ 36 ]
Fine-tunes an initial 3DGS
Video diffusion restores rendered novel views
Confidence scores guide enhanced-view fine-tuning
HierGF uses a separate fine field, selective inheritance, and explicit sparse-region completion.
IPSM-Gaussian [ 37 ]
Optimizes 3DGS with diffusion guidance
Inline priors rectify sparse-view score matching
Uses depth/geometric regularization in the IPSM objective
HierGF does not rectify SDS; it refines a selectively initialized fine field with enhanced pseudo-supervision and GCD.
Confidence weights filter enhanced losses; GCD allocates Gaussians in sparse regions
Two-field refinement with selective inheritance and coupled perceptual/geometric refinement.
TABLE I: Direct comparison with diffusion-based sparse-view reconstruction methods.
Fig. 2: Overview of the proposed HierGF framework with explicit selective coarse-to-fine transfer and clarified diffusion adaptation. The optimized coarse Gaussian field transfers only spatial means and RGB initialization to the separately optimized fine field, whereas opacity, anisotropic scale, rotation, and higher-order appearance attributes are reinitialized. Only the LoRA adapters are updated during scene-specific diffusion adaptation; the pre-trained SD/ControlNet backbone remains frozen, and the adapted enhancer is frozen during reconstruction.
Component / stage
Data used
Updated parameters
Fixed parameters
Output
Coarse Gaussian optimization
Sparse reference views
Coarse Gaussians
Diffusion module not used
Coarse Gaussian field
Pair construction
Reference views; perturbed coarse renderings
None
Optimized coarse Gaussians
Degraded/reference image pairs
LoRA adaptation
Degraded renderings and reference images
LoRA only
Pre-trained ControlNet/SD v1.5 backbone
Scene-adapted enhancer
Pseudo-view rendering
Additional camera poses
None
Coarse Gaussian field
Coarse pseudo-views
Pseudo-view enhancement
Coarse pseudo-views
None
Diffusion backbone and adapted LoRA
Enhanced pseudo-supervision
Fine-scale optimization
Reference views; enhanced pseudo-supervision
Fine Gaussians; confidence module
Diffusion backbone and adapted LoRA
Final Gaussian representation
TABLE II: Parameter status and data usage in the diffusion-assisted reconstruction pipeline.
Method
4-view
6-view
9-view
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
DVGO [ 72 ]
24.43
14.39
79.12
26.67
14.30
76.76
25.66
14.74
78.42
3DGS [ 16 ]
10.80
20.31
89.91
8.38
22.12
91.34
6.42
24.29
93.31
DietNeRF [ 52 ]
11.17
18.90
89.71
6.96
22.03
92.86
5.85
23.55
94.24
RegNeRF [ 73 ]
20.44
13.59
84.76
20.72
13.41
84.18
19.70
13.68
85.17
FreeNeRF [ 74 ]
16.83
13.71
85.34
6.84
22.26
93.32
5.51
27.66
94.85
TABLE III: Comparisons on the MipNeRF360 dataset with varying input views. LPIPS ∗ = LPIPS x 102 and SSIM ∗ = SSIM x 102 throughout the paper.
Method
4-view
6-view
9-view
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS* ↓
PSNR ↑
SSIM ∗ ↑
DVGO [ 72 ]
14.48
17.14
89.52
12.89
18.32
91.42
11.49
19.26
93.02
3DGS [ 16 ]
8.60
17.29
92.99
7.74
18.29
93.78
6.50
20.26
94.83
DietNeRF [ 52 ]
11.64
18.56
92.05
10.39
19.07
92.67
10.32
19.26
92.58
RegNeRF [ 73 ]
16.75
15.20
90.91
14.38
15.80
92.07
10.17
17.93
94.20
FreeNeRF [ 74 ]
8.28
17.78
94.02
7.32
19.02
94.64
7.25
20.35
94.67
TABLE IV: Comparisons on the OmniObject3D datasets with varying input views.
Fig. 3: Qualitative examples on the MipNeRF360 and OmniObject3D datasets with 4 input views.
Fig. 4: Qualitative examples on the OpenIllumination dataset with 4 input views.
Item
Setting
Representation
Per-scene optimization with a coarse Gaussian field and a separately optimized fine Gaussian field
Camera parameters
Official benchmark intrinsics and extrinsics; fixed during optimization; no camera-pose refinement
SfM initialization input
Selected sparse input views only; no held-out evaluation views
SfM initialization procedure
Sparse feature extraction and exhaustive matching among input views, followed by triangulation under fixed benchmark cameras
Gaussian centers initialized from triangulated SfM points; RGB initialized from observed colors in input views
TABLE V: Implementation and evaluation settings of HierGF.
Method
4-view
6-view
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
DVGO [ 72 ]
11.84
21.15
89.73
8.83
23.79
92.09
3DGS [ 80 ]
30.08
11.50
84.54
29.65
11.98
82.77
DietNeRF† [ 52 ]
10.66
23.09
93.61
9.51
24.20
94.01
RegNeRF† [ 73 ]
47.31
11.61
69.40
30.28
14.08
85.86
FreeNeRF† [ 74 ]
35.81
12.21
79.69
35.15
11.47
81.28
TABLE VI: Quantitative comparisons on the OpenIllumination dataset. Methods with †mean the metrics are from the ZeroRF paper [ 4 ] . We follow ZeroRF [ 4 ] on dataset sampling and processing.
TABLE VII: Experimental assumptions of directly related sparse-view reconstruction methods.
Method
Result source
3-view
6-view
9-view
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
IPSM-Gaussian [ 37 ]
Reported
20.70
20.44
70.20
13.50
23.94
81.80
11.10
25.13
85.50
CoR-GS [ 82 ]
Reported
19.60
20.45
71.20
11.50
24.49
83.70
8.90
26.06
87.40
NexusGS [ 83 ]
Reported / Re-run
17.70
21.07
73.80
10.65
24.83
84.92
8.35
26.24
88.00
HierGF (ours)
Ours
17.05
21.28
74.45
10.10
25.03
85.42
7.97
26.39
88.37
TABLE VIII: Quantitative comparison on LLFF under 3-view, 6-view, and 9-view settings.
Method
Result source
3-view
6-view
9-view
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
IPSM-Gaussian [ 37 ]
Reported
12.10
19.99
85.60
–
–
–
–
–
–
CoR-GS [ 82 ]
Reported
11.90
19.21
85.30
6.80
24.51
91.70
4.50
27.18
94.70
NexusGS [ 83 ]
Reported / Re-run
10.20
20.21
86.90
6.15
24.91
92.25
4.15
27.38
95.05
HierGF (ours)
Ours
9.68
20.40
87.45
5.74
25.10
92.68
3.89
27.52
95.36
TABLE IX: Quantitative comparison on DTU under 3-view, 6-view, and 9-view settings.
Fig. 5: Qualitative comparison on LLFF under the 3-view setting. Representative held-out views from fern , orchids , room , and fortress are shown. Columns compare the ground truth, IPSM-Gaussian, CoR-GS, NexusGS, and HierGF at aligned viewpoints; enlarged crops emphasize thin structures, boundaries, texture details, and weakly observed regions.
Fig. 6: Qualitative comparison on DTU under the 3-view setting. Representative held-out views from scans 34, 40, 103, and 8 are shown. Columns compare the ground truth, IPSM-Gaussian, CoR-GS, NexusGS, and HierGF at aligned viewpoints; enlarged crops emphasize object contours, fine geometry, surface texture, and sparsely observed regions.
Fig. 7: Ablation study on the two-stage hierarchical backbone. From top to bottom, Coarse-only, Fine-only, Single-GF Joint, and Full HierGF are shown.
Fig. 8: Ablation study on geometry-consistent densification. Red boxes highlight that our full model better preserves thin structures and closes gaps in under-sampled regions.
Setting/Dataset
MipNeRF360
OmniObject
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
Coarse-only
9.76
19.63
88.90
7.92
17.88
94.35
Fine-only
8.12
21.84
91.03
6.84
20.59
95.01
Single-GF Joint
5.24
25.06
92.95
3.43
27.06
96.90
Full HierGF
3.85
28.48
95.87
1.76
34.08
98.36
TABLE X: Ablation study over HierGF’s hierarchical architecture on the MipNeRF360 and OmniObject datasets with 4 input views.
Components
Metrics
SPS
VCS
PCS
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
9.40
20.33
89.78
✓
5.06
25.73
92.18
✓
✓
4.71
27.62
93.04
✓
✓
3.13
28.09
95.26
✓
✓
✓
3.85
28.48
95.87
TABLE XI: Ablation study over 2D-assisted Perceptual Enhancement on the MipNeRF360 and dataset with 4 input views. SPS , VCS , and PCS represent self-generated perceptual supervision, view-level confidence supervision and pixel-level confidence supervision, respectively.
Fine-field initialization strategy
Transferred parameters from coarse field
MipNeRF360
OmniObject3D
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
Reinitialize all attributes
None
4.46
27.21
94.67
2.21
31.65
97.53
Transfer all attributes
μ,q,s,σ,SH
4.18
27.82
95.14
1.98
32.93
97.95
Selective transfer, HierGF
μ and RGB only
3.85
28.48
95.87
1.76
34.08
98.36
TABLE XIII: Controlled ablation of the coarse-to-fine parameter-transfer rule under the 4-view setting.
Setting
MipNeRF360
OmniObject3D
LPIPS ∗↓
PSNR ↑
SSIM ∗↑
LPIPS ∗↓
PSNR ↑
SSIM ∗↑
Full HierGF w/o ZoeDepth
4.29
27.72
95.01
1.91
33.43
98.02
Full HierGF
3.85
28.48
95.87
1.76
34.08
98.36
TABLE XIV: Controlled ablation of ZoeDepth supervision under the 4-view setting.
Setting
LPIPS ∗↓
PSNR ↑
SSIM ∗↑
Reference views only; no pseudo-view supervision
10.07
19.86
89.05
Raw pseudo-views; no diffusion enhancement or confidence weighting
9.40
20.33
89.78
Frozen pre-trained 2D prior; without LoRA adaptation or confidence weighting
5.73
24.86
91.64
LoRA-adapted enhanced pseudo-views (SPS); without confidence weighting
5.06
25.73
92.18
SPS + view-level confidence supervision (VCS)
4.71
27.62
93.04
SPS + pixel-level confidence supervision (PCS)
3.13
28.09
95.26
TABLE XV: Controlled ablation of pseudo-view supervision, the diffusion prior, LoRA adaptation, and confidence weighting on MipNeRF360 under the 4-view setting.
Radius setting
Radius value
MipNeRF360
OmniObject3D
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
Small radius
1.5sv=0.03Drob
4.16
28.17
94.82
1.86
33.72
98.04
Default radius
2.0sv=0.04Drob
3.85
28.48
95.87
1.76
34.08
98.36
Large radius
3.0sv=0.06Drob
4.04
28.28
95.18
1.83
33.89
98.18
TABLE XVI: Ablation study on the neighborhood radius r under the 4-view setting.
Setting
MipNeRF360
OmniObject3D
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
w/o GCD, standard ADC
4.58
27.96
93.37
1.93
31.07
97.56
Distance-only interpolation
4.02
28.25
94.92
1.79
33.74
98.05
Additive: wnormal+wdistance
3.95
28.34
95.31
1.78
33.91
98.18
Multiplicative: wnormal⋅wdistance
3.85
28.48
95.87
1.76
34.08
98.36
TABLE XVII: Comparison of weighting strategies in geometry-consistent densification under the 4-view setting.
Method
Recon. time (s) ↓
Recon. speedup
Rendering time (s) ↓
Rendering speedup
ZeroRF [ 4 ]
1309
2.07×
78
1.11×
GaussianObject [ 9 ]
7028
11.12×
89
1.27×
WaveletGS [ 76 ]
1800
2.85×
95
1.36×
HierGF (ours)
632
–
70
–
TABLE XVIII: Efficiency comparison on the MipNeRF360 kitchen scene under the 4-view setting. All methods are evaluated at 779×520 on a single NVIDIA A100 GPU. The speedup values are calculated relative to HierGF.
Fig. 9: Stage-wise runtime breakdown of HierGF on the MipNeRF360 kitchen scene under the 4-view setting. Gaussian optimization remains the dominant cost, while diffusion-related processing is used only during reconstruction.
We introduce S2C-3D, a novel sparse-view 3D reconstruction framework for high-fidelity and complete scene reconstruction from as few as six to eight images. Our framework features three components: a specialized diffusion model for scene-specific image restoration, a training-free view-consistency conditioned sampling process in the diffusion model for refined Gaussian optimization, and a camera trajectory planning scheme to ensure comprehensive scene coverage. The specialized diffusion model is developed by finetuning a pretrained architecture on the input views and their corresponding degraded counterparts. The adaptation to the scene distribution allows the model to repair Gaussian renderings while effectively eliminating domain gaps. Meanwhile, the trajectory planning scheme optimizes scene coverage by connecting each newly sampled camera to its two nearest neighbors. By iteratively constructing paths and retaining only those that significantly enhance visibility, the scheme establishes a trajectory that covers the entire scene. To address multi-view conflicts, the view-consistency conditioned sampling process quantifies the consistency between neighboring repaired images. This information is injected as a condition into the sampling process of the frozen diffusion model, facilitating the generation of view-consistent images without additional training. Consequently, our approach produces high-fidelity 3D Gaussians that are robust to artifacts. Experimental results demonstrate that S2C-3D outperforms state-of-the-art methods, constructing high-quality scenes that are free from missing regions, blurring, or other artifacts with very sparse inputs. The source code and data are available at https://gapszju.github.io/S2C-3D.
Yiyang Shen, Yin Yang, Kun Zhou +1
State Key Lab of CAD&CG, Zhejiang University, China · University of Utah, USA · State Key Lab of CAD&CG, Zhejiang University, China and Hangzhou Research Institute of Holographic and AI Technology, China +1
3D Gaussian Splatting (3DGS) has emerged as a prominent paradigm for 3D reconstruction and novel view synthesis. However, it remains vulnerable to severe artifacts when trained under sparse-view constraints. While recent methods attempt to rectify artifacts in rendered views using image diffusion models, they typically rely on multi-view self-attention to retrieve information from reference images. We observe that this mechanism often fails when the rendered novel views output by 3DGS are heavily corrupted: damaged query features lead to erroneous cross-view retrieval, resulting in inconsistent rendering refinement. To address this, we propose GeoQuery, a geometry-guided diffusion framework that integrates generative priors with explicit geometric cues via a novel Geometry-guided Cross-view Attention (GCA) mechanism. First, by leveraging predicted depth maps and camera poses, we construct a geometry-induced correspondence field to sample reference features, forming a geometry-aligned proxy query that replaces the corrupted rendering features. Furthermore, we design a new cross-view feature aggregation pipeline, in which we restrict the cross-view attention to a local window around each proxy query to effectively retrieve useful features while suppressing spurious matches. GeoQuery can be seamlessly integrated into existing diffusion-based pipelines, enabling robust reconstruction even under extreme view sparsity. Extensive experiments on sparse-view novel view synthesis and rendering artifact removal demonstrate the effectiveness of our approach.
Xiao Cao, Yuze Li, Youmin Zhang +4
University of Electronic Science and Technology of China & Rawmantic AI, China · Rawmantic AI, China · Tianjin University, China +1
3D reconstruction from sparse views is a challenging task in 3D computer vision. Recent studies on 3D Gaussian Splatting (3DGS) have achieved remarkable results with sparse views in novel view synthesis, yet reconstructing high-quality geometric surfaces from sparse views remains a challenge, due to the limited geometry clues and the discreteness of Gaussians. In this paper, we propose a novel 3DGS-based method for high-fidelity surface reconstruction from sparse views. Our key insight is to introduce a normal-guided depth propagation approach, which can extend depth information from high-confidence regions to constrain the depth in low-confidence areas. Additionally, we propose an abnormal depth edge-aware regularization to address depth discontinuities caused by the discreteness of Gaussians. Extensive experiments on DTU and Tanks-and-Temples datasets demonstrate that our method outperforms the state-of-the-art methods in sparse view surface reconstruction. Project page: https://hanl2010.github.io/DP-GS.
Liang Han, Bangcai Wei, Junsheng Zhou +2
School of Software, Tsinghua University, Beijing, China · China Telecom · Department of Computer Science, Wayne State University, Detroit, USA