HierGF: Hierarchical Gaussian Fields via Geometry-perception Message Passing for Sparse-view 3D Reconstruction
Authors: Bi'an Du, Zhimin Zhang, Daizong Liu, Baoquan Chen, Wei Hu
Organizations: Wangxuan Institute of Computer Technology, Peking University, No. 128, Zhongguancun North Street, Beijing, China · Institute for Math & AI, Wuhan University, Wuhan, 430072, China · School of Intelligence Science and Technology, Peking University, Beijing 100871, China
Sparse view 3D reconstruction is an important and common scenario in multimedia applications, such as augmented reality/virtual reality (AR/VR) content creation, cultural heritage digitization, and certain robotic applications, where only a limited number of randomly captured views may be available. However, sparse views contain only limited 3D information, posing two major challenges:1) too few images are available for matching, making it difficult to build multi-view consistency; 2) insufficient view coverage leads to a lack of information in under-sampled regions, resulting in missing parts of object structure. Existing methods mostly still rely on limited reprojection errors and regularization terms, which are prone to overfitting to a single view and inconsistent appearances across views. In geometrically under-sampled regions, they often rely on heuristic density control, lacking reliable guidance and often resulting in blurring and structural holes.To address these issues, this paper proposes Hierarchical Gaussian Fields (HierGF), which revisits sparse-view reconstruction from a hierarchical geometry-perception perspective and converts limited observations into reliable self-generated supervision beyond fixed priors and heuristic density control. In particular, we transform coarse 3D geometric information and additional 2D generative priors into structured pseudo-supervision through a two-stage geometry-perception backbone network, thereby enhancing multi-view consistency with very few input views. In addition, we introduce a learnable confidence network to guide gradients toward cross-view consistent content, and a geometrically consistent densification module to improve the reconstruction of multi-view alignment and under-sampled regions.
Figures & tables
Fig. 1: Motivation of HierGF. (a) Sparse inputs in multimedia scenes provide weak multi-view constraints and leave large regions under-sampled. (b) Existing sparse-view methods rely on fixed regularizers or heuristic generative priors, causing holes and inconsistent textures. (c–d) Our hierarchical pipeline with 2D-assisted perceptual supervision turns few views into reliable pseudo-supervision and produces sharper, more consistent reconstructions than recent Gaussian-based baselines.
Method
Representation path
Diffusion use
Reliability / sparse-region handling
Main distinction from HierGF
3DGS-Enhancer [ 36 ]
Fine-tunes an initial 3DGS
Video diffusion restores rendered novel views
Confidence scores guide enhanced-view fine-tuning
HierGF uses a separate fine field, selective inheritance, and explicit sparse-region completion.
IPSM-Gaussian [ 37 ]
Optimizes 3DGS with diffusion guidance
Inline priors rectify sparse-view score matching
Uses depth/geometric regularization in the IPSM objective
HierGF does not rectify SDS; it refines a selectively initialized fine field with enhanced pseudo-supervision and GCD.
Confidence weights filter enhanced losses; GCD allocates Gaussians in sparse regions
Two-field refinement with selective inheritance and coupled perceptual/geometric refinement.
TABLE I: Direct comparison with diffusion-based sparse-view reconstruction methods.
Fig. 2: Overview of the proposed HierGF framework with explicit selective coarse-to-fine transfer and clarified diffusion adaptation. The optimized coarse Gaussian field transfers only spatial means and RGB initialization to the separately optimized fine field, whereas opacity, anisotropic scale, rotation, and higher-order appearance attributes are reinitialized. Only the LoRA adapters are updated during scene-specific diffusion adaptation; the pre-trained SD/ControlNet backbone remains frozen, and the adapted enhancer is frozen during reconstruction.
Component / stage
Data used
Updated parameters
Fixed parameters
Output
Coarse Gaussian optimization
Sparse reference views
Coarse Gaussians
Diffusion module not used
Coarse Gaussian field
Pair construction
Reference views; perturbed coarse renderings
None
Optimized coarse Gaussians
Degraded/reference image pairs
LoRA adaptation
Degraded renderings and reference images
LoRA only
Pre-trained ControlNet/SD v1.5 backbone
Scene-adapted enhancer
Pseudo-view rendering
Additional camera poses
None
Coarse Gaussian field
Coarse pseudo-views
Pseudo-view enhancement
Coarse pseudo-views
None
Diffusion backbone and adapted LoRA
Enhanced pseudo-supervision
Fine-scale optimization
Reference views; enhanced pseudo-supervision
Fine Gaussians; confidence module
Diffusion backbone and adapted LoRA
Final Gaussian representation
TABLE II: Parameter status and data usage in the diffusion-assisted reconstruction pipeline.
Method
4-view
6-view
9-view
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
DVGO [ 72 ]
24.43
14.39
79.12
26.67
14.30
76.76
25.66
14.74
78.42
3DGS [ 16 ]
10.80
20.31
89.91
8.38
22.12
91.34
6.42
24.29
93.31
DietNeRF [ 52 ]
11.17
18.90
89.71
6.96
22.03
92.86
5.85
23.55
94.24
RegNeRF [ 73 ]
20.44
13.59
84.76
20.72
13.41
84.18
19.70
13.68
85.17
FreeNeRF [ 74 ]
16.83
13.71
85.34
6.84
22.26
93.32
5.51
27.66
94.85
TABLE III: Comparisons on the MipNeRF360 dataset with varying input views. LPIPS ∗ = LPIPS x 102 and SSIM ∗ = SSIM x 102 throughout the paper.
Method
4-view
6-view
9-view
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS* ↓
PSNR ↑
SSIM ∗ ↑
DVGO [ 72 ]
14.48
17.14
89.52
12.89
18.32
91.42
11.49
19.26
93.02
3DGS [ 16 ]
8.60
17.29
92.99
7.74
18.29
93.78
6.50
20.26
94.83
DietNeRF [ 52 ]
11.64
18.56
92.05
10.39
19.07
92.67
10.32
19.26
92.58
RegNeRF [ 73 ]
16.75
15.20
90.91
14.38
15.80
92.07
10.17
17.93
94.20
FreeNeRF [ 74 ]
8.28
17.78
94.02
7.32
19.02
94.64
7.25
20.35
94.67
TABLE IV: Comparisons on the OmniObject3D datasets with varying input views.
Fig. 3: Qualitative examples on the MipNeRF360 and OmniObject3D datasets with 4 input views.
Fig. 4: Qualitative examples on the OpenIllumination dataset with 4 input views.
Item
Setting
Representation
Per-scene optimization with a coarse Gaussian field and a separately optimized fine Gaussian field
Camera parameters
Official benchmark intrinsics and extrinsics; fixed during optimization; no camera-pose refinement
SfM initialization input
Selected sparse input views only; no held-out evaluation views
SfM initialization procedure
Sparse feature extraction and exhaustive matching among input views, followed by triangulation under fixed benchmark cameras
Gaussian centers initialized from triangulated SfM points; RGB initialized from observed colors in input views
TABLE V: Implementation and evaluation settings of HierGF.
Method
4-view
6-view
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
DVGO [ 72 ]
11.84
21.15
89.73
8.83
23.79
92.09
3DGS [ 80 ]
30.08
11.50
84.54
29.65
11.98
82.77
DietNeRF† [ 52 ]
10.66
23.09
93.61
9.51
24.20
94.01
RegNeRF† [ 73 ]
47.31
11.61
69.40
30.28
14.08
85.86
FreeNeRF† [ 74 ]
35.81
12.21
79.69
35.15
11.47
81.28
TABLE VI: Quantitative comparisons on the OpenIllumination dataset. Methods with †mean the metrics are from the ZeroRF paper [ 4 ] . We follow ZeroRF [ 4 ] on dataset sampling and processing.
TABLE VII: Experimental assumptions of directly related sparse-view reconstruction methods.
Method
Result source
3-view
6-view
9-view
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
IPSM-Gaussian [ 37 ]
Reported
20.70
20.44
70.20
13.50
23.94
81.80
11.10
25.13
85.50
CoR-GS [ 82 ]
Reported
19.60
20.45
71.20
11.50
24.49
83.70
8.90
26.06
87.40
NexusGS [ 83 ]
Reported / Re-run
17.70
21.07
73.80
10.65
24.83
84.92
8.35
26.24
88.00
HierGF (ours)
Ours
17.05
21.28
74.45
10.10
25.03
85.42
7.97
26.39
88.37
TABLE VIII: Quantitative comparison on LLFF under 3-view, 6-view, and 9-view settings.
Method
Result source
3-view
6-view
9-view
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
IPSM-Gaussian [ 37 ]
Reported
12.10
19.99
85.60
–
–
–
–
–
–
CoR-GS [ 82 ]
Reported
11.90
19.21
85.30
6.80
24.51
91.70
4.50
27.18
94.70
NexusGS [ 83 ]
Reported / Re-run
10.20
20.21
86.90
6.15
24.91
92.25
4.15
27.38
95.05
HierGF (ours)
Ours
9.68
20.40
87.45
5.74
25.10
92.68
3.89
27.52
95.36
TABLE IX: Quantitative comparison on DTU under 3-view, 6-view, and 9-view settings.
Fig. 5: Qualitative comparison on LLFF under the 3-view setting. Representative held-out views from fern , orchids , room , and fortress are shown. Columns compare the ground truth, IPSM-Gaussian, CoR-GS, NexusGS, and HierGF at aligned viewpoints; enlarged crops emphasize thin structures, boundaries, texture details, and weakly observed regions.
Fig. 6: Qualitative comparison on DTU under the 3-view setting. Representative held-out views from scans 34, 40, 103, and 8 are shown. Columns compare the ground truth, IPSM-Gaussian, CoR-GS, NexusGS, and HierGF at aligned viewpoints; enlarged crops emphasize object contours, fine geometry, surface texture, and sparsely observed regions.
Fig. 7: Ablation study on the two-stage hierarchical backbone. From top to bottom, Coarse-only, Fine-only, Single-GF Joint, and Full HierGF are shown.
Fig. 8: Ablation study on geometry-consistent densification. Red boxes highlight that our full model better preserves thin structures and closes gaps in under-sampled regions.
Setting/Dataset
MipNeRF360
OmniObject
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
Coarse-only
9.76
19.63
88.90
7.92
17.88
94.35
Fine-only
8.12
21.84
91.03
6.84
20.59
95.01
Single-GF Joint
5.24
25.06
92.95
3.43
27.06
96.90
Full HierGF
3.85
28.48
95.87
1.76
34.08
98.36
TABLE X: Ablation study over HierGF’s hierarchical architecture on the MipNeRF360 and OmniObject datasets with 4 input views.
Components
Metrics
SPS
VCS
PCS
LPIPS ∗ ↓
PSNR ↑
SSIM ∗ ↑
9.40
20.33
89.78
✓
5.06
25.73
92.18
✓
✓
4.71
27.62
93.04
✓
✓
3.13
28.09
95.26
✓
✓
✓
3.85
28.48
95.87
TABLE XI: Ablation study over 2D-assisted Perceptual Enhancement on the MipNeRF360 and dataset with 4 input views. SPS , VCS , and PCS represent self-generated perceptual supervision, view-level confidence supervision and pixel-level confidence supervision, respectively.
Fine-field initialization strategy
Transferred parameters from coarse field
MipNeRF360
OmniObject3D
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
Reinitialize all attributes
None
4.46
27.21
94.67
2.21
31.65
97.53
Transfer all attributes
μ,q,s,σ,SH
4.18
27.82
95.14
1.98
32.93
97.95
Selective transfer, HierGF
μ and RGB only
3.85
28.48
95.87
1.76
34.08
98.36
TABLE XIII: Controlled ablation of the coarse-to-fine parameter-transfer rule under the 4-view setting.
Setting
MipNeRF360
OmniObject3D
LPIPS ∗↓
PSNR ↑
SSIM ∗↑
LPIPS ∗↓
PSNR ↑
SSIM ∗↑
Full HierGF w/o ZoeDepth
4.29
27.72
95.01
1.91
33.43
98.02
Full HierGF
3.85
28.48
95.87
1.76
34.08
98.36
TABLE XIV: Controlled ablation of ZoeDepth supervision under the 4-view setting.
Setting
LPIPS ∗↓
PSNR ↑
SSIM ∗↑
Reference views only; no pseudo-view supervision
10.07
19.86
89.05
Raw pseudo-views; no diffusion enhancement or confidence weighting
9.40
20.33
89.78
Frozen pre-trained 2D prior; without LoRA adaptation or confidence weighting
5.73
24.86
91.64
LoRA-adapted enhanced pseudo-views (SPS); without confidence weighting
5.06
25.73
92.18
SPS + view-level confidence supervision (VCS)
4.71
27.62
93.04
SPS + pixel-level confidence supervision (PCS)
3.13
28.09
95.26
TABLE XV: Controlled ablation of pseudo-view supervision, the diffusion prior, LoRA adaptation, and confidence weighting on MipNeRF360 under the 4-view setting.
Radius setting
Radius value
MipNeRF360
OmniObject3D
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
Small radius
1.5sv=0.03Drob
4.16
28.17
94.82
1.86
33.72
98.04
Default radius
2.0sv=0.04Drob
3.85
28.48
95.87
1.76
34.08
98.36
Large radius
3.0sv=0.06Drob
4.04
28.28
95.18
1.83
33.89
98.18
TABLE XVI: Ablation study on the neighborhood radius r under the 4-view setting.
Setting
MipNeRF360
OmniObject3D
LPIPS* ↓
PSNR ↑
SSIM* ↑
LPIPS* ↓
PSNR ↑
SSIM* ↑
w/o GCD, standard ADC
4.58
27.96
93.37
1.93
31.07
97.56
Distance-only interpolation
4.02
28.25
94.92
1.79
33.74
98.05
Additive: wnormal+wdistance
3.95
28.34
95.31
1.78
33.91
98.18
Multiplicative: wnormal⋅wdistance
3.85
28.48
95.87
1.76
34.08
98.36
TABLE XVII: Comparison of weighting strategies in geometry-consistent densification under the 4-view setting.
Method
Recon. time (s) ↓
Recon. speedup
Rendering time (s) ↓
Rendering speedup
ZeroRF [ 4 ]
1309
2.07×
78
1.11×
GaussianObject [ 9 ]
7028
11.12×
89
1.27×
WaveletGS [ 76 ]
1800
2.85×
95
1.36×
HierGF (ours)
632
–
70
–
TABLE XVIII: Efficiency comparison on the MipNeRF360 kitchen scene under the 4-view setting. All methods are evaluated at 779×520 on a single NVIDIA A100 GPU. The speedup values are calculated relative to HierGF.
Fig. 9: Stage-wise runtime breakdown of HierGF on the MipNeRF360 kitchen scene under the 4-view setting. Gaussian optimization remains the dominant cost, while diffusion-related processing is used only during reconstruction.
State Key Lab of CAD&CG, Zhejiang University, China · University of Utah, USA · State Key Lab of CAD&CG, Zhejiang University, China and Hangzhou Research Institute of Holographic and AI Technology, China +1