UGOD: Uncertainty-Guided Opacity and Dropout for Sparse-View 3D Gaussian Splatting
Authors: Zhihao Guo, Peng Wang, Zidong Chen, Xiangyu Kong, Yan Lyu, Guanyu Gao, Chenghao Qian, Ziyang Wang, +2 more
Organizations: Manchester Metropolitan University, Manchester, United Kingdom. · University of Surrey, Guildford, United Kingdom. · Imperial College London, London, United Kingdom. · University of Exeter, Exeter, United Kingdom. · Southeast University, Nanjing, China. · Nanjing University of Science and Technology, Nanjing, China. · University of Leeds, Leeds, United Kingdom. · Aston University, Birmingham, United Kingdom.
Sparse-view 3D Gaussian Splatting is prone to overfitting because limited observations leave many Gaussian primitives weakly constrained, yet their contributions are still accumulated through alpha blending. Without uncertainty estimation, the renderer cannot distinguish unreliable primitives from well-constrained ones, allowing their erroneous contributions to corrupt novel-view synthesis. We introduce UGOD, an uncertainty-guided framework that estimates a view-dependent uncertainty score for each Gaussian and uses it to regulate its rendering contribution. A lightweight uncertainty head conditioned on Gaussian attributes and viewing direction predicts this score, which then drives a differentiable opacity-modulation mechanism that attenuates high-uncertainty primitives before compositing. During training, a detached soft-dropout branch applies an uncertainty-controlled continuous keep mask to discourage the model from relying on poorly constrained Gaussians and thereby reduce overfitting. Crucially, detaching the uncertainty score prevents gradients from this stochastic regulariser from biasing or collapsing the uncertainty prediction. Experiments on Mip-NeRF~360 and LLFF show that UGOD improves sparse-view novel-view synthesis while producing more compact Gaussian representations than the compared methods. These results demonstrate that Gaussian uncertainty provides an effective rendering-time control for sparse-view reconstruction.
Figures & tables
Figure 1: Average PSNR and final-Gaussian reduction versus 3DGS* across MiP-NeRF 360 (24/8 views) and LLFF (3 views). Higher is better on every axis; UGOD leads all six.
Figure 2: Overview of UGOD. Sparse input views are reconstructed by SfM and initialise 3D Gaussian primitives (left). For each visible Gaussian, the uncertainty head combines the viewing direction, HashGrid position encoding, geometric attributes, and compressed SH appearance to predict a view-dependent uncertainty ui (lower left). This score drives two gradient-decoupled mechanisms (centre): normalised uncertainty differentiably gates opacity before compositing, whereas its detached copy determines the training-only Concrete soft-dropout probability. The resulting Gaussians are jointly optimised using supervised depth, pseudo-view depth, and photometric losses (right). At inference, soft dropout and relative-uncertainty gating are disabled, and deterministic raw-uncertainty opacity modulation is used.
Setting
Method
PSNR ↑
SSIM ↑
LPIPS ↓
#Gaussians
vs 3DGS*
LLFF (3-view)
3DGS*
14.6543
0.4379
0.3990
254,350
–
DropGS
14.6443
0.4557
0.3894
210,414
1.2 ×
FSGS (w/ depth)
19.0457
0.6230
0.2514
151,173
1.7 ×
Ours (full)
19.3300
0.6370
0.2449
102,280
2.5 ×
Mip360 (8-view)
3DGS*
11.9591
0.2747
0.6095
864,393
–
DropGS
12.7267
0.3193
0.6026
716,133
1.2 ×
Table 1: Scene-average results on LLFF (3 views) and MiP-NeRF 360 (8/24 views). Reduction is relative to 3DGS*; best and second-best values are bold and underlined.
Setting
Variant
N
PSNR ↑
SSIM ↑
LPIPS ↓
#Gaussians ↓
LLFF (3-view)
Ours
7
18.2386
0.5989
0.2701
180,905
w/o soft dropout
7
17.9529
0.5700
0.2877
183,460
Ours (full)
7
19.3300
0.6370
0.2449
102,280
Mip360 (8-view)
w/o soft dropout
5
13.9438
0.3672
0.6074
108,008
w/ soft dropout
5
12.9741
0.3475
0.6267
102,644
Ours (full)
6
14.1435
0.3682
0.6179
95,662
Table 2: Scene-average component ablation. N is the number of scenes averaged; the 8-view soft-dropout variants use five scenes, while the full model uses all six. Best and second-best values are bold and underlined.
Figure 3: Qualitative comparison on MiP-NeRF 360 (24/8 views) and LLFF (3 views). Each non-GT panel reports the displayed-view PSNR and final Gaussian count.
P
S
R
V
PSNR ↑
SSIM ↑
LPIPS ↓
5
0
0
0
12.8370
0.4092
0.6619
5
0
0
1
14.7173
0.4575
0.6400
6
0
0
0
15.0757
0.4628
0.6396
6
1
1
0
14.6207
0.4508
0.6461
7
0
0
0
10.9079
0.4070
0.6758
Table 3: Controlled HashGrid input-encoding ablation on MiP-NeRF 360 kitchen with 24 training views. P , S , R , and V are the numbers of encoding levels assigned to position, scale, rotation, and view direction. Bold indicates the best result.
Figure 4: Uncertainty dynamics for the LLFF fern scene with three training views: raw-uncertainty distributions (left), the fraction with u~i>0.8 (centre), and the uncertainty–opacity response at 6k iterations (right).
Figure 5: Photometric residuals for 3DGS*, DropGS, and FSGS, and UGOD’s predicted uncertainty. Rows show MiP-NeRF 360 bonsai (24 views) and LLFF horns and fern (3 views).
Symbol
Description
Value
κ
modulation slope
4.0
τg
modulation threshold
0.8
β
modulation floor (min. opacity multiplier)
0.70
Tg
modulation warm-up (iters)
1200
η
dropout scale
0.08
τd
dropout threshold
0.0
Table 1: Default hyper-parameters for the uncertainty head, opacity modulation, detached soft dropout, and early stopping.
Scene
Method
Init
Metrics
Final Gaussians
Reduction
PSNR ↑
SSIM ↑
LPIPS ↓
vs 3DGS*
bicycle_8
3DGS*
3
11.01
0.173
0.592
1,460,812
-
DropGS
3
12.24
0.208
0.597
1,000,235
1.5 × fewer
FSGS (w/ depth)
3
11.05
0.202
0.626
142,066
10.3 × fewer
Ours (full)
3
13.11
0.284
0.643
59,483
24.6 × fewer
bonsai_8
3DGS*
11
11.9641
0.3527
0.5958
565,053
-
Table 2: Per-scene results on MiP-NeRF 360 (8 views). The full model is shaded; best and second-best values are bold and underlined.
Scene
Method
Init
Metrics
Final Gaussians
Reduction
PSNR ↑
SSIM ↑
LPIPS ↓
vs 3DGS*
bicycle_24
3DGS*
21
15.3990
0.2815
0.5713
1,292,646
-
DropGS
21
15.8850
0.3172
0.5894
937,115
1.4 × fewer
FSGS (w/ depth)
21
15.8226
0.3482
0.6068
155,215
8.3 × fewer
Ours (full)
21
16.8980
0.3602
0.5954
141,530
9.1 × fewer
bonsai_24
3DGS*
5764
22.5979
0.8081
0.2986
686,826
-
Table 3: Per-scene results on MiP-NeRF 360 (24 views). The full model is shaded; best and second-best values are bold and underlined.
Scene
Method
Init
Metrics
Final
Compression
Points
PSNR ↑
SSIM ↑
LPIPS ↓
Gaussians
vs 3DGS*
Fern
3DGS*
1707
13.99
0.444
0.421
325,283
–
DropGS
1707
13.69
0.448
0.425
245,406
1.3 × fewer
FSGS (w/ depth)
1707
20.49
0.665
0.243
117,432
2.8 × fewer
Ours (full)
1707
20.74
0.680
0.231
93,766
3.5 × fewer
Flower
3DGS*
1804
16.81
0.479
0.371
233,646
–
Table 4: Per-scene results on LLFF (3 views). The full model is shaded; best and second-best values are bold and underlined.
Scene
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
Final Gaussians
vs 3DGS*
Ours (w/o soft dropout)
11.24
0.239
0.688
24,313
60.1 × fewer
Ours (w/ soft dropout)
10.56
0.246
0.689
14,400
101.4 × fewer
bicycle
Ours (full)
13.11
0.284
0.643
59,483
24.6 × fewer
Ours (w/o soft dropout)
13.9212
0.4432
0.5847
97,801
5.8 × fewer
Ours (w/ soft dropout)
12.6963
0.4385
0.6108
52,688
10.7 × fewer
bonsai
Ours (full)
13.7776
0.4309
0.5643
111,546
5.1 × fewer
Table 5: Per-scene component ablation on MiP-NeRF 360 (8 views). Best and second-best values are bold and underlined.
Scene
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
Final Gaussians
vs 3DGS*
bicycle
Ours (w/o soft dropout)
16.6135
0.3429
0.6026
84,962
15.2 × fewer
Ours (w/ soft dropout)
15.7267
0.3208
0.6139
76,494
16.9 × fewer
Ours (full)
16.8980
0.3602
0.5954
141,530
9.1 × fewer
bonsai
Ours (w/o soft dropout)
22.1005
0.7864
0.3320
337,290
2.0 × fewer
Ours (w/ soft dropout)
21.7632
0.7790
0.3339
344,817
2.0 × fewer
Ours (full)
20.1803
0.7038
0.4269
154,205
4.5 × fewer
Table 6: Per-scene component ablation on MiP-NeRF 360 (24 views). Best and second-best values are bold and underlined.
Scene
Variant
PSNR ↑
SSIM ↑
LPIPS ↓
Final Gaussians
vs 3DGS*
Fern
Ours
20.09
0.649
0.248
198,554
1.6 × fewer
Ours (w/o soft dropout)
19.06
0.605
0.277
204,963
1.6 × fewer
Ours (full)
20.74
0.680
0.231
93,766
3.5 × fewer
Flower
Ours
19.85
0.619
0.251
193,944
1.2 × fewer
Ours (w/o soft dropout)
18.97
0.591
0.271
192,774
1.2 × fewer
Ours (full)
19.96
0.616
0.264
93,361
2.5 × fewer
Table 7: Per-scene component ablation on LLFF (3 views). Best and second-best values are bold and underlined.
Advanced Micro Devices, Inc. · Dept. of Computer Science and Engineering University of Texas at Arlington · Dept. of Mathematics & Division of Data Science University of Texas at Arlington