The same Gaussian of a 3D Gaussian Splatting model is seen from many views, and these views do not always agree on the class it belongs to. The Gaussian may be occluded in some of them, and the confidence of the detector is not the same from one view to another. The ground truth, on the other hand, is given as an annotated mesh, because two training runs do not produce the same Gaussians. In this work, we propose a post-training lifting method that works with one target class at a time and combines the information coming from all the views. Target and non-target evidence are accumulated simultaneously, weighted by the visibility of each Gaussian in each view. After that, the Gaussians are filtered with two thresholds: a main threshold β selects the high-confidence seeds, and a lower one γβ adds the connected components around them. For the evaluation, the labels are transferred from the Gaussians to the mesh vertices that are both visible and annotated. With this design, we can separate three sources of error: the 2D detector, the lifting and the transfer between representations. The thresholds and the transfer operator are chosen on seven Replica validation scenes, and the method is evaluated on ten held-out ScanNet++ scenes with the same values for every scene and class. The mean mIoU on the validation scenes was 0.93 with masks from the dataset annotations and 0.65 with YOLO masks, and on the ScanNet++ test scenes it was 0.80 and 0.54. Compared with thresholding the evidence per view, as a previous version of the method did, the fraction improves the test mIoU by 0.24 and makes it possible to use a single threshold for all the classes and scenes of both datasets. Finally, the error analysis shows that most of the remaining error comes from the detector.
Figures & tables
Figure 2: The six stages of the method. Stages 2 and 3 do not depend on the class, and a change of β or γ only runs the pipeline again from stage 5.
Figure 3: Hysteresis on the graph of Gaussian centres. A component takes the class only when it contains a seed.
Figure 4: The two transfer operators at a vertex next to the border of an object. With θ below 0.5 , the radius vote gives the class to vertices where the nearest Gaussian does not, so the labelled region grows.
Parameter
Value
Model
Gaussian model iteration
30000
Masks
Detector score threshold
0.75
Mask binarisation
0.5
Voting
Table 1: Configuration fixed before the validation sweep. The last four rows only apply to the radius vote.
Dataset
Masks
Scenes
mIoU
95% CI
Precision
Recall
Reference
mIoUrel
Replica
annotation
7
0.93 ± 0.01
[0.92, 0.94]
0.96
0.97
0.97
0.956
YOLO
7
0.65 ± 0.13
[0.55, 0.74]
0.72
0.79
0.97
0.668
ScanNet++
annotation
10
0.80 ± 0.05
[0.77, 0.83]
0.86
0.92
0.91
0.876
YOLO
10
0.54 ± 0.09
[0.48, 0.60]
0.64
0.71
0.91
0.582
Table 2: Results at the selected point. The mIoU is the mean over scenes with its standard deviation, and the 95% confidence interval comes from the bootstrap of Section 5.4 .
Figure 5: Prediction against reference in two ScanNet++ test scenes close to the median, over all their evaluated classes. The mIoU is the one of the whole scene, measured on the mesh.
Replica
ScanNet++
Thresholded score
(β,γ)
annotation
YOLO
annotation
YOLO
Target evidence per view
(0.1,0.5)
0.64 ± 0.04
0.44 ± 0.07
0.56 ± 0.05
0.42 ± 0.04
Target evidence fraction
(0.7,0.8)
0.93 ± 0.01
0.65 ± 0.13
0.80 ± 0.05
0.54 ± 0.09
Table 3: mIoU of the baseline and of the method, each one at the point that the same rule chose for it.
Figure 6: Mean validation mIoU along the threshold grid of each score, with the annotation masks. The dotted line is the selected threshold.
Dataset
gdet
glift
grep
Replica
0.28
0.04
0.03
ScanNet++
0.26
0.11
0.09
Table 4: The three terms of Equation ( 21 ) at the selected point.
Figure 7: IoU of each class with each mask source and for the reference, averaged over the scenes that contain the class.
Figure 8: Pixel IoU of the YOLO masks against the 3D IoU that the method reaches from them, one marker per class and scene. Above the diagonal, the 3D result is better than the masks.
Figure 9: Mean validation mIoU along β for every γ , with the annotation masks. The grey band holds the candidates within 0.01 of the best mean.
Replica, validation
ScanNet++, test
Operator
(β,γ)
annotation
YOLO
reference
annotation
YOLO
reference
Nearest Gaussian
(0.7,0.8)
0.93 ± 0.01
0.65 ± 0.13
0.97
0.80 ± 0.05
0.54 ± 0.09
0.91
Radius vote
(0.9,0.7)
0.90 ± 0.04
0.63 ± 0.12
0.91
0.79 ± 0.05
0.55 ± 0.07
0.84
Table 5: Each transfer operator at the point that the rule chooses for it.
Configuration
mIoU
Difference
Reference
Radius vote at its selected point
0.90 ± 0.04
–
0.91
Hysteresis disabled
0.91 ± 0.04
+0.008
0.91
No competition, Gaussian to mesh
0.75 ± 0.05
-0.153
0.76
No competition, mesh to Gaussian
0.90 ± 0.04
+0.000
0.76
No competition, both directions
0.75 ± 0.05
-0.153
0.63
Opacity weighting disabled
0.90 ± 0.03
+0.002
0.91
Table 6: Ablations on the validation scenes with the annotation masks. Every row changes one factor of the radius vote at its selected point, and the difference is paired scene by scene.
Stage
Time
Peak memory
Dataset preparation
19.5 s
–
Mask generation
103 s
0.5 GB
Model training
617 s
1.9 GB
Vote accumulation
1359 s
0.8 GB
Threshold and hysteresis
6.4 s
–
Mesh transfer and metrics
257 s
–
Table 7: Cost of each stage for a whole test scene over the 8 scenes that ran every stage from scratch. The time is the median and the memory the highest peak. The last row runs the thirteen values of β from the cached votes.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Class
Annotation
YOLO
Reference
Scenes
Replica
chair
0.91
0.76
0.96
6
sofa
0.93
0.71
0.97
3
table
0.94
0.34
0.98
6
tv
0.95
0.92
0.98
2
plant
0.95
0.51
0.99
3
clock
0.93
0.92
0.96
3
Appendix
Table 8: IoU per class at the selected point, averaged over the scenes that contain the class.
Nearest Gaussian, (0.7,0.8)
Radius vote, (0.9,0.7)
Dataset
Scene
annotation
YOLO
reference
annotation
YOLO
reference
Replica
office_1
0.95
0.44
0.98
0.90
0.44
0.91
office_2
0.94
0.78
0.97
0.89
0.74
0.88
office_3
0.91
0.65
0.96
0.85
0.59
0.85
office_4
0.93
0.64
0.97
0.86
0.57
0.87
room_0
0.92
0.52
0.98
0.95
0.54
0.96
Appendix
Table 9: mIoU of every scene for each operator at its own selected point.
Figure 10: IoU of each class on the validation scenes along β , with the annotation masks, with and without hysteresis. The dotted line is β=0.7 .
Figure 11: Development sweep of the radius vote with the annotation masks. Solid lines are the prediction and dashed lines the reference.
Beijing University of Posts and Telecommunications, China · The University of Osaka, Japan · Beijing Key Laboratory of Multimodal Data Intelligent Perception and Governance, China +1