Gaussian splatting has emerged as a flexible representation for 3D reconstruction from posed images. However, existing methods are optimized primarily using rasterization-based losses, which supervise a splat only when it contributes to sampled camera rays. Gaussians that are occluded or contribute little to the sampled view therefore receive weak or no geometric gradients and may drift away from the underlying surface, producing undesired floaters. We introduce PCAsplat, a geometry-aware regularization framework for Gaussian splatting based on differentiable local principal component analysis (PCA). Our PCA regularizer acts directly on neighborhoods of Gaussian centers and can therefore update Gaussians that do not contribute to the current training view. We regularize the PCA eigenvalues to encourage Gaussians to move to the underlying surface with isotropic tangent-plane coverage. We also align each Gaussian normal with the PCA-estimated neighborhood normal to enforce consistent orientation. Experiments on DTU, Tanks and Temples, and NeRF Synthetic show that the splats produced by PCAsplat better approximate samples of the reference surface while substantially reducing undesired floaters. These surface-aligned splats enable downstream geometry-processing tasks, including point cloud segmentation, and direct Poisson reconstruction. Additionally, PCAsplat remains competitive under conventional novel view synthesis and mesh extraction tasks. Code will be released.
Figures & tables
Figure 1: We present PCAsplat , a geometric regularization of Gaussian splatting that aligns Gaussians with surfaces through differentiable local PCA. Compared with current methods, PCAsplat substantially reduces floaters (left) and produces cleaner point clouds with better surface coverage (center). This improved point-cloud enables downstream tasks such as instance segmentation (right).
Figure 2: Overview of PCAsplat. We regularize Gaussian neighborhoods directly in scene space. For each Gaussian, we estimate a local neighborhood and compute differentiable PCA statistics that capture the underlying tangent structure. We then regularize the local eigenspectrum to reduce neighborhood thickness and promote isotropic tangent-plane coverage. We also align Gaussian normals with PCA-derived normals, encouraging surface-aligned and well-localized splats. Our formulation is lightweight, differentiable, and directly compatible with existing splatting pipelines.
Figure 3: Balanced coverage with isotropy (left); elongated structures without it (right).
Figure 4: Qualitative comparison on DTU. PCAsplat produces surface-aligned Gaussian centers with substantially fewer floaters, sharper edges, and finer geometric detail than the baselines.
Method
24
37
40
55
63
65
69
83
97
105
106
110
114
118
122
Mean
Time
VGGT- Ω
0.931
1.217
1.502
1.302
1.995
1.468
1.162
5.626
3.363
4.084
2.000
1.263
1.219
2.131
1.754
2.068
63s
COLMAP-D
0.382
0.618
0.336
0.349
0.570
0.837
0.455
1.129
0.650
0.540
0.418
0.468
0.300
0.387
0.348
0.519
26.9m
3DGS
1.056
1.148
1.396
1.035
1.285
1.609
1.354
1.836
1.349
1.005
1.662
1.536
1.037
1.274
1.077
1.311
5.2m
2DGS
0.906
1.879
1.065
0.920
1.261
2.351
1.399
1.377
1.395
0.949
1.231
1.525
1.356
1.308
1.051
1.332
10.9m
GS-Pull
1.243
0.803
1.127
0.719
1.129
1.698
1.554
1.744
1.493
0.954
1.891
1.791
0.806
1.395
1.107
1.297
21.8m
PGSR
0.621
0.651
0.817
0.611
0.804
0.780
0.729
1.020
0.800
0.647
0.801
0.651
0.510
0.666
0.571
0.712
30.5m
Table 1: Quantitative comparison on DTU. We report Chamfer distance ( ↓ ) between reconstructed and ground-truth point clouds. We highlight the best , second-best , and third-best Gaussian-based results, with PCAsplat achieving the best overall results.
Figure 5: Qualitative comparison of cross-section slices on challenging DTU scenes.
Chamfer Dist. ( ↓ )
Method
w/o PCA →
w/ PCA
2DGS
1.332 →
1.211
PGSR
0.712 →
0.667
MILo
0.821 →
0.754
RaDe-GS
0.662 →
0.519
Table 2: Effect of local PCA regularization across splatting methods on DTU.
Tanks and Temples
NeRF Synthetic Mesh
Method
Barn
Cater.
Court.
Igna.
Meet.
Truck
Mean
Time
Chair
Drums
Ficus
Hotdog
Lego
Mats.
Mic
Ship
Mean
Time
VGGT- Ω *
0.041
0.049
0.033
0.058
0.054
0.116
0.059
2.8
0.112
0.121
0.130
0.168
0.199
0.161
0.091
0.164
0.143
1.5
COLMAP-D
0.555
0.494
0.359
0.712
0.332
0.639
0.515
168.7
0.654
0.543
0.632
0.496
0.489
0.240
0.428
0.409
0.486
25.3
3DGS
0.076
0.080
0.047
0.128
0.056
0.133
0.087
7.9
0.336
0.392
0.543
0.113
0.284
0.161
0.303
0.215
0.293
4.8
2DGS
0.074
0.075
0.032
0.136
0.026
0.144
0.081
12.3
0.253
0.243
0.306
0.051
0.206
0.076
0.221
0.159
0.189
5.5
GSPull
0.071
0.057
0.074
0.105
0.012
0.106
0.071
113
0.119
0.103
0.160
0.009
0.036
0.028
0.077
0.006
0.067
26.4
Table 3: Quantitative comparison on Tanks and Temples and NeRF Synthetic . We report F1-score ( ↑ ) between reconstructed and ground-truth point clouds. We highlight the best , second-best , and third-best Gaussian-based results. Time is in minutes.
Figure 6: Qualitative comparison on Tanks and Temples. Compared with MILo, RaDe-GS, and GaussianWrapping, PCAsplat concentrates Gaussian centers along scene surfaces, with substantially fewer floaters around objects and below ground. See Fig. 12 and Fig. 11 for additional views of the geometry above and below ground.
LPCA
Lnormal
CD ↓
0.655
✓
0.654
✓
0.520
✓
✓
0.519
Table 4: Ablation study .
Figure 7: Qualitative results of 3D semantic instance segmentation. We apply Mask3D directly to colored point clouds reconstructed by each method on ScanNet++. PCAsplat produces more coherent object masks in the examples shown.
COLMAP-S
COLMAP-D
RaDe-GS
VGGT- Ω
Ours
Oracle
Top-1 mIoU
0.02%
8.21%
7.77%
10.54%
24.35%
30.41%
Top-3 mIoU
0.21%
8.20%
7.81%
10.58%
24.84%
30.56%
Table 5: Semantic instance segmentation on ScanNet++. We report Top-1 and Top-3 mIoU ( ↑ ) using Mask3D on point clouds reconstructed by each method.
Figure 8: Using Gaussian centers and normals from PCAsplat, SPSR recovers finer statue details and a smoother ground surface, with substantially fewer spurious mesh components than with RaDe-GS or GaussianWrapping.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Visible meshes from the NeRF Synthetic eight scenes.
Figure 10: Qualitative comparison of object isolation . Previous methods leave floating primitives and subsurface artifacts (red arrows) upon cropping, needing further manual editing to erase them. In contrast, PCAsplat (Ours) achieves clean, artifact-free boundaries, yielding an asset-ready reconstruction without manual cleanup.
Figure 11: “Underworld” geometry comparison on Tanks and Temples. We visualize the Gaussian means of reconstructed scenes from beneath the ground. Baselines fail to control interior floaters in unobserved regions and behind structures. PCAsplat (Ours) suppresses all floaters, resulting in regularized, clean ground and wall boundaries.
Figure 12: Qualitative point cloud detail comparison on Tanks and Temples Across diverse outdoor scenes, baseline methods produce noisy clouds. PCAsplat reconstructs scenes with high geometric fidelity, cleanly resolving intricate structures (e.g., roofs, vehicle contours, and ground details) while removing floating outliers.
Methods
24
37
40
55
63
65
69
83
97
105
106
110
114
118
122
Mean
Time
Neural
VolSDF ( Yariv et al., 2021 )
1.14
1.26
0.81
0.49
1.25
0.7
0.72
1.29
1.18
0.7
0.66
1.08
0.42
0.61
0.55
0.857
12h¿
Neus ( Wang et al., 2021 )
1.00
1.37
0.93
0.43
1.10
0.65
0.57
1.48
1.09
0.83
0.52
1.20
0.35
0.49
0.54
0.837
12h¿
Neuralangelo ( Li et al., 2023 )
0.37
0.72
0.35
0.35
0.87
0.54
0.53
1.29
0.97
0.73
0.47
0.74
0.32
0.41
0.43
0.610
12h¿
Gaussian Based
3DGS ( Kerbl et al., 2023 )
2.14
1.53
2.08
1.68
3.49
2.21
1.43
2.07
2.22
1.75
1.79
2.55
1.53
1.52
1.50
1.966
5.2m
SuGaR ( Guédon & Lepetit, 2024 )
1.47
1.33
1.13
0.61
2.25
1.71
1.15
1.63
1.62
1.07
0.79
2.45
0.98
0.88
0.79
1.330
52m
2DGS ( Huang et al., 2024 )
0.50
0.72
0.39
0.39
0.93
0.85
0.78
1.21
1.12
0.67
0.66
1.14
0.43
0.67
0.50
0.731
10.9m
Appendix
Table 6: Quantitative comparison on the DTU dataset . We report the Chamfer Distance (CD) ↓ across 15 scenes and the average optimization time ↓ . Our method achieves the best mean accuracy with a competitive performance.
Figure 13: Qualitative comparison on the DTU Dataset. Our method produces cleaner meshes with more details (e.g. teddy bear plush or house windows) and reduced artifacts (e.g. scissor tip, fruit specularity or skull forehead).
Methods
Barn
Caterp.
Court.
Ignat.
Meet.
Truck
Mean
Time
Neural
Neus
0.29
0.29
0.17
0.83
0.24
0.45
0.38
12h+
Geo-Neus
0.33
0.26
0.12
0.72
0.20
0.45
0.35
12h+
Neuralangelo
0.70
0.36
0.28
0.89
0.32
0.48
0.50
12h+
Gaussian Based
3DGS
0.13
0.08
0.09
0.04
0.01
0.19
0.09
7.9m
SuGaR
0.14
0.16
0.08
0.44
0.16
0.26
0.21
73m
2DGS
0.36
0.23
0.13
0.44
0.16
0.26
0.26
12.3m
Appendix
Table 7: Quantitative comparison on the Tanks and Temples dataset . We report F1-score ↑ . We evaluate all results using the scores reported in the papers. PCAsplat achieves competitive results while producing a cleaner underlying point cloud. * The results presented are from RaDe-GS paper. However, due to compute-limiting constraints, we were unable to re-run RaDe-GS TSDF extraction and achieve the same results as they show in their table. As such, we also append here the average results we got from running their code on our end, 0.49 , thus, our 0.50 improves over 0.49, and makes our method competitive with SoTA results.
Indoor Scenes
Outdoor Scenes
Method
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
3DGS
30.41
0.920
0.189
24.64
0.731
0.234
Mip-Splatting
30.90
0.921
0.194
24.65
0.729
0.245
BakedSDF
27.06
0.836
0.258
22.47
0.585
0.349
SuGaR
29.43
0.906
0.225
22.93
0.629
0.356
2DGS
30.40
0.916
0.195
24.34
0.717
0.246
Appendix
Table 8: Quantitative comparison across indoor and outdoor scenes of the MipNeRF 360 dataset ( Barron et al., 2022 ) . We highlight the best , second-best , and third-best PSNR, SSIM, and LPIPS. Our method maintains competitive performance while significantly improving Gaussian localization.
Figure 14: Curvature (Blue-Red) and Density (Red) splatting. Our method enables local curvature and local density rendering.
3D Gaussians have become a powerful scene representation for real-time splatting and high-quality novel-view synthesis. This has motivated generalizable splatting -- methods that adapt feed-forward geometry prediction networks to produce per-pixel Gaussians from a set of images. However, most generalizable splatting pipelines are supervised primarily through a view-synthesis loss to predict Gaussian orientation, anisotropic scale, opacity, and appearance in addition to their locations. We show that this learning objective is under-constrained. Models trained with view synthesis alone produce splats whose orientations and scales have no geometric connotation. The result is that, while producing decent view-synthesis performance, nearly all generalizable splatting methods produce geometrically inaccurate and misaligned Gaussians. We introduce G3Splat, a geometry-consistent generalizable splatting framework that addresses these degeneracies through differentiable geometric priors on the predicted 3D Gaussians, making the learning problem well-posed. These priors encourage the per-pixel splats to remain on their viewing rays and to orient themselves in accordance with local surfaces. Our priors are architecture-agnostic and can be incorporated into any previously studied geometric backbone for generalizable splatting, as well as different scene representations. We test G3Splat with both DUSt3R-style and VGGT-style backbones to predict pixel-aligned full-rank 3DGS as well as surfel-like 2DGS. Trained on RE10K, G3Splat produces Gaussian splats with significantly higher geometric fidelity than baselines, providing state-of-the-art novel-view depth, mesh reconstruction, and relative pose estimation performance while preserving novel-view synthesis quality, as evaluated on datasets such as ACID and ScanNet. Code and pretrained models are released on our project page.
Mehdi Hosseinzadeh, Shin-Fang Chng, Yi Xu +3
Australian Institute for Machine Learning · Goertek Alpha Labs · MBZUAI
Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict one Gaussian per input pixel, tying the representation budget to camera resolution rather than scene complexity. A flat wall and a richly textured object thus produce equally many Gaussians despite very different geometric needs. We propose ZipSplat, a token-based feed-forward model that decouples Gaussian placement from the pixel grid. A multi-view backbone extracts dense visual tokens, and k-means clustering compresses them into a compact set of scene tokens. Cross- and self-attention refine these tokens, and a lightweight MLP decodes each into a group of Gaussians with unconstrained 3D positions. Because clustering is applied at inference, a single trained model spans the quality-efficiency curve without retraining. ZipSplat operates without ground-truth poses or intrinsics, yet sets a new state of the art on DL3DV and RealEstate10K with ∼6× fewer Gaussians than pixel-aligned methods, surpassing the best pose-free baseline by 2.1dB and 1.2dB PSNR, respectively. It further generalizes zero-shot to Mip-NeRF360 and ScanNet++, outperforming all comparable baselines. Our project page is at https://veichta.com/zipsplat.
We present StructSplat, a feed-forward and generalizable 3D Gaussian reconstruction framework that operates directly on uncalibrated images without requiring camera parameters. Existing methods either rely on per-scene optimization or assume known camera poses, and often entangle geometry and appearance within a unified backbone, limiting reconstruction fidelity and generalization. Our key idea is to adopt a structured representation that organizes geometry, semantic, and texture cues with explicit roles in the reconstruction process. Specifically, we introduce a pixel-aligned feature injection mechanism to enable accurate texture modeling from 2D observations, incorporate semantic-aware priors to improve global consistency, and design a camera alignment strategy to prevent information leakage and improve generalization. Experiments show that our method significantly outperforms prior approaches on challenging benchmarks. On DL3DV, our method achieves 28.045 PSNR, surpassing AnySplat (22.377) by +5.67 dB. In cross-dataset evaluation, our method achieves +1.94 dB over AnySplat on ACID and +1.72 dB on RealEstate10K. Project page: https://structsplat.github.io Code: https://github.com/J-C-Zhao/StructSplat
Jia-Chen Zhao, Beiqi Chen, Xinyang Chen +2
Harbin Institute of Technology (Shenzhen) · Great Bay University · Guangzhou CloudButterfly Technology Co., Ltd.