3D Gaussian Splatting (3DGS) enables real-time novel view synthesis, but existing general-purpose acceleration methods suffer severe rendering quality degradation when extended to more complex, large-scale scenes. To address this issue, we propose EffGS, a more general acceleration framework that improves training and rendering efficiency while maintaining reconstruction quality comparable to or better than vanilla 3DGS across bounded and large-scale scenes. EffGS combines frequency-aware guidance, localized density control, and adaptive primitive scale modulation. First, an importance scoring mechanism combines pixel-wise reconstruction errors with a difference-of-Gaussians mask scheduled over training to provide stage-dependent spatial guidance. Second, localized densification and pruning restricts density modifications to Gaussians with valid projected footprints in the sampled views. Third, learnable per-Gaussian scale modulation adjusts effective primitive extent during optimization while retaining the Compact Box rasterization rule. Extensive experiments on bounded and large-scale scene datasets demonstrate a favorable balance between reconstruction quality, training time, and primitive count. Component ablations and matched-primitive-budget comparisons further support the effectiveness of the framework.
Figures & tables
Figure 1: Compared with existing 3DGS acceleration methods, they can maintain high-quality rendering in small-scale scenes but suffer from degraded rendering quality when scaled up to city-level scenarios. In contrast, EffGS sustains high-fidelity rendering consistently, demonstrating superior scalability.
Figure 2: Spectral magnitude errors (left) and high-frequency discrepancy visualizations with ground-truth references (right) on Building ( Turki et al., 2022b ) ; see Sec. A.4 .
Figure 3: Overview of the proposed unified optimization pipeline. Given a set of sampled views, we first identify the visible Gaussians. We then compute a unified 2D supervision signal by fusing a pixel-wise error-aware mask with a frequency-aware mask (extracted via an annealing schedule). Finally, we project the visible Gaussians to 2D to accumulate importance scores from the unified mask, guiding localized densification and pruning.
Figure 4: Learnable primitive compactness. Compared with vanilla 3DGS ( Kerbl et al., 2023 ) , Speedy-Splat ( Hanson et al., 2025a ) , and FastGS ( Ren et al., 2025 ) , our method learns per-Gaussian compactness, reducing redundant pairs and improving fidelity.
Figure 5: Qualitative results of ours and other methods in image rendering on Deep Blending ( Park et al., 2019 ) , Mip-NeRF 360 ( Barron et al., 2022 ) and Tanks & Temples datasets ( Knapitsch et al., 2017 ) .
Figure 6: Qualitative results of ours and other methods in image rendering on Mill-19 ( Turki et al., 2022b ) , Urbanscene3D ( Lin et al., 2022 ) and GauU-Scene datasets ( Xiong et al., 2024 ) .
Method
Mip-NeRF 360
Deep Blending
Tanks & Temples
Time ↓
PSNR ↑
SSIM ↑
LPIPS ↓
NGS↓
FPS ↑
Time ↓
PSNR ↑
SSIM ↑
LPIPS ↓
NGS↓
FPS ↑
Time ↓
PSNR ↑
SSIM ↑
LPIPS ↓
NGS↓
FPS ↑
3DGS
20.93
27.53
0.812
0.221
2.63M
146
19.77
29.71
0.903
0.241
2.46M
158
11.34
23.71
0.850
0.170
1.57M
195
3DGS-LM
10.37
27.42
0.810
0.222
3.36M
220
9.83
29.61
0.905
0.248
2.71M
248
6.15
23.53
0.842
0.184
1.81M
289
Mini-Splatting
14.67
27.32
0.821
0.217
0.53M
567
13.35
29.99
0.907
0.244
0.56M
624
9.06
23.46
0.844
0.181
0.30M
756
Speedy-Splat
13.38
26.91
0.781
0.295
0.30M
552
10.75
29.42
0.898
0.272
0.25M
664
6.32
23.38
0.816
0.242
0.18M
691
Taming-3DGS
5.14
27.48
0.794
0.261
0.68M
221
3.06
29.50
0.894
0.278
0.29M
352
2.71
23.89
0.833
0.214
0.32M
379
Table 1: Quantitative comparison on Mip-NeRF 360 ( Barron et al., 2022 ) , Deep Blending ( Hedman et al., 2018 ) , and Tanks & Temples ( Knapitsch et al., 2017 ) .
Method
Mill-19
UrbanScene3D
GauU-Scene
Time (m) ↓
SSIM ↑
PSNR ↑
LPIPS ↓
NGS↓
FPS ↑
Time (m) ↓
SSIM ↑
PSNR ↑
LPIPS ↓
NGS↓
FPS ↑
Time (m) ↓
SSIM ↑
PSNR ↑
LPIPS ↓
NGS↓
FPS ↑
3DGS
196
0.735
23.06
0.292
10.76
<50
180
0.763
21.69
0.252
6.12
<50
176
0.736
23.66
0.268
6.58
<50
PGSR
212
0.603
20.12
0.447
12.41
<50
177
0.780
20.01
0.285
5.87
<50
187
0.593
20.57
0.421
5.98
<50
Mip-Splatting
187
0.703
22.35
0.334
18.44
<50
192
0.775
21.37
0.266
13.07
<50
184
0.660
21.40
0.341
14.50
<50
Taming-3DGS
21
0.551
21.22
0.510
0.75
88
23
0.633
19.74
0.460
0.48
73
29
0.714
23.77
0.308
5.92
84
Speedy-Splat
45
0.522
19.70
0.547
0.54
204
54
0.658
19.52
0.432
0.56
181
45
0.646
22.42
0.401
0.73
185
Table 2: Quantitative comparison of the averaged metrics on Mill-19 ( Turki et al., 2022b ) , UrbanScene3D ( Lin et al., 2022 ) , and GauU-Scene ( Xiong et al., 2024 ) .
Ablation Item
Rendering Quality
Efficiency Metrics
SSIM ↑
PSNR ↑
LPIPS ↓
Time (Min) ↓
Size (GB) ↓
GS (M) ↓
Mem (G) ↓
(a) w/o Frequency-aware Mask
0.752
24.12
0.278
22
1.15
4.51
15.6
(b) w/o Error-aware Mask
0.748
23.98
0.284
21
1.23
4.42
14.8
(c) w/o Learnable Compactness ( γi )
0.745
23.85
0.288
28
1.25
4.98
16.5
(d) w/o Local Densification and Pruning (LDP)
0.741
23.75
0.295
19
0.98
3.85
14.2
(e) w/o Frequency-aware Loss ( Lfreq )
0.753
24.16
0.276
19
1.05
4.20
15.3
Table 3: Component ablations on GauU-Scene ( Xiong et al., 2024 ) , evaluating reconstruction quality and computational cost.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Low-, mid-, and high-frequency spectral magnitude errors across scenes.
Method
Bicycle
Flowers
Garden
Stump
Treehill
Room
Counter
Kitchen
Bonsai
3DGS
25.14
21.30
27.34
26.64
22.59
31.71
29.16
31.54
32.37
Mini-Splatting
25.23
21.43
27.36
26.80
22.76
31.48
28.65
31.05
31.24
Speedy-Splat
24.79
21.21
26.69
26.67
22.48
30.83
28.22
30.09
31.16
Taming-3DGS
24.72
21.10
27.42
26.05
22.92
31.64
29.20
31.84
32.40
DashGaussian
25.31
21.78
27.57
27.17
22.94
31.81
29.11
31.69
32.15
FastGS
24.84
21.21
27.20
26.65
22.94
31.98
29.15
31.87
32.19
Appendix
Table 4: Quantitative PSNR comparison on Mip-NeRF 360 ( Barron et al., 2022 ) . The best, second-best, and third-best results are indicated by light red, light orange, and light yellow backgrounds, respectively.
Figure 8: Band-wise spectral magnitude errors over 30,000 training iterations on Rubble.
Method
Bicycle
Flowers
Garden
Stump
Treehill
Room
Counter
Kitchen
Bonsai
3DGS
0.748
0.586
0.857
0.768
0.636
0.927
0.915
0.932
0.946
Mini-Splatting
0.764
0.614
0.806
0.839
0.656
0.928
0.911
0.930
0.943
Speedy-Splat
0.704
0.560
0.814
0.765
0.590
0.903
0.876
0.895
0.925
Taming-3DGS
0.693
0.552
0.851
0.729
0.628
0.917
0.909
0.929
0.942
DashGaussian
0.763
0.604
0.857
0.783
0.640
0.924
0.911
0.927
0.945
FastGS
0.714
0.560
0.836
0.756
0.612
0.920
0.907
0.929
0.942
Appendix
Table 5: Quantitative SSIM comparison on Mip-NeRF 360 ( Barron et al., 2022 ) . The best, second-best, and third-best results are indicated by light red, light orange, and light yellow backgrounds, respectively.
Method
Bicycle
Flowers
Garden
Stump
Treehill
Room
Counter
Kitchen
Bonsai
3DGS
0.242
0.360
0.122
0.244
0.347
0.197
0.183
0.116
0.180
Mini-Splatting
0.241
0.341
0.215
0.161
0.326
0.190
0.181
0.120
0.177
Speedy-Splat
0.333
0.418
0.214
0.288
0.462
0.258
0.259
0.195
0.228
Taming-3DGS
0.332
0.416
0.138
0.324
0.395
0.227
0.200
0.128
0.193
DashGaussian
0.222
0.341
0.131
0.229
0.333
0.205
0.191
0.129
0.180
FastGS
0.310
0.406
0.174
0.297
0.429
0.217
0.204
0.127
0.191
Appendix
Table 6: Quantitative LPIPS comparison on Mip-NeRF 360 ( Barron et al., 2022 ) . The best, second-best, and third-best results are indicated by light red, light orange, and light yellow backgrounds, respectively.
Method
bicycle
flowers
garden
stump
treehill
room
counter
kitchen
bonsai
3DGS
27.97
18.98
26.78
21.77
20.33
18.78
17.58
21.30
14.90
Mini-Splatting
16.17
17.22
15.97
16.52
17.05
18.00
9.83
10.23
11.02
Speedy-Splat
15.87
13.38
15.73
13.77
12.90
12.05
12.10
13.25
11.37
Taming-3DGS
5.65
4.97
9.82
3.93
5.37
3.88
4.60
3.48
4.60
DashGaussian
9.93
7.05
8.27
6.57
8.20
4.00
3.95
5.52
3.95
FastGS
1.92
1.95
2.47
1.72
1.72
1.62
1.83
2.42
1.83
Appendix
Table 7: Quantitative results (Time) on Mip-NeRF 360 ( Barron et al., 2022 ) .
Method
Mill-19
UrbanScene3D
GauU-Scene
Building
Rubble
Residence
Sci-Art
Russian Building
Residence+
Modern Building
VastGaussian †
0.725
0.745
0.712
0.765
0.781
0.738
0.789
CityGaussian
0.776
0.814
0.810
0.835
0.801
0.755
0.791
CityGS-v2
0.661
0.724
0.771
0.808
0.792
0.741
0.762
CityGS-X*
0.771
0.803
0.802
0.829
0.793
0.742
0.774
3DGS
0.722
0.748
0.782
0.743
0.770
0.686
0.751
Appendix
Table 8: Quantitative SSIM comparison ( ↑ ) across different datasets. Building and Rubble are from Mill-19 ( Turki et al., 2022b ) ; Residence and Sci-Art are from UrbanScene3D ( Lin et al., 2022 ) ; Russian Building, Residence+, and Modern Building are from GauU-Scene ( Xiong et al., 2024 ) . The best, second-best, and third-best results are indicated by light red, light orange, and light yellow backgrounds, respectively. † denotes results obtained without decoupled appearance encoding, while * denotes results obtained without depth supervision.
Figure 9: Quantitative comparison of computational overhead on the Mill-19 ( Turki et al., 2022b ) , Urbanscene3D ( Lin et al., 2022 ) , and GauU-Scene datasets ( Xiong et al., 2024 ) . ∗ indicates multi-GPU training.
Method
Mill-19
UrbanScene3D
GauU-Scene
Building
Rubble
Residence
Sci-Art
Russian Building
Residence+
Modern Building
VastGaussian †
21.83
25.24
21.06
22.59
23.98
23.41
25.53
CityGaussian
21.56
25.79
22.03
22.42
24.11
23.65
26.01
CityGS-v2
19.85
24.02
21.23
20.71
24.04
23.43
25.78
CityGS-X*
21.78
25.45
22.11
22.32
24.18
23.91
25.95
3DGS
20.62
25.49
21.45
21.92
23.74
22.10
25.15
Appendix
Table 9: Quantitative PSNR comparison ( ↑ ) across different datasets. Building and Rubble are from Mill-19 ( Turki et al., 2022b ) ; Residence and Sci-Art are from UrbanScene3D ( Lin et al., 2022 ) ; Russian Building, Residence+, and Modern Building are from GauU-Scene ( Xiong et al., 2024 ) . The best, second-best, and third-best results are indicated by light red, light orange, and light yellow backgrounds, respectively. † denotes results obtained without decoupled appearance encoding, while * denotes results obtained without depth supervision.
Figure 10: Visualization of the frequency-aware mask at different schedule parameters f on a fixed input image. Top: continuous mask scores before thresholding. Middle: score heatmaps overlaid on the input. Bottom: binary masks obtained with a threshold of 0.5, where white indicates selected pixels. Scores are normalized independently for each setting.
Figure 11: Peak VRAM usage (top) and step-wise changes (bottom) during frequency extraction with varying resolutions and batch size 10 on Russian ( Xiong et al., 2024 ) .
Figure 12: Qualitative comparisons on the Mill-19, Urbanscene3D, and GauU-Scene datasets ( Turki et al., 2022b ; Lin et al., 2022 ; Xiong et al., 2024 ) .
Figure 13: Qualitative comparisons on the Deep Blending, Mip-NeRF 360, and Tanks & Temples datasets ( Park et al., 2019 ; Barron et al., 2022 ; Knapitsch et al., 2017 ) .
Method
Mill-19
UrbanScene3D
GauU-Scene
Building
Rubble
Residence
Sci-Art
Russian Building
Residence+
Modern Building
VastGaussian †
0.271
0.268
0.263
0.259
0.251
0.301
0.245
CityGaussian
0.267
0.226
0.213
0.232
0.223
0.289
0.253
CityGS-v2
0.382
0.311
0.235
0.265
0.234
0.294
0.254
CityGS-X*
0.256
0.224
0.221
0.241
0.220
0.284
0.234
3DGS
0.304
0.279
0.238
0.265
0.234
0.321
0.249
Appendix
Table 10: Quantitative LPIPS comparison ( ↓ ) across different datasets. Building and Rubble are from Mill-19 ( Turki et al., 2022b ) ; Residence and Sci-Art are from UrbanScene3D ( Lin et al., 2022 ) ; Russian Building, Residence+, and Modern Building are from GauU-Scene ( Xiong et al., 2024 ) . The best, second-best, and third-best results are indicated by light red, light orange, and light yellow backgrounds, respectively. † denotes results obtained without decoupled appearance encoding, while * denotes results obtained without depth supervision.
Figure 14: Quantitative comparison of computational overhead on the Deep Blending ( Park et al., 2019 ) , Mip-NeRF 360 ( Barron et al., 2022 ) , and Tanks & Temples datasets ( Knapitsch et al., 2017 ) .
Method
SSIM ↑
PSNR ↑
LPIPS ↓
FastGS
0.741
23.48
0.294
FastGS+LDP
0.758
23.89
0.257
Appendix
Table 11: Ablation study on the Russian scene from the GauU-Scene dataset ( Xiong et al., 2024 ) .
Method
Residence+
Russian Building
Modern Building
Time (min)
GS (M)
FPS
Time (min)
GS (M)
FPS
Time (min)
GS (M)
FPS
CityGaussian
261
8.14
68
214
7.02
55
217
7.90
57
CityGS-v2
203
8.04
46
182
6.97
33
188
7.90
35
Speedy-splat
50
0.50
195
43
0.82
189
43
0.87
170
FastGS
14
2.91
252
12
1.89
231
12
2.23
224
EffGS
19
4.86
131
21
3.45
127
23
4.37
115
Appendix
Table 12: Efficiency comparison on Residence+, Russian Building, and Modern Building scenes from the GauU-Scene dataset ( Xiong et al., 2024 ) .
Method
Deep Blending
Tanks & Temples
Time ↓
PSNR ↑
SSIM ↑
LPIPS ↓
NGS
Time ↓
PSNR ↑
SSIM ↑
LPIPS ↓
NGS
Uncontrolled Gaussian‑point count
EDGS
30.00
29.81
0.904
0.223
–
23.00
24.28
0.868
0.132
–
3DGS-MCMC
19.00
29.56
0.902
0.244
–
13.00
24.22
0.863
0.158
–
EffGS (Ours)
2.94
30.01
0.910
0.238
0.25
1.94
24.22
0.861
0.166
0.44
Matched Gaussian budgets: controlled comparison
Appendix
Table 13: Comparison on Deep Blending and Tanks & Temples. Top: reported settings. Bottom: Gaussian budgets matched to EffGS. Time is in minutes and NGS in millions; – denotes unavailable counts. Best quality and timing results within each block and dataset are bold.
Method
Mill-19
UrbanScene3D
Time ↓
PSNR ↑
SSIM ↑
LPIPS ↓
NGS
Time ↓
PSNR ↑
SSIM ↑
LPIPS ↓
NGS
Taming-3DGS
57
22.52
0.693
0.384
3.76
61
21.14
0.719
0.332
3.05
EDGS+3DGS
83
22.57
0.716
0.372
3.76
91
21.17
0.716
0.330
3.05
EffGS (Ours)
20
23.68
0.755
0.279
3.76
22
21.50
0.778
0.281
3.05
Appendix
Table 14: Matched-budget comparison on Mill-19 and UrbanScene3D. EDGS+3DGS uses matched initialization with densification disabled. Time is in minutes and NGS in millions. Best quality and timing results are bold.
Method
Frequency Error
Reconstruction Quality
NGS↓
Low ↓
Mid ↓
High ↓
PSNR ↑
SSIM ↑
LPIPS ↓
w/o Learnable Compactness ( γi )
86726.7
30836.8
7073.6
26.99
0.810
0.264
2.30
EffGS (Ours)
72713.2
25773.5
5678.1
28.51
0.855
0.194
2.02
Appendix
Table 15: Extended ablation of learnable primitive compactness, averaged over Garden, Room, Bonsai, Rubble, and Modern Building. NGS is reported in millions. The best results are shown in bold.
3D Gaussian Splatting achieves exceptional real-time rendering, but its substantial computational and storage demands hinder widespread deployment. Existing accelerated paradigms often aggressively prune primitives for rapid convergence, causing severe loss of high-frequency details. To address this, we tackle the fundamental problem of achieving both exceptional rendering quality and ultra-fast reconstruction speed. In this paper, we propose ACE-GS, a progressive optimization framework tailored for accurate, compressed, and efficient scene representation. We realize that precise primitive management is the key to breaking this trade-off. Therefore, we first design a momentum consistency-guided densification strategy, strictly constraining primitive growth onto authentic geometric manifolds to avoid computational waste while significantly accelerating convergence. Building upon this efficient initialization, we deploy a statistical sensitivity-driven sparsification mechanism to precisely prune redundant primitives, yielding a further compressed footprint. Finally, to thoroughly compensate for the risk of micro-structure loss caused by the aforementioned strict primitive control, we introduce a cross-dimensional residual frequency compensation scheme that explicitly back-injects high-frequency error energy into primitive attributes, perfectly restoring sharp geometric details. Extensive experiments validate our superiority. While maintaining a highly compact scene representation, our system achieves up to 3.7 times training acceleration against the rapid framework Speedy-Splat. Requiring only 3 to 5 minutes to converge, ACE-GS secures the highest structural similarity and achieves a peak PSNR improvement of up to 0.89 dB over the original 3DGS, establishing a new benchmark for ultra-fast and high-fidelity novel view synthesis.
Jijian Zhao
Huazhong University of Science and Technology, Wuhan, China
3D Gaussian Splatting (3DGS) achieves real-time novel-view synthesis by optimizing millions of anisotropic Gaussians, yet its training remains expensive, with the backward pass dominating runtime in the post-densification refinement phase. We observe substantial update redundancy in this phase: many sampled views have near-plateaued losses and provide diminishing gradient benefits, but standard training still runs full backpropagation. We propose SkipGS with a novel view-adaptive backward gating mechanism for efficient post-densification training. SkipGS always performs the forward pass to update per-view loss statistics, and selectively skips backward passes when the sampled view's loss is consistent with its recent per-view baseline, while enforcing a minimum backward budget for stable optimization. On Mip-NeRF 360, compared to 3DGS, SkipGS reduces end-to-end training time by 23.1%, driven by a 42.0% reduction in post-densification time, with comparable reconstruction quality. Because it only changes when to backpropagate without modifying the renderer, representation, or loss, SkipGS is plug-and-play and compatible with other complementary efficiency strategies, enabling additive speedups. Code is available at https://github.com/ASU-ESIC-FAN-Lab/SkipGS.
We present BlitzGS, a distributed 3DGS framework that reduces active Gaussian workload for fast city-scale reconstruction. BlitzGS manages this workload at three coupled levels. At the system level, the framework shards Gaussians across GPUs by index parity rather than spatial blocks. This approach mitigates the cross-block visibility redundancy inherent in spatial partitioning. Furthermore, it distributes each rendering step through a single cross-GPU exchange that routes projected Gaussians to their tile owners. At the model level, scheduled importance-scoring passes shrink the global Gaussian population. During these passes, the framework generates a per-Gaussian visibility weight to bias density-control updates toward contributing primitives and a per-view importance mask for the view-level renderer. At the view level, BlitzGS trims each camera's active set with a distance-based LOD gate to exclude excessively fine primitives for the current frustum and the importance-based culling mask to skip Gaussians with negligible cross-view contribution. On large-scale benchmarks, BlitzGS matches the rendering quality of recent large-scale baselines while delivering an order-of-magnitude speedup, training city-scale scenes in tens of minutes. Our code is available at https: //github.com/AkierRaee/BlitzGS.