Conventional 3D Gaussian Splatting (3DGS) requires depth sorting and ordered alpha blending to correctly render overlapping Gaussian primitives. Stochastic transparency enables sorting-free rendering by replacing fractional alpha contributions with discrete stochastic visibility samples, but produces substantial spatial and temporal noise at low sample counts. We refer to this conversion from continuous Gaussian splats to discrete visibility samples as \textit{Gaussian Stippling}. Based on this, we present an efficient order-independent rendering and reconstruction framework that operates directly on unmodified 3DGS assets. Our method adaptively integrates primitive-based and fragment-based stippling, leveraging their complementary strengths across different rendering regimes to significantly improve rendering throughput. To recover high-quality images from sparse stochastic samples, we further introduce a lightweight Gaussian-aware spatiotemporal reconstruction network. By exploiting the Gaussian attributes retained by each stipple, the network aggregates structured stochastic clues across both space and time, effectively suppressing stippling noise. Experiments show that our hybrid Gaussian stippling method, coupled with a spatiotemporal reconstruction network trained on diverse scenes, generalizes to unseen scenes and enables interactive, temporally stable, and visually plausible rendering on mobile devices without retraining or preprocessing. With scene-specific training and appropriately scaled sampling and network capacity, our method further outperforms the baselines in visual quality, offering a high-fidelity configuration for quality-prioritized applications.
Figures & tables
Figure 1: Sorting-free stochastic rendering achieves high throughput but produces substantial spatial noise (top). We augment its observations with Gaussian attributes and reconstruct them using a lightweight spatiotemporal network.
Figure 2: Overview of our proposed Gaussian stippling and reconstruction pipeline. In Stage 1, given a 3DGS asset containing N Gaussian primitives, a cost-aware routing strategy assigns NF primitives to the fragment-based stream and the remaining primitives to the primitive-based stream, which generates NP stochastic sample points. A shared depth test then selects accepted stipples together with their Gaussian-local descriptors to generate the current-frame observation. In Stage 2, the reconstruction network aggregates the current observation and forward-reprojected historical observations to produce the reconstructed image.
Figure 3: Time-cost ratio of two streams for Layer L=1 . The red solid line is the fitted boundary and the gray dotted line is the measured boundary.
Figure 4: Per-view latency. Ours has the lowest latency among the compared methods on the evaluated views and reduces the viewpoint-dependent variation observed with Gaussian Point Splatting.
Method
Mip-NeRF360
Tanks&Temples
Deep Blending
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
StochasticSplats
17.29
0.2427
0.8942
664.95
13.30
0.1754
0.9961
843.08
21.61
0.3562
0.8063
556.03
Gaussian Point Splatting
17.28
0.2413
0.8947
620.56
13.30
0.1749
0.9961
440.99
21.60
0.3559
0.8065
438.95
Ours (1-spp)
17.28
0.2407
0.8947
880.07
13.29
0.1738
0.9965
1001.78
21.61
0.3559
0.8065
1024.15
3DGS
28.90
0.8708
0.1856
122.18
23.39
0.8421
0.1837
113.19
29.52
0.9038
0.2459
111.18
MobileGS †
28.20
0.8570
0.2103
234.97
- ∗
- ∗
- ∗
- ∗
30.07
0.9109
0.2486
292.26
Table 1: Quality against ground truth at the training resolution; throughput at 1080p. The first three rows report renderer-only FPS; Ours-S/L rows include rendering and reconstruction. Best/second-best PSNR and FPS within each block are bold/underlined; ties share rank. ∗ MobileGS optimization produced NaNs on Tanks&Temples. † MobileGS FPS includes its view-dependent MLP (TensorRT) and rasterization; its released timing code omits the MLP.
Figure 5: Qualitative results of 1-spp Ours-S . A compact reconstruction network suppresses stochastic noise using Gaussian-aware observations from the current and historical frames.
Figure 6: Qualitative results of 16-spp Ours-L , illustrating image reconstruction with a higher observation budget and a larger network.
Variant
Bonsai
Counter
Train
Avg.
Full descriptor
30.749
28.263
21.450
26.820
RGB only
30.376
28.093
21.117
26.529
w/o alpha
30.732
28.253
21.439
26.808
w/o conic
30.653
28.197
21.373
26.741
w/o opacity
30.758
28.249
21.421
26.809
w/o dM2
30.710
28.263
21.453
26.809
Table 2: Ablation of temporal history and Gaussian-local descriptors using Ours-S . We evaluate variants that remove all historical samples, keep only a single history frame, or omit specific descriptor channels from the full input. All values are PSNR evaluated against the standard 3DGS target.
Method
Flow.
House
Pavil.
Plant
Stat.
Avg.
3DGS
22.73
23.10
21.25
22.76
21.62
22.29
Ours (1-spp)
13.47
13.45
12.67
13.92
13.27
13.36
SS (4-spp)
17.80
17.86
16.86
18.23
17.39
17.63
GPS (4-spp)
17.83
17.87
16.86
18.22
17.51
17.66
Ours (1-spp + S)
21.74
22.35
21.01
21.60
20.94
21.53
Table 3: PSNR on held-out DL3DV views with respect to photographic ground truth. Avg. is the arithmetic mean over the five scenes. The best and second-best values are shown in bold and underlined, respectively. SS (4-spp), GPS (4-spp), and Ours (1-spp + S) run at 353.32, 282.91, and 353.43 FPS, respectively.
Method
Truck
Room
DrJohnson
Avg.
3DGS-GL
5.46
12.8
3.2
7.15
SS
24.8
65.8
69.4
53.3
GPS
41.3
52.8
53.0
49.03
SortfreeGS
37.2
50.3
27.2
38.2
MobileGS
–
32.6
28.1
30.4
Ours(1-spp)
78.8
112.7
86.3
92.6
Table 4: Mobile 540p throughput (FPS) on vivo x300pro with MediaTek tianji9500, using one scene per benchmark. Our renderer uses Vulkan, and Ours-S runs in FP16 on the NPU. Avg. is the arithmetic mean over available scene results. The best and second-best values are shown in bold and underlined, respectively.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Truck
Room
DrJohnson
Avg.
3DGS-GL
5.46
12.8
3.2
7.15
SS
24.8
65.8
69.4
53.3
GPS
41.3
52.8
53.0
49.03
SortfreeGS
32.94
22.03
31.95
28.97
MobileGS
–
26.98
25.61
26.29
Ours (1-spp)
78.8
112.7
86.3
92.6
Appendix
Table A1: Mobile 540p throughput (FPS) on vivo x300pro with MediaTek tianji9500, using one scene per benchmark. Our renderer uses Vulkan, and Ours-S runs in FP16 on the NPU. Avg. is the arithmetic mean over available scene results. The best and second-best values are shown in bold and underlined, respectively.
Figure A1: Crossover boundaries ( tprimitive=tfragment ) for varying layer counts L∈{1,…,8} plotted over the opacity-area plane. The almost perfectly overlapping curves indicate that depth complexity has a negligible effect on the relative performance of the two streams.
Scene
Small (1 spp)
Large (1 spp)
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Bicycle
24.0400
0.6710
0.2600
24.6107
0.6996
0.2745
Bonsai
30.7500
0.9110
0.1600
32.1253
0.9284
0.2303
Counter
28.2700
0.8690
0.1890
28.9801
0.8927
0.2317
Garden
25.4000
0.7470
0.1750
26.0137
0.7769
0.1891
Kitchen
29.2100
0.8750
0.1550
30.0836
0.8948
0.1703
Appendix
Table A2: Per-scene and per-dataset reconstruction quality of the Small and Large variants at 1 spp, measured against ground-truth images. Dataset averages are arithmetic means over their constituent scenes; the final row is the macro-average over all 11 scenes.
Scene
Small (4 spp)
Large (4 spp)
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Bicycle
24.5220
0.7018
0.2799
24.9213
0.7212
0.2614
Bonsai
31.2651
0.9232
0.2346
32.4673
0.9317
0.2279
Counter
28.5755
0.8876
0.2357
29.1298
0.8981
0.2263
Garden
26.1494
0.7916
0.1903
26.5915
0.8104
0.1662
Kitchen
29.9029
0.8948
0.1701
30.5050
0.9057
0.1595
Appendix
Table A3: Per-scene and per-dataset reconstruction quality of the Small and Large variants at 4 spp, measured against ground-truth images.
Scene
Small (16 spp)
Large (16 spp)
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Bicycle
24.9336
0.7303
0.2550
25.1800
0.7470
0.2080
Bonsai
31.7702
0.9304
0.2228
32.8700
0.9350
0.1430
Counter
28.7965
0.8974
0.2180
29.2900
0.9000
0.1690
Garden
26.7682
0.8296
0.1487
27.0900
0.8450
0.1000
Kitchen
30.4227
0.9098
0.1461
30.7800
0.9150
0.1120
Appendix
Table A4: Per-scene and per-dataset reconstruction quality of the Small and Large variants at 16 spp, measured against ground-truth images. Dataset averages are arithmetic means over their constituent scenes; the final row is the macro-average over all 11 scenes.
Variant
Bonsai
Counter
Train
Avg.
Forward Nearest.
30.75
28.23
21.50
26.83
Forward Bilinear.
30.60
28.19
21.38
26.72
Backward Nearest.
30.59
28.21
21.40
26.73
Backward Bilinear.
30.68
28.25
21.40
26.78
Appendix
Table A5: Temporal transport under stochastic visibility. Backward gathering produces structured artifacts; bilinear transport spreads samples, whereas nearest-forward transport preserves discrete observations. The best and second-best PSNR (dB) values are shown in bold and underlined ; ties share the same rank.
Table A6: Scene names and corresponding dataset hash values in DL3DV.
Figure A2: After 1024spp rendering using GPS’s source code, the result was compared with that of Gaussian rendering, and there was a loss. The only modification to Gaussian Point Splatting is that we removed the upper limit on individual Gaussian sampling.
Figure A3: Rendering failure of Gaussian Point Splatting: its workload distribution loses much detail when aligning the 3DGS bounding box at conventional rendering resolution.
Configuration
Ours-S
Ours-L
Input / output channels
40/3
40/3
Down-/upsampling stages
3/3
4/4
Resolution pyramid
H to H/8
H to H/16
Encoder widths, fine to coarse
16,16,16,16
64,96,128,192,256
Encoder/bottleneck convolutions per resolution
2,1,1,2 (first stage: 1×1 , then 3×3 )
2,2,2,2,2 (all 3×3 )
Decoder transitions, coarse to fine
32→16→16 at H/4 ; 32→16→16 at H/2 ; 32→16→16 at H
384→192→192 at H/8 ; 288→128→128 at H/4 ; 192→96→96 at H/2 ; 128→64→64 at H
Appendix
Table A7: Detailed reconstruction-network architectures. A decoder transition a→b→c denotes the channel count after skip concatenation and the outputs of the two subsequent 3×3 convolutions. Both variants use concatenative encoder skips, a projected full-resolution input skip, max-pool downsampling, nearest-neighbor upsampling, and direct RGB prediction. Parameter counts include convolutional biases. FP16 kernel-weight sizes exclude biases, activations, and runtime workspace.
Figure A4: Temporal ablations in the evaluated stochastic-input setting. Backward reprojection exhibits structured artifacts, while forward reprojection leaves unfilled pixels.
Figure A5: Qualitative results of additional scenes of 16spp + Large network showing neural reconstruction quality
3D Gaussian Splatting (3DGS) has revolutionized novel-view synthesis with its fast and high-fidelity rendering. However, rendering at high FPS and low latency across various scenes remains a challenge, especially when large amounts of 3D Gaussian ellipsoids appear in the scene. To address this issue, we introduce TemporalGS, to the best of our knowledge, the first training-free plug-and-play algorithmic approach to accelerate 3DGS rendering without any post-training or post-processing, implemented on top of tile-based software rasterization. The key idea is that, instead of rendering frames independently as 3DGS, we leverage the temporal priors, represented by novel geometry and appearance buffers, etc., to reduce redundancy of Gaussian preprocessing, sorting, and rasterization operations of consecutive frames. Specifically, we propose two acceleration strategies: (1) temporal dynamic culling, which filters out Gaussians that contribute less to current frame rendering; (2) selective rendering, which renders only a small portion of tiles that cannot be approximated by the temporal priors. By adapting and interleaving these two strategies, TemporalGS yields a simple but effective plug-and-play solution for 3DGS rendering speed-up without any training. Extensive experiments show that TemporalGS achieves comparable or even better performance compared to existing state-of-the-art post-training or post-processing-based 3DGS rendering acceleration approaches. TemporalGS can significantly enhance the rendering speed of various 3DGS methods, achieving up to 1.48× acceleration, while maintaining competitive rendering quality. We further extend our TemporalGS to hardware rasterization-based 3DGS to show the portability of our algorithm.
Yuhongze Zhou, Zihao Yang, Xinxin Zuo +1
McGill University, Montr´eal, Canada · University of Waterloo, Waterloo, Canada · Concordia University, Montr´eal, Canada +1
We present BlitzGS, a distributed 3DGS framework that reduces active Gaussian workload for fast city-scale reconstruction. BlitzGS manages this workload at three coupled levels. At the system level, the framework shards Gaussians across GPUs by index parity rather than spatial blocks. This approach mitigates the cross-block visibility redundancy inherent in spatial partitioning. Furthermore, it distributes each rendering step through a single cross-GPU exchange that routes projected Gaussians to their tile owners. At the model level, scheduled importance-scoring passes shrink the global Gaussian population. During these passes, the framework generates a per-Gaussian visibility weight to bias density-control updates toward contributing primitives and a per-view importance mask for the view-level renderer. At the view level, BlitzGS trims each camera's active set with a distance-based LOD gate to exclude excessively fine primitives for the current frustum and the importance-based culling mask to skip Gaussians with negligible cross-view contribution. On large-scale benchmarks, BlitzGS matches the rendering quality of recent large-scale baselines while delivering an order-of-magnitude speedup, training city-scale scenes in tens of minutes. Our code is available at https: //github.com/AkierRaee/BlitzGS.
3D Gaussian Splatting (3DGS) enables real-time novel view synthesis, but existing general-purpose acceleration methods suffer severe rendering quality degradation when extended to more complex, large-scale scenes. To address this issue, we propose EffGS, a more general acceleration framework that improves training and rendering efficiency while maintaining reconstruction quality comparable to or better than vanilla 3DGS across bounded and large-scale scenes. EffGS combines frequency-aware guidance, localized density control, and adaptive primitive scale modulation. First, an importance scoring mechanism combines pixel-wise reconstruction errors with a difference-of-Gaussians mask scheduled over training to provide stage-dependent spatial guidance. Second, localized densification and pruning restricts density modifications to Gaussians with valid projected footprints in the sampled views. Third, learnable per-Gaussian scale modulation adjusts effective primitive extent during optimization while retaining the Compact Box rasterization rule. Extensive experiments on bounded and large-scale scene datasets demonstrate a favorable balance between reconstruction quality, training time, and primitive count. Component ablations and matched-primitive-budget comparisons further support the effectiveness of the framework.
Changbai Li, Shuo Yang, Yichen Yang +2
Beihang University · Nanyang Technological University