Conventional 3D Gaussian Splatting (3DGS) requires depth sorting and ordered alpha blending to correctly render overlapping Gaussian primitives. Stochastic transparency enables sorting-free rendering by replacing fractional alpha contributions with discrete stochastic visibility samples, but produces substantial spatial and temporal noise at low sample counts. We refer to this conversion from continuous Gaussian splats to discrete visibility samples as \textit{Gaussian Stippling}. Based on this, we present an efficient order-independent rendering and reconstruction framework that operates directly on unmodified 3DGS assets. Our method adaptively integrates primitive-based and fragment-based stippling, leveraging their complementary strengths across different rendering regimes to significantly improve rendering throughput. To recover high-quality images from sparse stochastic samples, we further introduce a lightweight Gaussian-aware spatiotemporal reconstruction network. By exploiting the Gaussian attributes retained by each stipple, the network aggregates structured stochastic clues across both space and time, effectively suppressing stippling noise. Experiments show that our hybrid Gaussian stippling method, coupled with a spatiotemporal reconstruction network trained on diverse scenes, generalizes to unseen scenes and enables interactive, temporally stable, and visually plausible rendering on mobile devices without retraining or preprocessing. With scene-specific training and appropriately scaled sampling and network capacity, our method further outperforms the baselines in visual quality, offering a high-fidelity configuration for quality-prioritized applications.
Figures & tables
Figure 1: Sorting-free stochastic rendering achieves high throughput but produces substantial spatial noise (top). We augment its observations with Gaussian attributes and reconstruct them using a lightweight spatiotemporal network.
Figure 2: Overview of our proposed Gaussian stippling and reconstruction pipeline. In Stage 1, given a 3DGS asset containing N Gaussian primitives, a cost-aware routing strategy assigns NF primitives to the fragment-based stream and the remaining primitives to the primitive-based stream, which generates NP stochastic sample points. A shared depth test then selects accepted stipples together with their Gaussian-local descriptors to generate the current-frame observation. In Stage 2, the reconstruction network aggregates the current observation and forward-reprojected historical observations to produce the reconstructed image.
Figure 3: Time-cost ratio of two streams for Layer L=1 . The red solid line is the fitted boundary and the gray dotted line is the measured boundary.
Figure 4: Per-view latency. Ours has the lowest latency among the compared methods on the evaluated views and reduces the viewpoint-dependent variation observed with Gaussian Point Splatting.
Method
Mip-NeRF360
Tanks&Temples
Deep Blending
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
PSNR ↑
SSIM ↑
LPIPS ↓
FPS ↑
StochasticSplats
17.29
0.2427
0.8942
664.95
13.30
0.1754
0.9961
843.08
21.61
0.3562
0.8063
556.03
Gaussian Point Splatting
17.28
0.2413
0.8947
620.56
13.30
0.1749
0.9961
440.99
21.60
0.3559
0.8065
438.95
Ours (1-spp)
17.28
0.2407
0.8947
880.07
13.29
0.1738
0.9965
1001.78
21.61
0.3559
0.8065
1024.15
3DGS
28.90
0.8708
0.1856
122.18
23.39
0.8421
0.1837
113.19
29.52
0.9038
0.2459
111.18
MobileGS †
28.20
0.8570
0.2103
234.97
- ∗
- ∗
- ∗
- ∗
30.07
0.9109
0.2486
292.26
Table 1: Quality against ground truth at the training resolution; throughput at 1080p. The first three rows report renderer-only FPS; Ours-S/L rows include rendering and reconstruction. Best/second-best PSNR and FPS within each block are bold/underlined; ties share rank. ∗ MobileGS optimization produced NaNs on Tanks&Temples. † MobileGS FPS includes its view-dependent MLP (TensorRT) and rasterization; its released timing code omits the MLP.
Figure 5: Qualitative results of 1-spp Ours-S . A compact reconstruction network suppresses stochastic noise using Gaussian-aware observations from the current and historical frames.
Figure 6: Qualitative results of 16-spp Ours-L , illustrating image reconstruction with a higher observation budget and a larger network.
Variant
Bonsai
Counter
Train
Avg.
Full descriptor
30.749
28.263
21.450
26.820
RGB only
30.376
28.093
21.117
26.529
w/o alpha
30.732
28.253
21.439
26.808
w/o conic
30.653
28.197
21.373
26.741
w/o opacity
30.758
28.249
21.421
26.809
w/o dM2
30.710
28.263
21.453
26.809
Table 2: Ablation of temporal history and Gaussian-local descriptors using Ours-S . We evaluate variants that remove all historical samples, keep only a single history frame, or omit specific descriptor channels from the full input. All values are PSNR evaluated against the standard 3DGS target.
Method
Flow.
House
Pavil.
Plant
Stat.
Avg.
3DGS
22.73
23.10
21.25
22.76
21.62
22.29
Ours (1-spp)
13.47
13.45
12.67
13.92
13.27
13.36
SS (4-spp)
17.80
17.86
16.86
18.23
17.39
17.63
GPS (4-spp)
17.83
17.87
16.86
18.22
17.51
17.66
Ours (1-spp + S)
21.74
22.35
21.01
21.60
20.94
21.53
Table 3: PSNR on held-out DL3DV views with respect to photographic ground truth. Avg. is the arithmetic mean over the five scenes. The best and second-best values are shown in bold and underlined, respectively. SS (4-spp), GPS (4-spp), and Ours (1-spp + S) run at 353.32, 282.91, and 353.43 FPS, respectively.
Method
Truck
Room
DrJohnson
Avg.
3DGS-GL
5.46
12.8
3.2
7.15
SS
24.8
65.8
69.4
53.3
GPS
41.3
52.8
53.0
49.03
SortfreeGS
37.2
50.3
27.2
38.2
MobileGS
–
32.6
28.1
30.4
Ours(1-spp)
78.8
112.7
86.3
92.6
Table 4: Mobile 540p throughput (FPS) on vivo x300pro with MediaTek tianji9500, using one scene per benchmark. Our renderer uses Vulkan, and Ours-S runs in FP16 on the NPU. Avg. is the arithmetic mean over available scene results. The best and second-best values are shown in bold and underlined, respectively.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Truck
Room
DrJohnson
Avg.
3DGS-GL
5.46
12.8
3.2
7.15
SS
24.8
65.8
69.4
53.3
GPS
41.3
52.8
53.0
49.03
SortfreeGS
32.94
22.03
31.95
28.97
MobileGS
–
26.98
25.61
26.29
Ours (1-spp)
78.8
112.7
86.3
92.6
Appendix
Table A1: Mobile 540p throughput (FPS) on vivo x300pro with MediaTek tianji9500, using one scene per benchmark. Our renderer uses Vulkan, and Ours-S runs in FP16 on the NPU. Avg. is the arithmetic mean over available scene results. The best and second-best values are shown in bold and underlined, respectively.
Figure A1: Crossover boundaries ( tprimitive=tfragment ) for varying layer counts L∈{1,…,8} plotted over the opacity-area plane. The almost perfectly overlapping curves indicate that depth complexity has a negligible effect on the relative performance of the two streams.
Scene
Small (1 spp)
Large (1 spp)
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Bicycle
24.0400
0.6710
0.2600
24.6107
0.6996
0.2745
Bonsai
30.7500
0.9110
0.1600
32.1253
0.9284
0.2303
Counter
28.2700
0.8690
0.1890
28.9801
0.8927
0.2317
Garden
25.4000
0.7470
0.1750
26.0137
0.7769
0.1891
Kitchen
29.2100
0.8750
0.1550
30.0836
0.8948
0.1703
Appendix
Table A2: Per-scene and per-dataset reconstruction quality of the Small and Large variants at 1 spp, measured against ground-truth images. Dataset averages are arithmetic means over their constituent scenes; the final row is the macro-average over all 11 scenes.
Scene
Small (4 spp)
Large (4 spp)
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Bicycle
24.5220
0.7018
0.2799
24.9213
0.7212
0.2614
Bonsai
31.2651
0.9232
0.2346
32.4673
0.9317
0.2279
Counter
28.5755
0.8876
0.2357
29.1298
0.8981
0.2263
Garden
26.1494
0.7916
0.1903
26.5915
0.8104
0.1662
Kitchen
29.9029
0.8948
0.1701
30.5050
0.9057
0.1595
Appendix
Table A3: Per-scene and per-dataset reconstruction quality of the Small and Large variants at 4 spp, measured against ground-truth images.
Scene
Small (16 spp)
Large (16 spp)
PSNR ↑
SSIM ↑
LPIPS ↓
PSNR ↑
SSIM ↑
LPIPS ↓
Bicycle
24.9336
0.7303
0.2550
25.1800
0.7470
0.2080
Bonsai
31.7702
0.9304
0.2228
32.8700
0.9350
0.1430
Counter
28.7965
0.8974
0.2180
29.2900
0.9000
0.1690
Garden
26.7682
0.8296
0.1487
27.0900
0.8450
0.1000
Kitchen
30.4227
0.9098
0.1461
30.7800
0.9150
0.1120
Appendix
Table A4: Per-scene and per-dataset reconstruction quality of the Small and Large variants at 16 spp, measured against ground-truth images. Dataset averages are arithmetic means over their constituent scenes; the final row is the macro-average over all 11 scenes.
Variant
Bonsai
Counter
Train
Avg.
Forward Nearest.
30.75
28.23
21.50
26.83
Forward Bilinear.
30.60
28.19
21.38
26.72
Backward Nearest.
30.59
28.21
21.40
26.73
Backward Bilinear.
30.68
28.25
21.40
26.78
Appendix
Table A5: Temporal transport under stochastic visibility. Backward gathering produces structured artifacts; bilinear transport spreads samples, whereas nearest-forward transport preserves discrete observations. The best and second-best PSNR (dB) values are shown in bold and underlined ; ties share the same rank.
Table A6: Scene names and corresponding dataset hash values in DL3DV.
Figure A2: After 1024spp rendering using GPS’s source code, the result was compared with that of Gaussian rendering, and there was a loss. The only modification to Gaussian Point Splatting is that we removed the upper limit on individual Gaussian sampling.
Figure A3: Rendering failure of Gaussian Point Splatting: its workload distribution loses much detail when aligning the 3DGS bounding box at conventional rendering resolution.
Configuration
Ours-S
Ours-L
Input / output channels
40/3
40/3
Down-/upsampling stages
3/3
4/4
Resolution pyramid
H to H/8
H to H/16
Encoder widths, fine to coarse
16,16,16,16
64,96,128,192,256
Encoder/bottleneck convolutions per resolution
2,1,1,2 (first stage: 1×1 , then 3×3 )
2,2,2,2,2 (all 3×3 )
Decoder transitions, coarse to fine
32→16→16 at H/4 ; 32→16→16 at H/2 ; 32→16→16 at H
384→192→192 at H/8 ; 288→128→128 at H/4 ; 192→96→96 at H/2 ; 128→64→64 at H
Appendix
Table A7: Detailed reconstruction-network architectures. A decoder transition a→b→c denotes the channel count after skip concatenation and the outputs of the two subsequent 3×3 convolutions. Both variants use concatenative encoder skips, a projected full-resolution input skip, max-pool downsampling, nearest-neighbor upsampling, and direct RGB prediction. Parameter counts include convolutional biases. FP16 kernel-weight sizes exclude biases, activations, and runtime workspace.
Figure A4: Temporal ablations in the evaluated stochastic-input setting. Backward reprojection exhibits structured artifacts, while forward reprojection leaves unfilled pixels.
Figure A5: Qualitative results of additional scenes of 16spp + Large network showing neural reconstruction quality