Stochastic order-independent transparency enables efficient and elegant rendering of primitive-based radiance fields like 3D Gaussian Splatting models, but remains impractical due to the inherent visible noise in the output. We propose a principled approach to minimize high-frequency noise, addressing its sources at the representation and image synthesis level. During stochastic rendering, our history-based spatial resampling scheme drastically accelerates image convergence, while temporal importance resampling ensures coherence under camera movement. During training, a color regularizer implicitly reduces the variance along view rays in the 3DGS models. With these properties, our optimized, Vulkan-based renderer effectively mitigates output noise at low and high sample counts, achieving a substantial 13~dB PSNR increase in quality over previous stochastic methods at 1 sample per pixel and quickly converging to sorted 3DGS with an average L1 error of less than 10−4.
Figures & tables
Figure 1 : Through optimized training and inference, our stochastic OIT renderer achieves significantly better image quality than prior art with the same number of random samples. Our approach makes single-sample rendering a viable candidate for applications that demand high quality, as well as extremely high frame rates. Measurements were performed using each method’s original implementation.
Figure 2 : Illustration of deterministic and stochastic rendering of 3D Gaussians.
Figure 3 : By default, 3DGS optimization blends disparate colors to match the ground truth. With λvar>0 , the blended colors align to reduce color variance.
Figure 4 : Overview of spatial and temporal sample reuse.
Figure 5 : Average quality metrics on the Mip-NeRF 360 dataset for varying λvar with deterministic ( Fig. 5(a) ) and stochastic ( Fig. 5(b) ) blending.
PSNR (dB) ↑
Method
Avg. PSNR (dB) ↑
SSIM ↑
LPIPS ↓
tPSNR (dB) ↑
Indoor
Outdoor
StochasticSplats [ 13 ]
18.336
0.293
0.667
15.235
19.562
16.701
StochasticSplats + Temporal [ 13 ]
21.954
0.552
0.524
19.450
23.313
20.141
Ours
31.437
0.845
0.345
28.410
32.978
29.381
Ours + Temporal
35.085
0.937
0.209
32.449
37.161
32.318
Table 1: Quantitative comparison on the camera path benchmark (21 trajectories, 3,152 frames across 7 Mip-NeRF 360 scenes). We ablate progressive pipeline components against the 1-sample StochasticSplats baseline ( λvar=0 ), temporal reuse (CMA, τ=0.20 ), our variance loss ( λvar=1 ) with spatial reuse, and our full spatio-temporal model ( Ours + Temporal , δˉ=0.7 ). All variants are evaluated against sorted α -blending ground truth.
Figure 6 : Average residual of test quality metrics on the Mip-NeRF 360 dataset trained with λvar=1 , comparing static blur kernels against spatial reuse weights with varying K . Log-scale is used to visualize small differences.
Figure 7 : Average frame time for rendering all test views of the Mip-NeRF 360 dataset. We average the frame time for 200 frames per view.
Figure 8 : Excess Quality Factor on the test views of the Mip-NeRF 360 dataset for PSNR and SSIM quality metrics.
Figure 9 : Average residual of test PSNR on the Mip-NeRF 360 dataset trained with λvar=1 rendered with and without STBN compared against sorted rendering. Log-scale is used to visualize small differences.
Figure 10 : Visual Ablations.
Lvar
Spatial Reuse
Temporal Reuse
tPSNR (dB) ↑
SSIM ↑
LPIPS ↓
Avg. PSNR (dB) ↑
—
—
—
15.235
0.293
0.667
18.336
—
✓
—
21.720
0.586
0.498
24.774
—
—
✓
22.204
0.619
0.522
25.227
—
✓
✓
27.188
0.817
0.365
29.946
✓
—
—
22.423
0.613
0.525
25.561
✓
✓
—
28.410
0.845
0.345
31.437
Table 2: Component ablation on camera path benchmark. We ablate Lvar , Spatial Reuse, and Temporal Reuse combinations across 21 camera trajectories on the Mip-NeRF 360 dataset (3 per scene) against sorted rendering. For Spatial Reuse we set K=4 , and Temporal Reuse evaluates EMA filtering with our soft history rejection at fixed δˉ=0.70 . STBN was used for all evaluations.
Figure 11 : Log-scale histogram of primitive counts with respect to activated opacity over all Mip-NeRF 360 scenes trained with various values of λvar .
Figure 12 : Histogram of primitive counts with respect to Gaussian size (log-scale) over all Mip-NeRF 360 scenes trained with various values of λvar .
Figure 13 : Test view of the bicycle scene, rendered deterministically with models trained with varying λvar .
Figure 14 : Increasing λvar also reduces variance in stochastically rendered depth. Images were rendered at 1 sample per pixel.
Figure 15 : Distribution of primitive base colors in RGB space for scenes of Mip-NeRF 360 trained with and without variance loss.
Figure 16 : Effect of Lvar on alpha-blending.
Figure 17 : Comparison of quality metrics and variance for trained models with varying λvar on scenes from the Mip-NeRF 360 dataset.
Figure 18 : Average residual of L1 score and values for LPIPS on the Mip-NeRF 360 dataset trained with λvar=1 , comparing static blur kernels against spatial reuse weights with varying K . Log-scale is used for L1 to visualize small differences.
Figure 19 : Excess Quality Factor on the test views of the Mip-NeRF 360 dataset for the L1 quality metric.
Figure 20 : Heatmaps of quality metrics over sample counts and buffer sizes on Mip-NeRF 360.
Figure 21 : Pareto-optimal combinations of sample count and K for L1, PSNR, SSIM, and LPIPS. Stoch denotes stochastic rendering without spatial reuse and blur denotes stochastic rendering with a Gaussian cross-blur kernel.
Figure 22 : Average residual of test quality metrics on the Mip-NeRF 360 dataset trained with λvar=1 , comparing spatial reuse of 4- and 8-neighborhood kernels with varying K . Log-scale is used to visualize small differences.
Figure 23 : Average quality comparison of StochasticSplats and Gaussian Point Splatting with and without spatial reuse on the Mip-NeRF 360 dataset. Models were trained with λvar=1 .
Figure 24 : Comparison of uniform random noise sampling (PCG32) against per-primitive spatio-temporal blue noise sampling.