We revisit the role of appearance modeling in 3D Gaussian Splatting (3DGS) and show that limited expressiveness in view-dependent reflectance is a key driver of representation redundancy. In standard 3DGS, low-order spherical harmonics (SH) are used, restricting the splats' ability to model high-frequency directional effects, which is typically compensated by increasing the number of splats. We propose \emph{NRF-GS: Neural Residual Fields for Gaussian Splatting}, a hybrid representation that replaces per-splat SH-bases with a shared neural residual field. Each Gaussian encodes a compact set of appearance features and a lambertian base color, while a lightweight \emph{global scene-level MLP} predicts view-dependent residuals conditioned on viewing direction, distance, and per-splat features. This formulation enhances directional reflectance modeling by combining diffuse per-splat reflectance representations with a shared global function for high-frequency details, enabling both higher expressiveness and parameter sharing across splats. Our key insight is that by accurately capturing high-frequency directional reflectance, especially in specular regions, the GS-representation becomes more expressive, reducing the need for geometrically redundant splats. As a result, NRF-GS achieves comparable or better rendering quality while reducing the number of Gaussians by up to 50%, and produces visibly improved specular and high-frequency details.
Figures & tables
Figure 1 : Novel view from Mip-NeRF 360 dataset - Garden scene. Our method (d) achieves better specular representation on the foreground table as well as sharper details on the background wall with significantly lower number of points compared to other methods.
Figure 2 : NRF-GS architecture. Per Gaussian geometric attributes (xi,Σi) and appearance attributes ( cibase , αibase , and a latent feature zi ). The viewing direction di is frequency-encoded and decomposed into low ( γlow(di) ) and high-frequency ( γhigh(di) ) components. Two MLP branches ( flow , fhigh ) predict a low-frequency residual color and opacity residual (Δcilow,δiα) and a high-frequency color residual with a learned modulation weight (Δcihigh,wi) . The final color and opacity are then rendered using the unchanged standard Gaussian splatting rasterizer.
Table 3
Figure 3 : Simulation results for synthetic degree-5 SH (top row) and real BTF data Sattler et al. (2003) (bottom row). Each row shows ground truth, degree-3 SH and NRF reconstruction, and corresponding error maps.
Figure 4 : Point count growth for the Kitchen scene (MipNeRF360). NRF-GS converges to fewer Gaussians under identical densification/pruning.
Dataset
Method
Points (M) ↓
Time (min) ↓
MS-SSIM ↑
LPIPS ↓
L1 ↓
PSNR ↑
MipNeRF360
3DGS
3.359
26.46
0.813
0.184
0.030
27.40
VDGS
3.548
41.27
0.813
0.186
0.029
27.65
NRF-GS (ours)
1.656
31.38
0.814
0.191
0.029
27.86
MipNeRF360*
3DGS
2.605
24.22
0.812
0.199
0.030
27.74
VDGS
2.828
36.70
0.813
0.200
0.029
28.03
GSNB
2.667
59.06
0.800
0.201
0.031
27.71
Table 3 : Average test-set results on MipNeRF360, DL3DV, and Tanks and Temples. Points are reported in millions. MipNeRF360* excludes Bicycle and Garden , where GSNB goes out-of-memory. Best and second-best values in each dataset are highlighted.
Figure 5 : Qualitative comparison across multiple scenes. Top to Bottom: MipNeRF360 - scenes ’kitchen’ and ’room’, TnT - scenes ’palace’ and ’panther’, DL3DV - scenes ’387e’ and ’06da’.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Scene
Method
MS-SSIM ↑
LPIPS ↓
L1 ↓
PSNR ↑
Train
Test
Train
Test
Train
Test
Train
Test
Bicycle
Single
0.864
0.715
0.162
0.234
0.026
0.047
27.34
23.71
Ours
0.843
0.759
0.170
0.203
0.029
0.041
26.11
25.06
Bonsai
Single
0.950
0.942
0.138
0.144
0.012
0.013
34.54
33.23
Ours
0.950
0.944
0.136
0.137
0.012
0.014
34.10
33.24
Counter
Single
0.926
0.906
0.151
0.173
0.013
0.018
32.03
29.81
Appendix
Table 4 : Per-scene comparison of Single-branch vs Ours (NRF-GS) on MipNeRF360. Train/Test metrics are reported as sub-columns.
Dataset
Method
Points (M) ↓
Time (min) ↓
MS-SSIM ↑
LPIPS ↓
L1 ↓
PSNR ↑
MipNeRF360
3DGS
3.359 ± 1.986
26.46 ± 05.02
0.813 ± 0.130
0.184 ± 0.094
0.030 ± 0.015
27.40 ± 3.95
VDGS
3.548 ± 1.869
41.27 ± 09.95
0.813 ± 0.132
0.186 ± 0.096
0.030 ± 0.015
27.65 ± 4.09
NRF-GS (ours)
1.656 ± 1.015
31.38 ± 07.12
0.814 ± 0.131
0.191 ± 0.097
0.029 ± 0.016
27.86 ± 4.27
MipNeRF360*
3DGS
2.605 ± 1.503
24.22 ± 02.66
0.813 ± 0.147
0.200 ± 0.098
0.029 ± 0.017
27.74 ± 4.46
VDGS
2.828 ± 1.372
36.69 ± 04.58
0.814 ± 0.149
0.200 ± 0.101
0.029 ± 0.017
28.04 ± 4.59
GSNB
2.667 ± 1.315
59.06 ± 20.68
0.800 ± 0.162
0.201 ± 0.102
0.032 ± 0.020
27.71 ± 5.13
Appendix
Table 5 : Average test-set results (mean ± std) on MipNeRF360, DL3DV, and Tanks and Temples. MipNeRF360* excludes Bicycle and Garden scenes where GSNB runs out-of-memory. Points are reported in millions.
Figure 6 : Additional qualitative comparison across multiple scenes. Top to Bottom: MipNeRF360 - scenes ’bonsai’, TnT - scenes ’horse’ and ’lighthouse’, DL3DV - scenes ’5c3a’ and ’669c’.
Scene
Method
Points (M) ↓
Time (min) ↓
MS-SSIM ↑
LPIPS ↓
L1 ↓
PSNR ↑
Bicycle
3DGS
6.137 ± 0.076
34.08 ± 0.13
0.765 ± 0.001
0.175 ± 0.001
0.039 ± 0.001
25.17 ± 0.04
VDGS
6.441 ± 0.069
58.88 ± 0.29
0.755 ± 0.002
0.195 ± 0.002
0.039 ± 0.001
25.14 ± 0.14
NRF-GS
3.187 ± 0.024
44.12 ± 0.28
0.759 ± 0.003
0.203 ± 0.006
0.041 ± 0.001
25.06 ± 0.16
Bonsai
3DGS
1.257 ± 0.005
20.12 ± 0.10
0.943 ± 0.001
0.133 ± 0.001
0.015 ± 0.001
32.23 ± 0.10
VDGS
1.617 ± 0.012
29.70 ± 0.33
0.945 ± 0.002
0.134 ± 0.005
0.014 ± 0.000
32.72 ± 0.15
GSNB
1.477 ± 0.042
39.15 ± 0.18
0.945 ± 0.002
0.127 ± 0.003
0.014 ± 0.001
33.19 ± 0.05
Appendix
Table 6 : Per-scene test-set results on MipNeRF360.
Scene
Method
Points (M) ↓
Time (min) ↓
MS-SSIM ↑
LPIPS ↓
L1 ↓
PSNR ↑
06da
3DGS
1.330 ± 0.007
12.85 ± 0.09
0.905 ± 0.001
0.107 ± 0.001
0.024 ± 0.001
28.21 ± 0.01
VDGS
1.687 ± 0.012
22.30 ± 0.05
0.906 ± 0.002
0.110 ± 0.002
0.023 ± 0.001
28.41 ± 0.13
GSNB
1.404 ± 0.007
31.57 ± 0.50
0.908 ± 0.002
0.102 ± 0.003
0.021 ± 0.001
29.34 ± 0.01
NRF-GS
0.703 ± 0.009
15.13 ± 0.13
0.928 ± 0.000
0.098 ± 0.001
0.017 ± 0.001
30.91 ± 0.03
093e
3DGS
0.897 ± 0.011
11.05 ± 0.13
0.952 ± 0.001
0.073 ± 0.001
0.016 ± 0.000
31.21 ± 0.02
VDGS
0.957 ± 0.004
16.28 ± 0.10
0.955 ± 0.002
0.071 ± 0.001
0.016 ± 0.000
31.46 ± 0.15
Appendix
Table 7 : Per-scene test-set results on DL3DV.
Scene
Method
Points (M) ↓
Time (min) ↓
MS-SSIM ↑
LPIPS ↓
L1 ↓
PSNR ↑
Auditorium
3DGS
0.735 ± 0.005
10.92 ± 0.08
0.871 ± 0.002
0.205 ± 0.005
0.044 ± 0.001
24.28 ± 0.03
VDGS
0.856 ± 0.002
17.42 ± 0.08
0.858 ± 0.003
0.209 ± 0.001
0.047 ± 0.001
24.10 ± 0.06
GSNB
0.765 ± 0.004
21.17 ± 0.06
0.867 ± 0.002
0.199 ± 0.001
0.046 ± 0.001
24.25 ± 0.04
NRF-GS
0.375 ± 0.003
12.17 ± 0.06
0.886 ± 0.002
0.193 ± 0.001
0.038 ± 0.001
25.69 ± 0.11
Ballroom
3DGS
3.077 ± 0.068
22.48 ± 0.08
0.824 ± 0.001
0.102 ± 0.003
0.038 ± 0.000
24.15 ± 0.06
VDGS
3.802 ± 0.007
43.83 ± 0.13
0.830 ± 0.002
0.105 ± 0.001
0.035 ± 0.001
24.83 ± 0.58
Appendix
Table 8 : Per-scene test-set results on Tanks and Temples.
Scene
Method
Points (M) ↓
Time (min) ↓
MS-SSIM ↑
LPIPS ↓
L1 ↓
PSNR ↑
Lighthouse
3DGS
0.870 ± 0.025
12.40 ± 0.22
0.843 ± 0.002
0.161 ± 0.001
0.052 ± 0.001
22.31 ± 0.02
VDGS
1.201 ± 0.031
22.37 ± 0.10
0.841 ± 0.002
0.162 ± 0.001
0.051 ± 0.001
22.54 ± 0.01
GSNB
0.819 ± 0.006
21.45 ± 0.13
0.846 ± 0.004
0.165 ± 0.002
0.052 ± 0.001
22.57 ± 0.01
NRF-GS
0.522 ± 0.008
15.32 ± 0.18
0.854 ± 0.001
0.158 ± 0.001
0.045 ± 0.001
23.52 ± 0.01
M60
3DGS
1.652 ± 0.006
14.73 ± 0.06
0.903 ± 0.001
0.120 ± 0.001
0.023 ± 0.001
27.82 ± 0.05
VDGS
2.349 ± 0.003
27.78 ± 0.08
0.904 ± 0.001
0.119 ± 0.001
0.022 ± 0.001
28.05 ± 0.05
Appendix
Table 9 : Per-scene test-set results on Tanks and Temples, continued.
Three-dimensional Gaussian Splatting (3DGS) combines explicit primitives with efficient rasterization, yet recent systems increasingly use neural networks to generate or share Gaussian parameters. We characterize this trend along five axes: attribute decoding, spatial sharing, view-conditioned decoding, topology generation, and amortized inference. An analysis of 19 representative methods shows that these choices address different limitations and cannot be reduced to a binary neural label. We also isolate three forms of neural parameterization in a controlled mip-NeRF 360 study. Sharing appearance and opacity improves reconstruction quality, while decoding geometric structure offers no further gain. The evidence favors selective neuralization: shared functions help when they capture reusable correlations without sacrificing the local geometric freedom of explicit splats.
Gaussian Splatting has significantly improved the quality of novel view synthesis with explicit Gaussian representation. However, we observed that existing 3D Gaussian Splatting methods (3DGS) often suffer from surface collapse issues on reflective regions, and thus produce inferior geometry and low-quality specular. In this work, we propose a physically-based deferred rendering framework, named Reflection-aware Gaussian Splatting (RGS), that can accurately model specular regions and improve novel view synthesis performance. Specifically, we found that a powerful 3D foundation model can provide a strong 3D geometric prior to foster correct geometric modeling. Based on this, we propose a cross-view shape consistency regularization to regularize the geometry surface with the large model prior and cross-view constraints. In this manner, our RGS can produce smoother geometric surfaces on reflective regions while reducing geometric hollows. To further improve rendering results on reflective regions, we present a reflection-aware densification strategy that is designed to capture specular variations across various views. With this strategy, our RGS is able to render novel views of objects in higher quality. Extensive experiments demonstrate our method consistently renders high-quality reflective objects, achieving state-of-the-art performance.
Xiaobiao Du, Yida Wang, Cheng Bi +2
University of Technology Sydney · Li Auto Inc. · Adelaide University
While 3D Gaussian Splatting (3DGS) has demonstrated impressive real-time rendering performance, its efficacy remains constrained by a reliance on heuristic density control. Despite numerous refinements to these handcrafted rules, such methods inherently lack the flexibility to adapt to diverse scenes with complex geometries. In this paper, we propose a paradigm shift for density control from rigid heuristics to fully learnable policies. Specifically, we introduce \textbf{LeGS}, a framework that reformulates density control as a parameterized policy network optimized via Reinforcement Learning (RL). Central to our approach is the tailored effective reward function grounded in sensitivity analysis, which precisely quantifies the marginal contribution of individual Gaussians to reconstruction quality. To maintain computational tractability, we derive a closed-form solution that reduces the complexity of reward calculation from O(N2) to O(N). Extensive experiments on the Mip-NeRF 360, Tanks & Temples, and Deep Blending datasets demonstrate that \textbf{LeGS} significantly outperforms state-of-the-art methods, striking a superior balance between reconstruction quality and efficiency. The code will be released at https://github.com/AaronNZH/LeGS
Zhenhua Ning, Xin Li, Jun Yu +3
Pengcheng Laboratory, Shenzhen · Harbin Institute of Technology, Shenzhen