Standard 3D Gaussian Splatting (3DGS) learns geometry and appearance jointly from RGB supervision, making it difficult to isolate how luminance and chroma contribute to the learned representation. We study this by training models under different channel supervision, freezing their non-appearance parameters (position, scale, rotation, and opacity), and re-estimating appearance with the same solver before comparing held-out reconstruction. Across eleven benchmark scenes with four independent runs each, geometry learned from luminance alone supports held-out reconstruction 0.085 dB below RGB-trained geometry on average. If chroma is deleted from a trained model, a sufficiently expressive solver can re-fit it on the frozen geometry to the original quality or slightly better. Higher-order spherical harmonics contribute much more reconstruction quality to luminance than to chroma, improving PSNR by 1.44 dB versus 0.19 dB on average, although on mirror-like surfaces hue does still change with viewpoint. The luminance advantage is even larger when geometry is being formed. Chroma-only supervision produces geometry 3.9-5.5 dB worse than luminance-only supervision after the same appearance solve; densification explains part of this gap. Overall, geometry formation in standard 3DGS is strongly luminance-dominated but not luminance-exclusive, and much of the chromatic appearance can be recovered after spatial support has formed.
Figures & tables
Figure 1 : Separating geometry supervision from appearance fitting. ( a ) One geometry is trained with full RGB supervision and the other with luminance-only supervision; after freezing geometry, the same full-appearance solver A∗ is applied to both. ( b ) The Delayed-Chroma observation that motivated the study: chromatic supervision is withheld during an initial luminance phase and introduced later. This training behavior motivates the geometry test but is not itself evidence of geometry equivalence.
PSNR
Geometry
dB
vs. Q1
SSIM ↑
LPIPS ↓
Trained appearance
GRGB
28.112
–
0.8694
0.1469
Common appearance solve A∗
GRGB
28.555
0.000
0.8717
0.1469
GY
28.457
−0.098
0.8706
0.1458
Table 1: Reconstruction before and after the common appearance solve. Means are over eleven seed-0 scenes; this is the single-seed realization. The four-run results supporting the headline are in Table 2 . The same A∗ is applied to the last three rows. The basis control trains the YCbCr SH parameterization under the stock RGB loss, so that only the appearance basis differs from RGB training. SSIM and LPIPS are computed on the same held-out views and render path as PSNR.
Figure 2 : Luminance-trained geometry approaches RGB-trained geometry under the same appearance solve, and chroma can be recovered once geometry is fixed. (A) Each point compares held-out RGB reconstruction after applying the same full-appearance solve A∗ to RGB-trained and luminance-trained geometry. The dashed line marks exact parity, and the shaded region is the prespecified ±0.15 dB investigation band. Every point is the condition mean over four independent runs per condition, so the mean difference of the points is the eleven-scene mean; the labels give the condition-mean differences for the three departures. Eight of eleven scenes lie inside the band; kitchen , dr. johnson , and playroom retain RGB-geometry advantages. (B) Chroma recovery from fixed RGB-trained geometry improves as the appearance solve moves from a single scene-wide tint to independent per-Gaussian DC, coupled DC, and coupled degree-3 SH. The dashed line at RCbCr=1 marks the reconstruction quality of the appearance stored in the trained model.
Figure 3 : Matched appearance fitting exposes both the typical near-equivalence and a scene-specific departure. Each row compares ground truth with the identical canonical full-appearance solve A∗ applied to RGB-trained geometry GRGB and luminance-trained geometry GY . (a) On garden , the two geometries support nearly identical held-out reconstructions on both the weakest and strongest views shown: 20.39 versus 20.36dB , and 30.63 versus 30.61dB , respectively. (b) kitchen illustrates why benchmark-wide near-equivalence is not universal. Its weakest shown view favors RGB-trained geometry by 1.35dB ( 27.38 versus 26.03dB ), while the strongest shown view is effectively tied ( 35.77 versus 35.80dB ). All reconstructions use the same appearance-fitting procedure; only the frozen geometry differs.
Figure 4 : Higher-order SH contributes far more held-out reconstruction quality to luminance than to chroma. (A) Recovery as the SH degree cap increases across the eleven benchmark scenes, measured by re-solving at each cap. Chroma is already close to its full recovery at degree 0, while luminance gains substantially more from higher-order SH. (B) Held-out PSNR loss after removing SH degrees 1–3 from luminance or chroma on diffuse, synthetic-shiny, and real-shiny scenes, measured by the collapse test of Sec. 3 , which zeroes the trained coefficients without re-solving. Higher-order luminance is more important in every case, although shiny scenes also show measurable chromatic losses. (C) Two ground-truth views of the chrome ball separated by approximately 90∘ . Reflected colors move substantially across the surface with viewpoint, motivating the fixed-surface-point chromaticity analysis.
Figure 5 : Chromatic supervision does not substitute for luminance when learning reconstruction-supporting geometry. (a) Held-out reconstruction after the identical canonical full-appearance solve A∗ . Luminance-trained geometry remains close to RGB-trained geometry, whereas raw chroma-only geometry loses 3.9 – 5.5dB . The five bars per scene are, in legend order, A∗(GRGB) and A∗(GY) from the reruns of this section, raw chroma-only A∗(GCbCr) , chroma-only with the loss multiplied by s ( A∗(GCbCrs) ), and luminance-only with the loss multiplied by 1/s ( A∗(GY1/s) ), the reduced-luminance control of Sec. 3 ; on garden the two loss-scaled arms are not valid controls (supplement). (b) Reconstruction quality versus Gaussian count for the channel interventions and capacity controls, including the threshold-matched and densification-free chroma conditions. Lowering chroma’s densification threshold increases both primitive count and reconstruction quality but does not close the gap to luminance. On counter and kitchen , luminance-supervised models remain substantially better even when they contain fewer primitives. Dashed lines join the densification-free pairs, which keep the same set of primitives; chroma-only geometry still trails by 1.4 – 2.5dB . “Y-only, down-weighted” is the reduced-luminance control of Sec. 3 .
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Scene
A∗(GRGB)
A∗(GY)
Δgeom
SE
Status
bicycle
25.409
25.432
+0.023
0.008
in band
bonsai
33.133
33.023
−0.110
0.035
in band
counter
29.621
29.634
+0.013
0.014
in band
garden
27.660
27.600
−0.060
0.009
in band
kitchen
32.056
31.772
−0.284
0.103
departure
room
32.185
32.182
−0.003
0.059
in band
Appendix
Table 2: Four-run main comparison per scene. Condition means of held-out PSNR after A∗ over four independent training runs per condition. Δgeom is Y minus RGB; its standard error is sRGB2/4+sY2/4 from the per-run values. A scene is a departure when ∣Δgeom∣>0.15 dB; all three departures also exceed twice their standard error. The standard error in the mean row is the between-scene standard deviation divided by 11 .
SSIM ↑
LPIPS ↓
Scene
Q0
Q1
Q2
Q2−Q1
Q0
Q1
Q2
Q2−Q1
bicycle
0.7427
0.7429
0.7449
+0.0020
0.2094
0.2113
0.2079
−0.0034
bonsai
0.9427
0.9459
0.9447
−0.0011
0.1262
0.1225
0.1221
−0.0005
counter
0.9068
0.9130
0.9129
−0.0000
0.1516
0.1463
0.1444
−0.0020
garden
0.8554
0.8546
0.8536
−0.0009
0.0903
0.0946
0.0932
−0.0013
kitchen
0.9300
0.9329
0.9310
−0.0020
0.0885
0.0893
0.0893
−0.0001
Appendix
Table 3: SSIM and LPIPS per scene for the main comparison. Seed 0, with the same rows and evaluation path as Table 1 ; Q0 is evaluated through the common YCbCr render-to-RGB path. SSIM is higher-better and LPIPS lower-better, so a positive LPIPS Q2−Q1 favors RGB geometry.
Group
Cap
PSNR ↑
SSIM ↑
LPIPS ↓
Y
0
26.925
0.8392
0.1683
1
28.059
0.8615
0.1541
2
28.291
0.8661
0.1518
3
28.366
0.8673
0.1514
CbCr
0
28.049
0.8679
0.1484
1
28.206
0.8711
0.1445
Appendix
Table 4: Channel capacity in PSNR, SSIM, and LPIPS. Eleven-scene means of the full-appearance solve restricted to SH degree ≤ cap for one channel group; the other channel group keeps its trained appearance, and this sweep uses tolerance 5×10−4 and a 70-iteration cap, so its degree-3 rows are not the A∗ rows of Table 1 .
Scene
Q0
Q1
Q2
Q1−Q0
Q2−Q1
MSE( Q0 ) ×10−3
bicycle
25.070
25.411
25.448
+0.341
+0.037
3.836
bonsai
32.376
33.192
33.070
+0.817
−0.123
0.7725
counter
28.752
29.628
29.663
+0.876
+0.035
1.486
garden
27.261
27.649
27.609
+0.387
−0.039
2.254
kitchen
31.542
32.166
31.714
+0.624
−0.452
0.7777
room
31.673
32.156
32.089
+0.482
−0.067
0.8860
Appendix
Table 5: Per-scene Q0 , Q1 , and Q2 at seed 0. PSNR in dB on the benchmark test views. Q0 is the trained RGB appearance evaluated through the common YCbCr render-to-RGB path; Q1 and Q2 are A∗ on GRGB and GY . The last column is the pooled held-out MSE of the trained reference; pooled MSE was not recorded for the seed-0 A∗ solves.
RCbCr (PSNR domain)
PSNR den.
MSE ( ×10−3 )
Scene
tint
indep. DC
coupled DC
coupled d3
(dB)
ref.
removed
den.
bicycle
0.427
0.956
0.990
1.009
3.096
3.836
6.741
2.906
bonsai
0.007
0.858
0.971
1.022
7.884
0.7725
3.703
2.931
counter
0.231
0.942
1.007
1.036
9.860
1.486
13.28
11.79
garden
0.458
0.920
0.986
1.007
8.051
2.254
12.14
9.885
kitchen
0.137
0.814
0.946
1.010
12.993
0.7777
14.13
13.36
Appendix
Table 6 : Operator ladder per scene, with the denominators of both recovery ratios. RCbCr is the PSNR-domain ratio of Eq. ( 5 ) for the four solves of Fig. 2 B. The PSNR denominator is Q(Aref)−Q(ACbCr∅) . The MSE columns give the pooled held-out MSE of the trained reference and of the chroma-removed state, and their difference, which is the denominator of RMSE .
R at cap
RMSE at cap
PSNR den.
MSE den.
Scene
0
1
2
3
0
1
2
3
(dB)
( ×10−3 )
Luminance ( Y )
bicycle
0.928
1.009
1.024
1.027
0.987
1.003
1.005
1.006
11.635
42.71
bonsai
0.872
0.985
1.014
1.024
0.994
1.000
1.001
1.002
21.478
82.30
counter
0.882
0.988
1.015
1.023
0.989
0.999
1.001
1.002
17.554
74.97
garden
0.900
0.998
1.018
1.024
0.983
1.000
1.003
1.004
12.432
31.23
Appendix
Table 7 : Capacity sweep in both domains. Recovery when one channel group is re-solved through SH degree ≤ cap while the other keeps its trained appearance. R is the PSNR-domain ratio of Eq. ( 5 ) and RMSE its MSE-domain counterpart (Sec. 3 ), each with its denominator. Solver settings: ridge 10−6 , tolerance 5×10−4 , 70-iteration cap.
Figure 6 : Delayed-chroma training dynamics and boundary controls in stock 3DGS. (A) Training begins with luminance-only supervision and switches to full YCbCr supervision at 15k iterations. Chroma release produces a 1.84dB increase in held-out PSNR over the following 2k iterations, after which the delayed schedule converges near the RGB reference. (B) Moving chroma release moves the reconstruction jump with it: releases at 10k, 15k, and 20k produce approximately 1.77 , 1.84 , and 1.79dB steps, respectively, despite the fixed densification boundary at 15k. (C) Boundary controls separate chroma release from coincident optimization events. Chroma release produces 1.74 – 1.84dB changes across release times, whereas stopping densification produces only 0.10 – 0.13dB changes. Each bar spans 2,000 iterations from its boundary, except the 25k release, which spans 2,500. Releasing higher-order SH without changing chromatic supervision produces smaller and non-systematic effects. Black horizontal marks show the change in an RGB run from the same boundary; the RGB run was logged at 2,500-iteration spacing beyond 5k, so its marks span 2,500 iterations where the bars span 2,000. (D) The four delayed-release schedules finish within 0.074dB of one another. Including the standard RGB reference, the full five-condition spread is 0.122dB . All panels use the stock 3DGS implementation; the independent FastGS transfer is reported in Sec. A.6 .
After the success of 3D Gaussian Splatting (3DGS) for novel view synthesis, many works have explored how to also use it for geometric surface representation. However, extracting accurate geometric information directly from 3DGS remains challenging and can often reduce the appearance rendering quality. In this work, we show that 3DGS in its default form is inheritedly unsuited to represent texture and geometry at the same time, by training with complete ground-truth texture and geometry information. We also propose a simple solution by applying a single additional geometry opacity parameter to each splat, together with an optional transparency-curated optimization pipeline. Our experiments, both with ground-truth and vision foundation model geometric input, show that this change leads to improved rendering and geometry performance on a wide variety of dataset, and especially complex scenes with transparent objects benefit significantly from our method.
While feed-forward 3D Gaussian Splatting (3DGS) enables efficient 3D reconstruction, achieving high-fidelity rendering remains challenging. Existing pixel-aligned approaches suffer from spatial inflexibility and massive structural redundancy, whereas query-based methods lack 3D priors and entangle geometry with appearance, yielding blurry, pose-dependent results. To overcome these deficiencies, we propose \textbf{QuerySplat}, a feed-forward 3DGS framework driven by geometric priors and explicit appearance decoupling. Specifically, we design a dual-branch query-based decoder: the geometry branch leverages a pretrained Vision Geometric Model for spatial understanding, which intrinsically endows QuerySplat with pose-free modeling capabilities, while the appearance branch recovers high-frequency details through a dedicated pathway separated from geometric attribute regression. Extensive experiments demonstrate that QuerySplat mitigates the blurry rendering issues of early query-based models and consistently outperforms pixel-aligned approaches in rendering fidelity. On the challenging DL3DV benchmark, it achieves state-of-the-art novel view synthesis performance, with average PSNR gains of 2.30 dB and 1.04 dB over the best pose-free and pose-required baselines, respectively. Project Page: https://inspatio.github.io/querysplat.
Yinglong Li, Donghui Shen, Xiaoyu Zhang +5
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University · InSpatio Research · State Key Lab of CAD&CG, Zhejiang University
Standard 3D Gaussian Splatting (3DGS) assumes that every input image faithfully samples scene radiance. However, mixed-quality JPEG images violate this assumption because compression-induced blocking and ringing artifacts can corrupt updates to Gaussians shared across views. To address this problem, we propose JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views (JSGS). JSGS uses luminance and chrominance quantization tables stored in each JPEG file to construct a view-specific JPEG observation operator. This operator encodes and decodes each rendered view for domain-matched comparison with the corresponding decoded input image. The luminance quantization table supplies continuous weights within a fixed middle frequency band. A loss in the low frequency band anchors coarse structure, while the weighted middle frequency loss redistributes supervision among the selected DCT coordinates. The resulting block disagreement also guides the Gaussian Controller to regularize small primitives with high opacity in disagreement regions. Across seven scenes and three mixed-quality schedules, JSGS achieves the lowest mean LPIPS and the highest mean SSIM under every schedule while rendering at approximately 150 FPS. Code: https://github.com/Jayden-Cui/JSGS.
Jinhua Cui, Anhong Wang, Kai Hu +5
1Taiyuan University of Science and Technology · 2Chn Energy Digital Intelligence Technology Development (Beijing)