A surface photographed under even light presents nearly the same appearance from every angle; the same surface under uneven light does not. Exposure changes between views, illumination varies within a single image, and locally strong light sources leave one region bright and its neighbor in shadow. Multi-view reconstruction methods such as 3D Gaussian Splatting treat these lighting artifacts as if they were properties of the scene, entangling capture-specific illumination with the geometry and color they recover. We present EvenSplat, a framework that separates the two. EvenSplat couples an image-space illumination decomposition with an illumination field carried by the Gaussians, so that the same explanation of the lighting is shared between the two-dimensional and three-dimensional views of the scene; a camera-response network and a local exposure-compensation module absorb the global and residual differences that remain across training images. Through extensive experiments across multiple datasets and diverse forms of uneven illumination (cross-view exposure, spatial illumination variation, and high-contrast lighting) on both real-world captured and simulated benchmarks, EvenSplat generally outperforms state-of-the-art methods, particularly under high-contrast illumination.
Figures & tables
Figure 1: EvenSplat reconstructs base appearance from multi-view images with cross-view exposure variation (CEV), spatial illumination variation (SIV), and high-contrast illumination (HCI). Our method generally outperforms the state-of-the-art methods in PSNR, SSIM, and LPIPS on simulated and real-world 3D Gaussian Splatting novel view synthesis.
Figure 2: EvenSplat couples image-space decomposition with Gaussian-level illumination. Illumination alignment and image recombination connect the branches. CRN and ILEC absorb global and spatial training-image residuals. Base appearance is distinguished from the observation-fitting path.
Figure 3: Image-space regularization, from left to right: adaptive curve constraints on base-appearance intensity; edge-aware illumination smoothness; and white preservation using a bright-achromatic soft mask.
Figure 4: Gallery of the three illumination settings in the real-world and simulated datasets. CEV changes exposure across views, SIV introduces spatial illumination variation within images, and HCI produces a strong bright–dark imbalance.
Figure 5: Real-world comparisons. Columns show 3DGS, 3DGS+CHROMA, GS-W, Bilateral Grid, PPISP, Luminance-GS, EvenSplat, and the captured reference. Rows 1–2 show HCI, rows 3–4 SIV, and rows 5–6 CEV. EvenSplat most clearly separates appearance from illumination in the spatially uneven HCI and SIV.
Figure 6: Simulated HCI, SIV, and CEV comparisons. EvenSplat reduces illumination leakage while retaining scene texture; HCI remains the most challenging setting because contrast amplification clips information.
Figure 7: Cross-lighting appearance consistency. The capture sets differ only in left- versus right-dominant illumination. Paired appearance renderings and color-chart crops show the resulting reconstructions. PSNR is computed between corresponding left- and right-dataset appearance renderings, using one as the reference for the other; higher values indicate greater consistency across lighting conditions. PSNR (dB): 3DGS 10.5628; GS-W 10.6427; Bilateral Grid 10.3909; PPISP 10.3641; Luminance-GS 12.7311 (second-best); EvenSplat 15.5723 (best).
Figure 8: Ceiling digitization under uneven illumination ( Edwards et al., 2025 ) . Each panel pairs a held-out photograph (left) with an EvenSplat rendered novel view (right).
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Architecture of the image-space network for base-appearance and illumination decomposition.
Figure 10: Additional real-world qualitative comparisons. Columns show 3DGS, 3DGS+CHROMA, GS-W, Bilateral Grid, PPISP, Luminance-GS, EvenSplat, and the reference. The first two rows show PlasticCart (SIV), the next two show Rocks (SIV), and the final two show the extreme FourLogs case (CEV). For FourLogs , the last row applies a per-image, per-channel affine correction, yc=acxc+bc , to the preceding outputs. This diagnostic removes global color mismatch and exposes the remaining spatial reconstruction artifacts.
Figure 11: Qualitative comparison on simulated scenes under Cross-View Exposure Variation (CEV). Columns show 3DGS, 3DGS+CHROMA, GS-W, Bilateral Grid, PPISP, Luminance-GS, EvenSplat, and the reference; each row is a held-out view from a different scene. CEV changes the global exposure between views, testing whether each method can recover a consistent scene appearance without retaining view-dependent brightness and color shifts.
Figure 12: Qualitative comparison on simulated scenes under Spatial Illumination Variation (SIV). Columns show 3DGS, 3DGS+CHROMA, GS-W, Bilateral Grid, PPISP, Luminance-GS, EvenSplat, and the reference; each row is a held-out view from a different scene. The illumination changes within each image, producing adjacent bright and dark regions that cannot be removed by a single global exposure correction.
Figure 13: Qualitative comparison on simulated scenes under High-Contrast Illumination (HCI). Columns show 3DGS, 3DGS+CHROMA, GS-W, Bilateral Grid, PPISP, Luminance-GS, EvenSplat, and the reference; each row is a held-out view from a different scene. Strong highlights and deep shadows create clipping and severe local imbalance, emphasizing whether a method can restore shadow detail without flattening or overexposing bright regions.
Figure 14: Qualitative comparison of Rendered Observation on the four HDR-NeRF scenes. Columns show 3DGS, GS-W, Bilateral Grid, PPISP, Luminance-GS, EvenSplat, and the reference. The reference retains the severe but view-consistent illumination used to generate the observations; this comparison therefore measures complete-scene reconstruction fidelity rather than illumination removal.
Figure 15: Qualitative comparison of Base Appearance recovery on the four HDR-NeRF scenes. Columns show 3DGS, GS-W, Bilateral Grid, PPISP, Luminance-GS, EvenSplat, and the reference rendered under parallel uniform light. EvenSplat directly renders its decomposed base appearance, Luminance-GS produces enhanced renderings, whereas the other comparison methods contribute their standard scene renderings because they do not expose a separate base-appearance output.
Novel-view synthesis and 3D reconstruction from sparse posed images are central to robotics and AR/VR. Yet, feed-forward 3D Gaussian reconstruction fails under lowlight due to noise, color shifts, and unreliable correspondence. We propose DelowlightSplat, a lowlight-aware feed-forward Gaussian splatting framework for clean novel-view rendering. We build a controllable multi-view lowlight benchmark by degrading only context views while keeping target views clean. We introduce a lightweight Lowlight Adapter for residual enhancement to improve matchability, and couple it with cost-volume-based multi-view inference to directly predict clean 3D Gaussians. Experiments show that DelowlightSplat significantly outperforms previous feed-forward method and two-stage pipeline under lowlight conditions.
Fuzhen Jiang, Zengtian Xie, Zhuoran Li
Hangzhou Dianzi University · Hangzhou, China · Zhuhai College of Science and Technology +1
We present StructSplat, a feed-forward and generalizable 3D Gaussian reconstruction framework that operates directly on uncalibrated images without requiring camera parameters. Existing methods either rely on per-scene optimization or assume known camera poses, and often entangle geometry and appearance within a unified backbone, limiting reconstruction fidelity and generalization. Our key idea is to adopt a structured representation that organizes geometry, semantic, and texture cues with explicit roles in the reconstruction process. Specifically, we introduce a pixel-aligned feature injection mechanism to enable accurate texture modeling from 2D observations, incorporate semantic-aware priors to improve global consistency, and design a camera alignment strategy to prevent information leakage and improve generalization. Experiments show that our method significantly outperforms prior approaches on challenging benchmarks. On DL3DV, our method achieves 28.045 PSNR, surpassing AnySplat (22.377) by +5.67 dB. In cross-dataset evaluation, our method achieves +1.94 dB over AnySplat on ACID and +1.72 dB on RealEstate10K. Project page: https://structsplat.github.io Code: https://github.com/J-C-Zhao/StructSplat
Jia-Chen Zhao, Beiqi Chen, Xinyang Chen +2
Harbin Institute of Technology (Shenzhen) · Great Bay University · Guangzhou CloudButterfly Technology Co., Ltd.
In this paper, we propose UniqueSplat, a view-conditioned feed-forward 3D Gaussian Splatting model to reconstruct customized 3D radiance fields for each view query. Existing feed-forward methods such as pixelSplat and MVSplat aim to generate fixed Gaussians across all views of each scene by minimizing the error between rendered views and ground-truth images. However, such fixed Gaussians generally render images from all views and lack the ability to adapt to specific viewpoints, as they do not incorporate target view information when predicting Gaussians. To address this, our UniqueSplat learns the view-conditioned information as a prior and incorporates this knowledge into network parameters, so that Gaussians are dynamically adjusted in accordance with different views. Specifically, we propose a two-branch view-conditioned hyperNetwork to simultaneously learn view-agnostic embeddings and view-specific knowledge, which not only explores the shareable knowledge from various views, but also adapts the model to specific views at test time. Extensive experiments on widely-used datasets including RealEstate10K, ACID and DTU demonstrate the superiority of UniqueSplat over the state-of-the-art methods. Moreover, UniqueSplat encouragingly outperforms existing methods in cross-dataset evaluation, showing its notable generalization ability.