Illuminant estimation is a fundamental problem in computational photography, as it enables the correction of color shifts induced by varying lighting conditions. While learning-based methods have demonstrated strong performance, their progress is hindered by the limited availability of large-scale datasets with accurate illuminant ground-truth. In this work, we propose a general and reusable pipeline to derive dense illuminant chromaticity maps from physically based 3D-rendered scenes. By repurposing an existing 3D scene collection, our approach enables the systematic generation of pixel-wise illuminant annotations under controlled lighting conditions, effectively lowering the barrier to data acquisition for learning-based illuminant estimation. Using this pipeline, we generate a large-scale synthetic set of 74,321 images, which we employ for pre-training both single- and multi-illuminant estimation models. Extensive experiments with state-of-the-art architectures show that synthetic pre-training consistently improves performance, with gains of up to 28% for single-illuminant estimation and up to 57% for multi-illuminant estimation, particularly in data-scarce regimes. These findings demonstrate that synthetic data generation pipelines offer an effective and scalable solution for the pre-training of illuminant estimation methods.
Figures & tables
Figure 1: Sample images from the dataset. (a) illustrates the individual image components. (b) presents the input images along with their corresponding ground-truth illuminant images.
Hypersim-WB
LSMI-Galaxy [ 13 ]
MIMO [ 2 ]
Method
Pre-training Data
Mean
B-25
W-25
95-P
Mean
B-25
W-25
95-P
Mean
B-25
W-25
95-P
GW [ 22 ]
-
5.45
1.88
11.10
13.2
4.18
0.95
8.47
10.30
2.33
0.89
4.20
4.83
UNet [ 13 ]
-
3.18
1.1
6.35
8.00
1.94
0.65
3.88
4.63
3.30
2.06
5.48
6.08
Hypersim-WB
-
-
-
-
1.83
0.61
3.65
4.06
2.13
0.83
4.06
4.71
HDRNet [ 9 ]
-
3.77
1.45
7.57
9.20
2.36
0.67
4.98
5.75
3.41
2.31
4.87
5.34
Hypersim-WB
-
-
-
-
2.19
0.59
4.65
5.68
3.16
1.71
5.19
5.73
Table 1: Quantitative results, in terms of angular error, of the considered methods on three multi-illuminant estimation datasets: the proposed Hypersim-WB, LSMI-Galaxy [ 13 ] , and MIMO [ 2 ] . For each learning-based method we compare a version trained from-scratch and one that is fine-tuned after pre-training on the Hypersim-WB dataset. We highlight in bold the per-method best and underline the per-dataset best.
Figure 2: Qualitative results of the methods on the considered datasets: Hypersim-WB, LSMI-Galaxy [ 13 ] , MIMO [ 9 ] , and ColorChecker [ 11 ] . The learning based methods from columns (3-5) are pretrained on Hypersim-WB and fine-tuned on the target dataset. The mean angular error for each image is also reported.
Figure 3: Mean angular error for each considered training set size. The upper and lower bound represent respectively the mean of the worst quartile (W-25) and best quartile (B-25) results. The dotted lines correspond to the mean 95th percentile.
ColorChecker [ 11 ]
Method
Pre-training Data
Mean
B-25
W-25
95-P
GW [ 22 ]
-
4.80
1.72
7.11
13.20
UNet [ 13 ]
-
3.14
0.69
6.72
8.84
Hypersim-WB
2.94
0.63
6.76
7.55
HDRNet [ 9 ]
-
3.53
0.90
7.20
9.25
Hypersim-WB
3.40
0.82
7.46
9.50
Table 2: Quantitative results, in terms of angular error, of the considered methods for the task of single-illuminant estimation on the ColorChecker dataset [ 11 ] . For each learning-based method we compare a version trained from-scratch and one that is fine-tuned after pre-training on the Hypersim-WB dataset. We highlight in bold the per-method best and underline the per-dataset best.
We present PIXLRelight, a feed-forward approach for physically controllable single-image relighting. Existing methods either provide limited lighting control (e.g. through text or environment maps), accumulate errors when chaining inverse and forward rendering, or require costly per-image optimization. Our key idea is to bridge physically based rendering (PBR) and learned image synthesis through a shared intrinsic conditioning that can be obtained from either real photographs or PBR renders. At training time, paired multi-illumination photographs are decomposed into albedo, diffuse shading, and non-diffuse residuals, which condition the model. At inference time, the same conditioning is computed from a path-traced render of a coarse 3D reconstruction of the input under user-specified PBR lights. A transformer-based neural renderer then applies the target illumination to the source photograph, preserving fine image detail through a per-pixel affine modulation. PIXLRelight enables arbitrary PBR-style lighting control, achieves state-of-the-art relighting quality, and runs in under a tenth of a second per image. Code and models are available at https://mlfarinha.github.io/pixl-relight/.
Miguel Farinha, Ronald Clark
Department of Computer Science University of Oxford
Accurate modelling of illumination is central to realistic image synthesis and scene understanding. Yet, there is little exploration into whether image generative models are good at this task or whether physical plausibility remains a key challenge for them. Clearly, significant progress has been made in realistic image synthesis, but do models truly understand lighting in a physically accurate manner? To answer this question, this work proposes a benchmark to assess the lighting understanding and harmonisation capabilities of generative models. Our key insight is that evaluating lighting understanding for such models only requires testing how well they insert novel objects into real photographs whilst maintaining consistent illumination. To do so, we use a multi-illumination dataset with images containing simple objects serving as ``light probes'', and prompt models to inpaint the same object onto the original image, then compare the generated results against the ground-truth light probes. We then estimate the lighting direction, colour and radiance distribution from the inpainted probes, providing a quantitative measure of illumination accuracy and photometric realism. Our work establishes a scalable evaluation protocol to systematically assess how well generative models capture and reproduce real-world lighting, offering a foundation for benchmarking the photometric accuracy of any future models. All code and data are available at https://lvsn.github.io/SheddingLight/ .
Justine Giroux, Jack Oliver Hilliard, Yannick Hold-Geoffroy +2
Université Laval, Canada · Adobe Research, USA · Computer Vision Center & Universitat Autònoma de Barcelona, Spain
While synthetic data generation resolves the manual labeling bottleneck in computer vision, minimizing the syn-to-real domain gap requires optimizing rendering variables. This paper presents a systematic study analyzing the impact of lighting configurations and background complexity on object detection performance. We introduce SmartSDG, an automated, reproducible pipeline built on NVIDIA Isaac Sim using Physically-Based Shading (PBS), alongside ILLUM_INTRUCK, a new multi-object industrial benchmark dataset. Through 18 controlled experiments utilizing a state-of-the-art YOLOv12 framework, we demonstrate that complex, indirect lighting configurations paired with domain-relevant background variability significantly increase visual cue richness. Our quantitative findings show that avoiding direct specular peaks preserves crucial surface textures, mitigates the domain gap, reduces false positives, and accelerates model convergence compared to using conventional direct-light synthetic data. Ultimately, we provide actionable virtual scene design guidelines to maximize object detection robustness in industrial automation.
Hooman Tavakoli Ghinani, Tatjana Legler, Martin Ruskowski
German Research Center for Artificial Intelligence (DFKI), Kaiserslautern, Germany · RPTU Kaiserslautern-Landau, Kaiserslautern, Germany