Authors: David Serrano-Lozano, Anand Bhattad, Luis Herranz, Jean-François Lalonde, Javier Vazquez-Corral
Organizations: Computer Vision Center Universitat Autònoma de Barcelona · Johns Hopkins University · Universidad Politécnica de Madrid · Université Laval
We present SyncLight, a method to enable consistent, parametric control over light sources across multiple uncalibrated views of a static scene conditioned on a single view. While single-view relighting has advanced significantly, existing generative approaches struggle to maintain the rigorous lighting consistency essential for multi-camera broadcasts, stereoscopic cinema, and virtual production. SyncLight addresses this by enabling precise control over light intensity and color across a multi-view capture of a scene, conditioned on a single reference edit. Our method leverages a multi-view diffusion transformer trained using a latent bridge matching formulation, achieving high-fidelity relighting of the entire image set in a single inference step. To facilitate training, we introduce a large-scale hybrid dataset comprising diverse synthetic environments -- curated from existing sources and newly designed scenes -- alongside high-fidelity, real-world multi-view captures under calibrated illumination. Though trained only on image pairs, SyncLight generalizes zero-shot to an arbitrary number of viewpoints, effectively propagating lighting changes across all views, without requiring camera pose information. SyncLight enables practical relighting workflows for multi-view capture systems.
Figures & tables
Figure 1: Multi-view light editing with SyncLight. Given an uncalibrated multi-view capture of a static scene (1), the user picks a reference view , clicks one or more visible light sources (circle markers), and sets their intensity and chromaticity ; the other views never need to be inspected . SyncLight (2) relights all input views, reference included (3), in a single feedforward pass via latent bridge matching. Per-view light control baselines (bottom-left) hallucinate, lack precise control, and are inconsistent across views. SyncLight is geometrically consistent, needs no camera poses, generalizes to any number of views, and is 10 × faster than current models.
Figure 2: SyncLight formulates multi-view relighting as a conditional flow matching problem in latent space. (Left) Input scenes under source ( xsrc ) and target lighting ground truth ( xtar ) are encoded into latents ( zsrc,ztar ) via a VAE encoder. This is done for both the “Reference” (0) and “Other” (1) views. (Middle) During training, we sample a timestep t and construct a Bridge Matching (BM) interpolant zt . Our backbone, “Multi-View SD”, is conditioned on a user-specified “Lightmap” (encoding color as intensity and chromaticity) derived from the “Reference view”. To enforce consistency, the backbone processes both views simultaneously. (Right) The model predicts the target latents, which are decoded into relit images ( x^T ). The network is optimized using a hybrid objective L combining latent flow matching loss ( Llbm ) with pixel-level reconstruction losses ( Lpix ) for each view to ensure high-fidelity, consistent relighting.
Infinigen
BlenderKit
Real captures
Method
PSNR
SSIM
ΔE00
LPIPS
PSNR
SSIM
ΔE00
LPIPS
PSNR
SSIM
ΔE00
LPIPS
VGGTm
Time (s)
Ref. view
ScribbleLight
10.89
.518
30.67
.382
11.02
.528
29.26
.372
11.78
.483
28.36
.375
-
58.2 ± 1.14
Flux.2-dev
14.23
.672
20.31
.338
16.71
.704
17.22
.291
18.32
.757
14.28
.264
-
55.6 ± 1.02
LightLab*
29.86
.941
3.08
.131
26.38
.907
4.92
.147
28.42
.880
4.61
.199
-
1.17 ± 0.02
GR3EN
25.92
.893
4.38
.158
21.74
.803
8.93
.173
22.31
.782
9.31
.238
-
72.3 ± 3.51
SyncLight
31.32
.950
2.47
.119
27.16
.915
4.02
.134
30.34
.895
3.73
.196
-
-
Table 1: Quantitative multi-view image relighting results on our SyncLight test set. Each section of the table indicates which view is used to compute metrics: the “Reference (Ref.) view” (where edits are specified), “Second (Sec.) view” (another view in the set), and “Additional (Add.) views” (all but the reference view). Best results, achieved by SyncLight in all cases, are highlighted. See text for details on baselines. As SyncLight is always run on more than one image; no inference time is reported for the “Ref. view.” VGGTm (%) measures cross-view consistency against the ground truth (see text); it is undefined for the reference view alone.
Figure 3: Qualitative results showcasing SyncLight’s controllability and generalization. Top-left: precise control over light color, varying the chromaticity of a table lamp across three target hues. Top-right: continuous control over light intensity, from off ( 0.0 ) to fully on ( 1.0 ). Bottom-left: robustness to large camera angle changes, with input views rotated by up to 180∘ relative to the reference. Bottom-right: generalization across diverse scenes and unseen light types. SyncLight produces consistent relighting across all viewpoints.
Figure 4: Results in the Real Split from our dataset. SyncLight relights the images correctly. Flux.2-dev struggles with cross-view consistency, producing visually plausible but geometrically inconsistent results (e.g., the window in the second row). Meanwhile, the LightLab*+LumiNet baseline maintains better consistency by explicitly relighting each view, but fails to propagate lighting effects.
Figure 6
Figure 6: Relighting a scene with seven different views. Top row: Input images. Bottom row: SyncLight results. Note how all the views are consistently modified. We want to emphasize views 2 and 4, in which the lamp turned on is not visible, yet the effects in the scene match those of the other views.
Figure 7: Qualitative results on out-of-distribution images from the RealEstate10K dataset Google Research (2018) . SyncLight outperforms all the other methods, including a per-view informed version of LightLab*, demonstrating that the joint processing of all views is beneficial even for single-light relighting.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Examples from the SyncLight dataset. Samples were selected randomly with random light colors.
Table 10Table 11Table 12
Figure 11: Multi-view relighting examples when training SyncLight without one of the splits of the dataset at a time.
Figure 12: Instructions given to the users for the User Study.
Figure 13: User study results.
Figure 14: SyncLight applied to: (a) video relighting, and (b) novel-view synthesis with radiance fields. Please see the supplementary video on our webpage.
Figure 15: Qualitative results on the SyncLight dataset on the Infinigen (1–2), and BlenderKit (3–4) splits. Real split is shown in the main paper.
Figure 16: Inference time (left) and peak memory (right) of SyncLight as a function of the number of views N on an NVIDIA A40 (BlenderKit, 32 rendered views per scene). Memory grows linearly; compute grows quadratically.
N
PSNR
SSIM
ΔE00
LPIPS
VGGTm
Inf. time (s)
2
30.23
.897
3.56
.197
95.8
1.58 ± 0.02
4
30.40
.904
3.55
.194
94.8
2.17 ± 0.02
6
30.45
.907
3.56
.194
94.9
3.01 ± 0.03
7
30.47
.906
3.56
.192
95.3
3.83 ± 0.03
Appendix
Table 9: Scaling with the number of views N on the Real test split.
Method
Interaction per edited light
Cost per additional view
SyncLight
one click + color + intensity
none
LightLab [ 49 ]
segmentation mask + color + intensity
one mask per view
GR3EN [ 63 ]
segmentation mask + color + intensity
one mask per view + camera poses
LuxRemix [ 42 ]
OLAT harmonization + color + intensity
Plücker embeddings + harmonization
Appendix
Table 10: User input required by each method.
Ours vs. GT
ΔE00 vs. input
Region
PSNR
ΔE00
GT
Ours
Weak edits ( ∣ΔL∣<0.5 )
0–1 r (marker)
30.6
3.12
7.9
8.2
1–2 r
30.9
2.84
3.1
3.3
2–4 r
31.8
2.62
1.6
1.7
> 4 r
32.1
2.71
0.9
1.0
Appendix
Table 11: Local versus global changes on the Real split. Weak and strong edits are separated by the requested lightness change ∣ΔL∣ ; pixels are grouped by their distance to the edited source in multiples of the marker radius r . The last row reports all edits over the full image, as in table 1 .
Figure 17: Additional results on the RealEstate10K dataset.
Figure 18: Additional results on the RealEstate10K dataset
Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.
Hejun Wang, Jinxi Li, Junwei Jiang +4
Shenzhen Research Institute, The Hong Kong Polytechnic University, Hong Kong
Indoor scene relighting demands photorealism, precise spatial control, and strict multi-view consistency. While diffusion-based image editing models enable semantic lighting manipulation via text prompts, enforcing exact 3D light placement often disrupts their generative priors. We propose Lume-Palette, a progressive framework that leverages semantic lighting priors for spatially controllable multi-view indoor relighting. The approach decouples relighting into two stages: (1) illumination distillation, which extracts canonical illumination palettes from a pretrained diffusion model to preserve realistic material-light interactions, and (2) illumination casting, which explicitly maps target spatial lighting conditions defined from coarse 3D geometry. To efficiently handle dense multi-view and multi-modal inputs, we introduce an asymmetric multi-view conditioning strategy that selectively injects essential spatial context. Experiments on diverse synthetic scenes and real-world scenes demonstrate that Lume-Palette produces photorealistic, spatially controllable, and multi-view consistent relighting results. Project Page: https://cjeen.github.io/lumepalette
Chenjian Gao, Linning Xu, Tianfan Xue
Multimedia Laboratory, The Chinese University of Hong Kong · Shanghai AI Laboratory · CPII under InnoHK
We present LiveLight, the first diffusion-based framework for real-time streaming video relighting with interactive 3D lighting control. Achieving this is non-trivial, as it requires overcoming three critical challenges: effectively injecting dynamic 3D lighting into a diffusion model, maintaining high-fidelity generation under an extremely low NFE (Number of Function Evaluations) budget for real-time speed, and facilitating continuous streaming for interactive control. To address these pain points, we propose three key designs. First, for accurate lighting injection, we propose a lightweight adapter that feeds Multi-Plane Light Irradiance (MPLI) conditions-depth-aware irradiance maps encoding 3D lighting geometry-directly into the diffusion backbone. Second, to prevent rendering quality degradation at low NFEs towards real-time distillation, we introduce a geometry-guided feedback branch. This training-time constraint leverages a frozen geometry estimator to enforce depth- and normal-consistent relighting, ensuring geometrically plausible shading without adding inference overhead. Finally, to enable streaming interaction, we develop a progressive rolling-window strategy that maintains a denoising ladder of latent chunks at varying noise levels. By propagating intermediate states, this strategy guarantees temporal coherence and supports arbitrarily long video relighting with per-frame reference refresh. Extensive experiments on real-world and synthetic benchmarks demonstrate that LiveLight achieves state-of-the-art relighting quality while running at real-time speed, significantly outperforming offline baselines in temporal stability, lighting controllability, and user preference. To foster real-time interactive relighting research, we will publicly release our models, training data, and synthetic data generator.