Adapting a pre-trained 3D Gaussian Splatting (3DGS) road scene to a new time of day requires learning appearance changes from a few anchor images while preserving consistent, real-time rendering. We propose PAM-ToD, a lightweight plug-in that learns color corrections while keeping the pre-trained 3DGS parameters fixed. PAM-ToD scales each Gaussian's existing color to model illumination changes and uses an additive term for additional brightness, such as when street lamps turn on at night. Under a simplified image formation model, unchanged surface albedo can be eliminated from the relation between source and target appearances, allowing us to learn these corrections without separately estimating albedo and illumination. The model corrects colors across the scene while allowing the corrections to vary by location and by Gaussian. To guide learning from a few anchor images, it discourages abrupt spatial changes in these corrections. We also introduce CARLA-ToD, a benchmark with matching geometry, camera poses, and moving-object trajectories across three times of day. A few target-time anchor images are used to train each plug-in, while separate views are used for evaluation. Across the static and dynamic settings, PAM-ToD achieves higher PSNR and lower LPIPS than the baselines, even when the anchor images come from a single synchronized capture across multiple cameras.
Figures & tables
Figure 1: Overview of PAM-ToD .Given a frozen 3DGS scene from source time s and sparse anchor images at target time t , PAM-ToD adapts scene appearance in real time using a multiplicative illumination term and an additive emission term.
Reconstruction
Novel View Synthesis
FPS ↑
Methods
SSIM ↑
PSNR ↑
LPIPS ↓
SSIM ↑
PSNR ↑
LPIPS ↓
GS-W
0.748
16.76
0.312
0.729
16.58
0.324
22.9
GS-IR
0.663
16.42
0.372
0.671
16.57
0.361
61.8
GI-GS
0.663
16.37
0.353
0.667
16.60
0.342
16.3
LumiGauss
0.491
11.29
0.597
0.467
10.89
0.622
152.7
StreetGS
0.471
10.17
0.487
0.463
10.13
0.486
66.7
Table 1: Quantitative relighting comparison on the CARLA-ToD Static Dataset. The first, second, and third best performances are highlighted in First , Second , and Third , respectively.
Figure 2: Qualitative comparison on the CARLA-ToD Static Dataset.
Reconstruction
Novel View Synthesis
FPS ↑
Methods
SSIM ↑
PSNR ↑
LPIPS ↓
SSIM ↑
PSNR ↑
LPIPS ↓
LumiNet
0.569
14.44
0.422
0.560
14.45
0.421
0.08
DiffusionRenderer
0.541
14.49
0.496
0.529
14.49
0.566
0.08
UniRelight
0.597
15.89
0.484
0.560
15.49
0.577
0.07
StreetGS
0.467
10.54
0.486
0.462
10.51
0.481
59.3
+Ours ( K=1 )
0.715
21.29
0.318
0.706
21.15
0.311
49.4
Table 2: Quantitative relighting comparison on the CARLA-ToD Dynamic Dataset.
Figure 3: Qualitative comparison on the CARLA-ToD Dynamic Dataset.
Table 6
K=1
K=8
Method
SSIM ↑
PSNR ↑
LPIPS ↓
SSIM ↑
PSNR ↑
LPIPS ↓
Base
0.471
10.17
0.487
0.471
10.17
0.487
w/o Illumination term
0.565
12.83
0.431
0.659
16.73
0.375
w/o Emission term
0.717
20.86
0.329
0.783
22.69
0.311
Global Only
0.597
16.91
0.439
0.786
23.86
0.303
Residual Only
0.714
20.60
0.399
0.790
23.58
0.289
Table 5: Ablation on the correction channels and hierarchy levels with K=1,8 .
Figure 4: Qualitative comparison on the Waymo Open Dataset.
Q=32 bins, classes with <2,000 anchor px or <Q Gaussians skipped, 65,536 samples/iter
Seen-unseen
edges {2,4,8,16,32} m, min. 500 Gaussians/cell, update every 200 iters
Appendix
Table 6: Hyperparameters of PAM-ToD, shared by all experiments.
Sensor
Loc (x,y,z)
Rot (roll,pitch,yaw)
Front camera (ID 0)
(0,0,0)
(0,0,0)
Front-left camera (ID 1)
(−0.045,0.116,−0.001)
(−0.05,−0.58,55.00)
Front-right camera (ID 2)
(−0.049,−0.070,0.001)
(−0.55,−0.44,−55.00)
LiDAR
(−1.541,0.021,−0.300)
(−0.70,0.42,−0.01)
Appendix
Table 7: Sensor mounting relative to the front camera, in a camera basis with x forward, y left, and z up. Location in m, rotation in deg.
Method
Gaussians
Params
Size
GS-IR
1.44 M
122.9 M
492 MB
GI-GS
1.58 M
125.9 M
504 MB
GS-W
1.91 M
173.7 M
695 MB
StreetGS
1.48 M
106.2 M
425 MB
+ Ours
1.48 M
113.9 M (+7.7 M)
456 MB (+31 MB)
Appendix
Table 8: Model size on CARLA-ToD Static.
ToD
Metric
GI-GS
GS-IR
LumiGauss
GS-W
Ours ( K=1 )
Ours ( K=8 )
Mo → No
PSNR ↑
11.91
12.37
7.50
19.47
20.93
24.03
SSIM ↑
0.697
0.689
0.428
0.815
0.777
0.827
LPIPS ↓
0.362
0.382
0.609
0.265
0.251
0.237
No → Mo
PSNR ↑
19.81
18.65
7.66
15.63
25.30
27.00
SSIM ↑
0.736
0.707
0.441
0.823
0.798
0.858
LPIPS ↓
0.306
0.318
0.587
0.287
0.278
0.234
Appendix
Table 9: Quantitative comparison of reconstruction on the CARLA-ToD Static Dataset. Mo, No, and Ni denote Morning, Noon, and Night, respectively.
ToD
Metric
GI-GS
GS-IR
LumiGauss
GS-W
Ours ( K=1 )
Ours ( K=8 )
Mo → No
PSNR ↑
12.74
12.98
7.51
16.41
20.70
23.56
SSIM ↑
0.719
0.715
0.400
0.780
0.757
0.806
LPIPS ↓
0.339
0.352
0.634
0.310
0.252
0.235
No → Mo
PSNR ↑
20.42
19.11
7.99
15.65
25.04
27.00
SSIM ↑
0.744
0.731
0.437
0.780
0.788
0.848
LPIPS ↓
0.289
0.301
0.601
0.310
0.273
0.223
Appendix
Table 10: Quantitative comparison of novel view synthesis (NVS) on the CARLA-ToD Static Dataset. Mo, No, and Ni denote Morning, Noon, and Night, respectively.
Figure 5: Qualitative comparison on of the number of anchors.
Figure 6: Qualitative comparison on our key illumination and emission term.
Figure 7: Additional qualitative results for NVS on CARLA-ToD static scenes.
Figure 8: Additional qualitative results for NVS on CARLA-ToD dynamic scenes.
Figure 9: Additional qualitative results for NVS on WOD.
We propose WildSplatter, a feed-forward 3D Gaussian Splatting (3DGS) model for unconstrained images with unknown camera parameters and varying lighting conditions. 3DGS is an effective scene representation that enables high-quality, real-time rendering; however, it typically requires iterative optimization and multi-view images captured under consistent lighting with known camera parameters. WildSplatter is trained on unconstrained photo collections and jointly learns 3D Gaussians and appearance embeddings conditioned on input images. This design enables flexible modulation of Gaussian colors to represent significant variations in lighting and appearance. Our method reconstructs 3D Gaussians from sparse input views in under one second, while also enabling appearance control under diverse lighting conditions. Experimental results demonstrate that our approach outperforms existing pose-free 3DGS methods on challenging real-world datasets with varying illumination.
3D Gaussian Splatting (3DGS) has recently emerged as a powerful explicit representation enabling fast, high-fidelity rendering, making it a promising foundation for closed-loop simulators and perception models in autonomous driving. However, conventional 3DGS implicitly assumes consistent exposure and tone mapping across views. Real driving data violates this assumption due to heterogeneous camera pipelines and dynamic outdoor illumination, baking exposure discrepancies and sensor noise into the radiance field and producing artifacts and inconsistent illumination especially in static backgrounds crucial for realistic simulation. These issues are amplified in autonomous driving, where sparse viewpoints, varying exposures, and outdoor lighting interact, while prior work mainly targets dynamic-object reconstruction and overlooks cross-view photometric consistency. To address this limitation, we introduce P2GS, a physically consistent Gaussian Splatting framework that jointly decomposes a view-invariant linear HDR radiance field, per-view exposure scales, and tone-mapping functions from only LDR images without HDR supervision. P2GS employs a unified optimization strategy grounded in the physical image-formation process, enforcing relative-exposure consistency and HDR-domain radiance regularization. This yields a radiance field robust to inter-camera illumination differences while preserving the real-time efficiency of standard 3DGS. Experiments across real and simulated driving environments show that P2GS matches or surpasses prior methods in LDR reconstruction while providing substantially improved photometric consistency, reliable exposure normalization, and physically coherent illumination across diverse scenes.
Kota Shimomura, Hidehisa Arai, Tsubasa Takahashi +2
Adaptive density control in 3D Gaussian Splatting (3DGS) repeatedly grows the Gaussian population through fixed-cardinality random splitting to discover useful scene structure. However, in vanilla 3DGS, its binary split operator requires many densification rounds to expose fine details, making it a bottleneck for efficient training schedules with fewer iterations. We introduce AdpSplit, an error-driven adaptive split operator that determines the number of split children and initializes the child parameters from L1-pixel-error region statistics, enabling fewer densification iterations, thus reduced training time, while preserving the rendering quality of full-schedule training. Across the MipNeRF360, Deep-Blending, and Tanks&Temples datasets, AdpSplit reduces the training time of multiple accelerated 3DGS pipelines by 9.2%-22.3% as a simple drop-in replacement for the standard split operator. With FastGS, AdpSplit matches the full-schedule PSNR on MipNeRF360 while reducing training time by 16.4%, corresponding to a 12.6x acceleration over vanilla 3DGS.
Yongjae Lee, Jingxing Li, Abhay Kumar Yadav +2
Arizona State University · Johns Hopkins University