Generating realistic cast shadows for inserted foreground objects requires reasoning about scene geometry and illumination. However, most learning-based approaches treat shadow generation as an image translation problem and capture these physical relationships only implicitly. This often results in misaligned or implausible shadows. Motivated by the physics of shadow formation, we introduce explicit geometric guidance for outdoor shadow generation. Given a composite image and a foreground object mask, we recover approximate scene geometry and estimate a dominant light direction to derive a coarse shadow estimate via geometric reasoning. While coarse, this estimate provides a spatial anchor for shadow placement. Because illumination cannot always be uniquely inferred from a single image, we predict confidence scores for both lighting and shadow cues and use them to regulate their influence during generation. These cues (shadow mask, light direction, and their confidence scores) condition a diffusion-based generator that refines the estimate into a realistic shadow. Experiments on DESOBAV2 show substantially improved shadow-region fidelity and localization, with an overall 23% lower shadow-region RMSE and 30% lower shadow-mask BER than the prior state-of-the-art method.
Figures & tables
Figure 1: Shadow estimates from approximate geometry and light direction. Given a monocular RGB image and a foreground object mask, we recover an approximate 3D point map from MoGe-2 ( Wang et al., 2025 ) and estimate a single dominant light direction, then infer a shadow estimate using object points. This shadow estimate serves as a condition for our shadow generator.
Figure 2: Framework overview. (a) Shadow mask estimation. A light predictor estimates a dominant 3D light direction from the input image, object mask, and 3D point map. Together with the point map and object mask, this direction is used to construct a coarse physics-guided soft shadow prior through 3D ray-direction alignment. A mask corrector refines the prior into a full-resolution shadow mask. (b) Quality estimation. A quality estimator predicts confidence scores qlight and qmask for the predicted light direction and shadow mask, respectively. (c) Shadow generation. A frozen diffusion backbone is augmented with two trainable conditioning pathways. The confidence-weighted shadow mask provides dense spatial control, while the predicted light direction and its confidence modulate scene features injected through cross-attention for scene-level lighting guidance.
Method
BOS
BOS-free
GR ↓
LR ↓
GS ↑
LS ↑
GB ↓
LB ↓
GR ↓
LR ↓
GS ↑
LS ↑
GB ↓
LB ↓
ShadowGAN ( Zhang et al., 2019b )
8.681
70.459
0.961
0.174
0.470
0.938
19.146
87.149
0.903
0.052
0.483
0.961
Mask-SG ( Hu et al., 2019 )
10.450
73.776
0.938
0.186
0.485
0.966
19.662
89.637
0.895
0.054
0.488
0.972
AR-SG ( Liu et al., 2020 )
8.873
69.336
0.957
0.190
0.463
0.922
19.594
84.939
0.896
0.057
0.468
0.925
SGRNet ( Hong et al., 2022 )
9.017
71.582
0.961
0.189
0.446
0.887
20.883
85.841
0.894
0.056
0.450
0.881
SGDiffusion ( Liu et al., 2024 )
7.366
51.830
0.962
0.332
0.218
0.433
14.664
54.931
0.914
0.168
0.179
0.348
Table 1: Comparison with state-of-the-art methods on DESOBAV2 under BOS and BOS-free settings. GR/LR, GS/LS, and GB/LB denote global/local RMSE, SSIM, and BER, respectively. GAN-based methods are pretrained on DESOBA. Diffusion methods use best-of-five selection; italics denote single-sample results averaged over five independent draws. Best scores are in bold .
Figure 3: Qualitative comparison with state-of-the-art methods. Results in both BOS (with reference background object–shadow pairs) and BOS-free (with a single object–shadow pair) settings. We compare generated images and predicted shadow masks from SGDiffusion ( Liu et al., 2024 ) , GPSDiffusion ( Zhao et al., 2025 ) , and our method with the ground truth. Our method produces more accurately localized shadow masks and visually coherent shadows across both settings.
Figure 4: Challenging cases. ( Top ) Spatially varying illumination makes a single dominant light direction ambiguous. ( Bottom ) An out-of-frame shadow truncates the reference object–shadow pair.
Light
N
Mean
p50 / p90
BOS
375
6.780
5.90 / 11.40
BOS-free
235
10.110
7.86 / 20.94
Table 2: Predictor statistics on the DESOBAV2 test set. N denotes the number of evaluated samples. Light-direction error is reported in degrees on samples with valid directions, while mask quality is reported as full-resolution Dice on all samples.
Score
Split
N
Pearson r
Top 20% Error
Bottom 20% Error
qmask
BOS
500
-0.751
0.368
0.743
BOS-free
250
-0.556
0.430
0.649
qlight
BOS
375
-0.355
4.583
10.461
BOS-free
235
-0.518
6.308
16.240
Table 3: Reliability of predicted quality scores on DESOBAV2 test set. We report Pearson correlation with the corresponding cue error and mean errors among the highest- and lowest-scoring 20% of samples. Mask error is 1−Dice , and light error is angular error in degrees. Negative correlations and lower top-group errors indicate reliable ranking.
Figure 5: Quality-guided generation. We compare forcing qmask=qlight=1 with using their predicted values. Full-strength conditioning can propagate errors from inaccurate cues, whereas the predicted scores attenuate unreliable guidance and produce better-aligned shadows and masks.
Config.
PointMap
Light
Mask
Quality
GRMSE ↓
LRMSE ↓
GSSIM ↑
LSSIM ↑
GBER ↓
LBER ↓
1
–
–
–
–
18.207
52.602
0.906
0.178
0.147
0.271
2
✓
–
–
–
10.041
37.010
0.933
0.317
0.082
0.157
3
✓
†
†
–
10.295
37.976
0.932
0.316
0.089
0.170
4
✓
✓
–
✓
9.873
37.651
0.934
0.319
0.087
0.168
5
✓
✓
✓
✓
9.785
36.126
0.934
0.322
0.079
0.151
Table 4: Ablation of the predictor and quality-network design on the DESOBAV2 BOS-free setting. Configuration 1 is GPSDiffusion. † marks predictors updated during generator training; Configuration 3 applies their outputs without quality gating. Best displayed scores are in bold .
Figure 6: Ablations. Point-map guidance provides coarse shadow placement, while ungated predictor conditioning can introduce errors in shadow extent or direction. Our full model modulates these cues using their predicted quality, improving shadow shape and localization.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Light-direction predictions. Predicted and pseudo-labeled light directions are shown as arrows pointing toward the light source on a hemisphere defined in the input-view coordinate frame. Azimuth and elevation are shown for each direction, together with qlight for the prediction.
Figure 8: Shadow-mask predictions. The geometry-derived soft shadow prior and refined shadow-mask prediction are shown as probability heatmaps, while the ground-truth mask is binary. We also report the scalar quality score qmask for each refined prediction.
q setting
GRMSE ↓
LRMSE ↓
GSSIM ↑
LSSIM ↑
GBER ↓
LBER ↓
Fixed 0
11.00
38.84
0.9244
0.2939
0.0803
0.1535
Fixed 0.5
10.97
38.69
0.9246
0.2964
0.0799
0.1528
Fixed 1
11.24
39.26
0.9240
0.2950
0.0824
0.1573
Predicted
9.78
36.13
0.9335
0.3225
0.0793
0.1514
Appendix
Table 5: Sensitivity to quality conditioning on DESOBAV2 BOS-free. Fixed values are applied to both qmask and qlight at inference. All results use the best-of-five protocol. Best scores are in bold .
Config.
PointMap
Light
Mask
Quality
GRMSE ↓
LRMSE ↓
GSSIM ↑
LSSIM ↑
GBER ↓
LBER ↓
1
–
–
–
–
18.207
52.602
0.906
0.178
0.147
0.271
2
✓
–
–
–
10.041
37.010
0.933
0.317
0.082
0.157
3
✓
†
†
–
10.295
37.976
0.932
0.316
0.089
0.170
4
✓
✓
–
✓
9.873
37.651
0.934
0.319
0.087
0.168
5
✓
bg
✓
✓
9.916
36.996
0.933
0.316
0.083
0.160
6
✓
✓
✓
✓
9.785
36.126
0.934
0.322
0.079
0.151
Appendix
Table 6: Additional conditioning ablations on the DESOBAV2 BOS-free setting. Rows 1–6 are evaluated on all 250 samples. Row 6∗ reevaluates the full model on the 235 samples with valid light-direction pseudo-labels, providing a matched comparison with the oracle in Row 7. bg denotes background-mask conditioning, PL denotes the offline pseudo-labeled light direction, and GT denotes the ground-truth shadow mask. Best learned-setting scores on the full test set are in bold . † indicates that the module is not frozen during training.
Figure 9: Representative limitations. Each group shows the input, our result, and ground truth. ( Left ) Incomplete foreground geometry can lead to an incorrect shadow shape. ( Right ) Without an explicit attenuation cue, the generated shadow can have plausible placement but inaccurate intensity.
Figure 10: Valid approximate light. Examples where the automatically estimated light direction produces a rendered shadow density that closely matches the observed foreground shadow mask. The cast shadows from the object points align well with the observed shadow location and shape, indicating a plausible approximate light direction.
Figure 11: Invalid approximate light. Examples rejected by our manual check. Here, the approximate light either yields cast shadows that noticeably deviate from the observed shadow region (in location or shape), or the observed shadow is too partial to support a reliable elevation estimate. Such cases are discarded from the supervision.
Method
GR ↓
LR ↓
GS ↑
LS ↑
GB ↓
LB ↓
GPSDiffusion ( Zhao et al., 2025 )
6.16
39.50
0.9692
0.5029
0.1513
0.2969
Ours
4.28
33.83
0.9749
0.5637
0.1261
0.2497
Appendix
Table 7: Performance on the 140 DESOBAV2 test tuples without valid light-direction pseudo-labels. GR/LR, GS/LS, and GB/LB denote global/local RMSE, SSIM, and BER, respectively. All results use best-of-five selection. Best scores are in bold .
Figure 12: Additional qualitative results in the BOS-free setting. We compare the generated images and predicted shadow masks from SGDiffusion ( Liu et al., 2024 ) , GPSDiffusion ( Zhao et al., 2025 ) , and our method with the ground truth. Across these examples, our method produces shadows with more accurate placement, shape, and extent.
Figure 13: Additional qualitative results in the BOS setting. We compare the generated images and predicted shadow masks from SGDiffusion ( Liu et al., 2024 ) , GPSDiffusion ( Zhao et al., 2025 ) , and our method with the ground truth. Across these examples, our method produces shadows with more accurate placement, shape, and extent.