ODDR: One-Step Deshadow Diffusion via Reward Guidance
Authors: Junseong Shin, Kijun Kim, Minseong Kim, Dongjin Kim, Tae Hyun Kim
Organizations: Department of Artificial Intelligence, Hanyang University · AX Future Technology Institute, KT · Department of Computer Science, Hanyang University
Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world paired datasets, which are costly to collect and often limited in scene diversity, leading to limited generalization. To address these limitations, we propose One-step Deshadow Diffusion via Reward guidance (ODDR), a new framework that achieves efficient and high-fidelity shadow removal without relying on real-world paired supervision. Our method begins with One-step Deshadow Diffusion (ODD), a baseline model trained on synthetic shadow data for efficient one-step shadow-free reconstruction. We further adapt ODD into ODDR using ShadowReward. In contrast to traditional, annotation-heavy approaches, ShadowReward is the first reward model for shadow removal trained entirely without human annotation. It learns to mimic human perceptual judgments by ranking synthetically generated images with controlled degradations, such as texture distortion and boundary artifacts. This reward-guided fine-tuning enables ODDR to close the synthetic-to-real domain gap. Extensive experiments show that ODD achieves strong performance without relying on real-world paired supervision, and ODDR further improves the results, narrowing the gap to fully supervised methods trained on real-world paired data while maintaining higher computational efficiency as a single-step model.
Figures & tables
Figure 1: We compare our method with diffusion-based shadow removal models, DeS3 [ 22 ] and SSR [ 62 ] on SRD [ 44 ] , LRSS [ 9 ] , and UIUC [ 14 ] datasets, and analyze computational efficiency. Our ODDR achieves the highest restoration quality and simultaneously demands the least FLOPs, inference time, and NFE (Number of Function Evaluations).
Figure 2: Overall training pipeline. (a) The baseline ODD is trained on synthetic shadow images ( xsyn ), optimizing the ODD-LoRA and ODD-Conv modules. (b) ODDR is then fine-tuned from ODD parameters using real-world shadow images ( xreal ) guided by ShadowReward ( rθ ) model, optimizing ODDR-LoRA module.
Figure 3: An example list {x(j)}j=03 of the human-annotation-free dataset. Larger j means more degradation.
Figure 4: Overview of ShadowReward.
D
Methods
Mask-free
Shadow Region
Non-Shadow Region
All Image
D
Methods
Mask-free
PSNR ↑
SSIM ↑
RMSE ↓
PSNR ↑
SSIM ↑
RMSE ↓
PSNR ↑
SSIM ↑
RMSE ↓
Real-world Shadow Image
–
20.83
0.927
39.01
37.46
0.985
2.40
20.46
0.894
8.40
w/ real-world pair
DHAN [ 7 ]
✗
32.92
0.988
9.6
27.15
0.971
7.4
25.66
0.956
7.8
SP+M-Net [ 27 ]
✗
37.60
0.990
6.3
36.02
0.976
3.0
32.94
0.962
3.5
Fu et al. [ 8 ]
✗
36.04
0.978
6.7
31.16
0.892
3.8
29.45
0.861
4.2
BMNet [ 68 ]
✗
37.87
0.991
5.62
37.51
0.985
2.45
33.98
0.972
2.97
Table 1: Quantitative comparison on AISTD [ 27 ] dataset. Red and blue denote the best and second-best results, evaluated among the methods trained without real-world pairs. ✓ denotes mask-free methods and ✗ indicates methods that require shadow masks. D indicates whether real-world paired data is used for training.
Figure 5: Visual comparison on AISTD [ 27 ] . Top: non-shadow preservation (yellow arrows mark competitor failures). Bottom: shadow removal. Ours performs well in both.
SRD [ 44 ]
LRSS [ 9 ]
UIUC [ 14 ]
Method
PSNR ↑
SSIM ↑
RMSE ↓
PSNR ↑
SSIM ↑
RMSE ↓
PSNR ↑
SSIM ↑
RMSE ↓
S3R-Net [ 25 ]
22.51
0.878
10.60
22.54
0.786
10.45
23.92
0.857
8.53
DCS [ 21 ]
21.99
0.868
11.47
22.10
0.774
11.63
23.94
0.852
9.69
G2R [ 36 ]
24.23
0.898
8.40
22.79
0.778
9.64
27.03
0.867
6.23
DeS3 [ 22 ]
22.44
0.900
10.29
21.10
0.767
12.75
19.02
0.822
11.23
SSR [ 62 ]
25.86
0.922
8.25
24.47
0.790
9.89
27.79
0.873
7.61
Table 2: Results on SRD, LRSS, and UIUC. The best results are highlighted in red .
Figure 6: Visual comparisons on SRD [ 44 ] , LRSS [ 9 ] , and UIUC [ 14 ] (top to bottom).
Table 9Table 10
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Type
Operation
j=1
j=2
j=3
T
Gaussian blur
3×3
5×5
7×7
N
Gaussian noise
σ=0.004
σ=0.008
σ=0.012
C
Hue shift
Δh=±0.01
Δh=±0.02
Δh=±0.03
Saturation scale
s∈[0.95,1.05]
s∈[0.90,1.10]
s∈[0.85,1.15]
Brightness scale
b∈[0.97,1.03]
b∈[0.94,1.06]
b∈[0.92,1.08]
L
Gamma correction
γ∈[0.97,1.03]
γ∈[0.94,1.06]
γ∈[0.90,1.10]
Appendix
Table 7: Parameter ranges used to synthesize cumulative degradations (T, N, C, L, B) for the Human-Annotation-Free Dataset. Higher j corresponds to stronger degradation ( j=0 is the clean image). Here, r denotes the boundary dilation radius, Δh the hue shift, s the saturation scaling factor, b the brightness scaling factor, and γ the luminance gamma parameter. All parameters are uniformly sampled within the specified ranges. Larger values (or wider ranges) imply more severe distortions.
List length ( J )
Shadow
All
List length ( J )
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
1
37.29
5.99
33.70
3.06
2
37.78
5.74
33.96
3.01
3
38.37
5.36
34.20
2.97
4
38.17
5.45
34.15
2.97
Appendix
Table 8: Impact of the degradation list length J on shadow removal performance.
Figure 7: User study interface for human preference acquisition. Annotators rank three randomized results (A, B, C) against the input image (red border).
Figure 8: Comparison of reward model ranking accuracy. ShadowReward correctly ranks the shadow removal results in alignment with human preference ( A>B>C ), while existing reward models fail to match human perception.
Figure 9: PSNR is unreliable under GT misalignment, while our ShadowReward (labeled ‘Reward’) provides perceptually consistent evaluation.
Backbone
Shadow
All
Backbone
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
CLIP
37.27
5.91
33.67
3.04
MoCov3
37.72
5.71
33.91
3.01
DINOv2
38.37
5.36
34.20
2.97
Appendix
Table 9: Impact of different feature extractor backbones for ShadowReward.
Method
Shadow
All
Method
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
Mean
37.66
5.77
33.90
3.02
Max
37.62
5.56
33.89
3.02
Top- K (0.05)
38.28
5.48
34.19
2.97
Top- K (0.15)
38.08
5.51
34.12
2.98
Top- K (0.1)
38.37
5.36
34.20
2.97
Appendix
Table 10: Impact of different spatial reduction methods for ShadowReward.
Guidance Method
Shadow
ALL
Guidance Method
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
Shadow Detection
36.69
6.26
33.33
3.11
Conditional GAN
36.79
6.40
33.36
3.12
DINO Discriminator
37.79
5.74
33.91
3.03
ShadowReward (ours)
38.37
5.36
34.20
2.97
Appendix
Table 11: Comparison with alternative guidance strategies on the AISTD dataset. All methods use the same ODD baseline and unpaired fine-tuning protocol.
Method
ODD Train
Shadow
ALL
Method
ODD Train
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
ODD
AISTD
36.67
6.27
33.32
3.11
ODDR
AISTD
38.37
5.36
34.20
2.97
ODD
USR
36.08
6.70
32.86
3.24
ODDR
USR
37.58
5.80
33.71
3.13
Appendix
Table 12: Effect of training data source overlap on the AISTD dataset. USR denotes a disjoint shadow-free dataset used to train ODD.
Figure 10: Visual comparisons of our ODD and ODDR on the AISTD [ 27 ] (top), LRSS [ 9 ] (middle), and UIUC [ 14 ] (bottom) datasets.
Figure 11: Detailed illustration of VAE skip-connection.
Figure 12: Example of failure case on highly textured/non-flat surfaces.
Figure 13: Additional visual comparison on the AISTD [ 27 ] dataset.
Figure 14: Additional visual comparisons on the LRSS [ 9 ] (top four rows) and UIUC [ 14 ] (bottom four rows) datasets.
Figure 15: Additional visual comparison of ODD and ODDR on the AISTD [ 27 ] dataset.
Figure 16: Additional visual comparison of ODD and ODDR on the unseen LRSS [ 9 ] dataset.
Figure 17: Additional visual comparison of ODD and ODDR on the unseen UIUC [ 14 ] dataset.
School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China · Pazhou Lab, Guangzhou 510005, China · Polytech Nantes, Universit´e de Nantes, Nantes 44306, France