ODDR: One-Step Deshadow Diffusion via Reward Guidance
Authors: Junseong Shin, Kijun Kim, Minseong Kim, Dongjin Kim, Tae Hyun Kim
Organizations: Department of Artificial Intelligence, Hanyang University · AX Future Technology Institute, KT · Department of Computer Science, Hanyang University
Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world paired datasets, which are costly to collect and often limited in scene diversity, leading to limited generalization. To address these limitations, we propose One-step Deshadow Diffusion via Reward guidance (ODDR), a new framework that achieves efficient and high-fidelity shadow removal without relying on real-world paired supervision. Our method begins with One-step Deshadow Diffusion (ODD), a baseline model trained on synthetic shadow data for efficient one-step shadow-free reconstruction. We further adapt ODD into ODDR using ShadowReward. In contrast to traditional, annotation-heavy approaches, ShadowReward is the first reward model for shadow removal trained entirely without human annotation. It learns to mimic human perceptual judgments by ranking synthetically generated images with controlled degradations, such as texture distortion and boundary artifacts. This reward-guided fine-tuning enables ODDR to close the synthetic-to-real domain gap. Extensive experiments show that ODD achieves strong performance without relying on real-world paired supervision, and ODDR further improves the results, narrowing the gap to fully supervised methods trained on real-world paired data while maintaining higher computational efficiency as a single-step model.
Figures & tables
Figure 1: We compare our method with diffusion-based shadow removal models, DeS3 [ 22 ] and SSR [ 62 ] on SRD [ 44 ] , LRSS [ 9 ] , and UIUC [ 14 ] datasets, and analyze computational efficiency. Our ODDR achieves the highest restoration quality and simultaneously demands the least FLOPs, inference time, and NFE (Number of Function Evaluations).
Figure 2: Overall training pipeline. (a) The baseline ODD is trained on synthetic shadow images ( xsyn ), optimizing the ODD-LoRA and ODD-Conv modules. (b) ODDR is then fine-tuned from ODD parameters using real-world shadow images ( xreal ) guided by ShadowReward ( rθ ) model, optimizing ODDR-LoRA module.
Figure 3: An example list {x(j)}j=03 of the human-annotation-free dataset. Larger j means more degradation.
Figure 4: Overview of ShadowReward.
D
Methods
Mask-free
Shadow Region
Non-Shadow Region
All Image
D
Methods
Mask-free
PSNR ↑
SSIM ↑
RMSE ↓
PSNR ↑
SSIM ↑
RMSE ↓
PSNR ↑
SSIM ↑
RMSE ↓
Real-world Shadow Image
–
20.83
0.927
39.01
37.46
0.985
2.40
20.46
0.894
8.40
w/ real-world pair
DHAN [ 7 ]
✗
32.92
0.988
9.6
27.15
0.971
7.4
25.66
0.956
7.8
SP+M-Net [ 27 ]
✗
37.60
0.990
6.3
36.02
0.976
3.0
32.94
0.962
3.5
Fu et al. [ 8 ]
✗
36.04
0.978
6.7
31.16
0.892
3.8
29.45
0.861
4.2
BMNet [ 68 ]
✗
37.87
0.991
5.62
37.51
0.985
2.45
33.98
0.972
2.97
Table 1: Quantitative comparison on AISTD [ 27 ] dataset. Red and blue denote the best and second-best results, evaluated among the methods trained without real-world pairs. ✓ denotes mask-free methods and ✗ indicates methods that require shadow masks. D indicates whether real-world paired data is used for training.
Figure 5: Visual comparison on AISTD [ 27 ] . Top: non-shadow preservation (yellow arrows mark competitor failures). Bottom: shadow removal. Ours performs well in both.
SRD [ 44 ]
LRSS [ 9 ]
UIUC [ 14 ]
Method
PSNR ↑
SSIM ↑
RMSE ↓
PSNR ↑
SSIM ↑
RMSE ↓
PSNR ↑
SSIM ↑
RMSE ↓
S3R-Net [ 25 ]
22.51
0.878
10.60
22.54
0.786
10.45
23.92
0.857
8.53
DCS [ 21 ]
21.99
0.868
11.47
22.10
0.774
11.63
23.94
0.852
9.69
G2R [ 36 ]
24.23
0.898
8.40
22.79
0.778
9.64
27.03
0.867
6.23
DeS3 [ 22 ]
22.44
0.900
10.29
21.10
0.767
12.75
19.02
0.822
11.23
SSR [ 62 ]
25.86
0.922
8.25
24.47
0.790
9.89
27.79
0.873
7.61
Table 2: Results on SRD, LRSS, and UIUC. The best results are highlighted in red .
Figure 6: Visual comparisons on SRD [ 44 ] , LRSS [ 9 ] , and UIUC [ 14 ] (top to bottom).
Table 9Table 10
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Type
Operation
j=1
j=2
j=3
T
Gaussian blur
3×3
5×5
7×7
N
Gaussian noise
σ=0.004
σ=0.008
σ=0.012
C
Hue shift
Δh=±0.01
Δh=±0.02
Δh=±0.03
Saturation scale
s∈[0.95,1.05]
s∈[0.90,1.10]
s∈[0.85,1.15]
Brightness scale
b∈[0.97,1.03]
b∈[0.94,1.06]
b∈[0.92,1.08]
L
Gamma correction
γ∈[0.97,1.03]
γ∈[0.94,1.06]
γ∈[0.90,1.10]
Appendix
Table 7: Parameter ranges used to synthesize cumulative degradations (T, N, C, L, B) for the Human-Annotation-Free Dataset. Higher j corresponds to stronger degradation ( j=0 is the clean image). Here, r denotes the boundary dilation radius, Δh the hue shift, s the saturation scaling factor, b the brightness scaling factor, and γ the luminance gamma parameter. All parameters are uniformly sampled within the specified ranges. Larger values (or wider ranges) imply more severe distortions.
List length ( J )
Shadow
All
List length ( J )
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
1
37.29
5.99
33.70
3.06
2
37.78
5.74
33.96
3.01
3
38.37
5.36
34.20
2.97
4
38.17
5.45
34.15
2.97
Appendix
Table 8: Impact of the degradation list length J on shadow removal performance.
Figure 7: User study interface for human preference acquisition. Annotators rank three randomized results (A, B, C) against the input image (red border).
Figure 8: Comparison of reward model ranking accuracy. ShadowReward correctly ranks the shadow removal results in alignment with human preference ( A>B>C ), while existing reward models fail to match human perception.
Figure 9: PSNR is unreliable under GT misalignment, while our ShadowReward (labeled ‘Reward’) provides perceptually consistent evaluation.
Backbone
Shadow
All
Backbone
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
CLIP
37.27
5.91
33.67
3.04
MoCov3
37.72
5.71
33.91
3.01
DINOv2
38.37
5.36
34.20
2.97
Appendix
Table 9: Impact of different feature extractor backbones for ShadowReward.
Method
Shadow
All
Method
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
Mean
37.66
5.77
33.90
3.02
Max
37.62
5.56
33.89
3.02
Top- K (0.05)
38.28
5.48
34.19
2.97
Top- K (0.15)
38.08
5.51
34.12
2.98
Top- K (0.1)
38.37
5.36
34.20
2.97
Appendix
Table 10: Impact of different spatial reduction methods for ShadowReward.
Guidance Method
Shadow
ALL
Guidance Method
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
Shadow Detection
36.69
6.26
33.33
3.11
Conditional GAN
36.79
6.40
33.36
3.12
DINO Discriminator
37.79
5.74
33.91
3.03
ShadowReward (ours)
38.37
5.36
34.20
2.97
Appendix
Table 11: Comparison with alternative guidance strategies on the AISTD dataset. All methods use the same ODD baseline and unpaired fine-tuning protocol.
Method
ODD Train
Shadow
ALL
Method
ODD Train
PSNR ↑
RMSE ↓
PSNR ↑
RMSE ↓
ODD
AISTD
36.67
6.27
33.32
3.11
ODDR
AISTD
38.37
5.36
34.20
2.97
ODD
USR
36.08
6.70
32.86
3.24
ODDR
USR
37.58
5.80
33.71
3.13
Appendix
Table 12: Effect of training data source overlap on the AISTD dataset. USR denotes a disjoint shadow-free dataset used to train ODD.
Figure 10: Visual comparisons of our ODD and ODDR on the AISTD [ 27 ] (top), LRSS [ 9 ] (middle), and UIUC [ 14 ] (bottom) datasets.
Figure 11: Detailed illustration of VAE skip-connection.
Figure 12: Example of failure case on highly textured/non-flat surfaces.
Figure 13: Additional visual comparison on the AISTD [ 27 ] dataset.
Figure 14: Additional visual comparisons on the LRSS [ 9 ] (top four rows) and UIUC [ 14 ] (bottom four rows) datasets.
Figure 15: Additional visual comparison of ODD and ODDR on the AISTD [ 27 ] dataset.
Figure 16: Additional visual comparison of ODD and ODDR on the unseen LRSS [ 9 ] dataset.
Figure 17: Additional visual comparison of ODD and ODDR on the unseen UIUC [ 14 ] dataset.
In the field of deep learning, has seen significant advancements; however, shadow removal remains a persistent challenge owing to the variable sizes and colors of shadows influenced by lighting conditions. This study proposes a novel shadow feature refinement network (SFR-Net), which leverages supervised learning, feature refinement loss, and knowledge distillation to enhance shadow removal performance. A dedicated post-processing algorithm is further introduced to restore natural color consistency in the generated shadow-free images. We evaluated our method on two public datasets: the adjusted image shadow triplet dataset (ISTD+) and the shadow removal dataset (SRD), which demonstrate strong generalization capabilities under diverse conditions. On ISTD+, our model achieved a root mean square error (RMSE) of 3.4627 and structural similarity index measure (SSIM) of 0.9382 across the entire image. On SRD, it recorded an RMSE of 4.3781 and an SSIM of 0.9341. These comprehensive results show that our approach performs competitively across both shadow and non-shadow regions while setting a promising direction for robust and perceptually natural shadow removal. Code is available at https://github.com/DongHyun99/SFRNet.
Existing supervised and unsupervised shadow removal methods often suffer from limited generalization due to the insufficient diversity of available training datasets, while zero-shot methods tend to produce artifacts and require time-consuming test-time optimization. To address these issues, we propose FreeShadow, a training-free shadow removal method built upon pretrained diffusion models, which exploits diffusion priors for shadow removal without any training or optimization. For illumination recovery, we propose an illumination transfer attention (ITA), which re-weights the self-attention maps in diffusion model to transfer illumination cues from non-shadow to shadow regions. For content preservation, we analyze the effects of illumination variations on self-attention maps and latent high-frequency features in diffusion model, and selectively preserve illumination-invariant components to maintain content fidelity while suppressing residual shadows. We further propose local texture-preserving relighting (LTPR) to mitigate local texture misalignment caused by VAE compression. Extensive experiments demonstrate that our method achieves strong generalization and produces realistic shadow-free images.
Yinan Wang, Yan Huang, Yong Xu +1
School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China · Pazhou Lab, Guangzhou 510005, China · Polytech Nantes, Universit´e de Nantes, Nantes 44306, France
We present a three-stage progressive shadow-removal pipeline for the CVPR2026 NTIRE WSRD+ challenge. Built on OmniSR, our method treats deshadowing as iterative direct refinement, where later stages correct residual artefacts left by earlier predictions. The model combines RGB appearance with frozen DINOv2 semantic guidance and geometric cues from monocular depth and surface normals, reused across all stages. To stabilise multi-stage optimisation, we introduce a contraction-constrained objective that encourages non-increasing reconstruction error across the cascade. A staged training pipeline transfers from earlier WSRD pretraining to WSRD+ supervision and final WSRD+ 2026 adaptation with cosine-annealed checkpoint ensembling. On the official WSRD+ 2026 hidden test set, our final ensemble achieved 26.680 PSNR, 0.8740 SSIM, 0.0578 LPIPS, and 26.135 FID, ranked first overall, and won the NTIRE 2026 Image Shadow Removal Challenge. The strong performance of the proposed model is further validated on the ISTD+ and UAV-SC+ datasets.
Lorenzo Beltrame, Jules Salzinger, Filip Svoboda +4
1Austrian Institute of Technology · 2Technical University of Munich · University of Cambridge +1