In the field of deep learning, has seen significant advancements; however, shadow removal remains a persistent challenge owing to the variable sizes and colors of shadows influenced by lighting conditions. This study proposes a novel shadow feature refinement network (SFR-Net), which leverages supervised learning, feature refinement loss, and knowledge distillation to enhance shadow removal performance. A dedicated post-processing algorithm is further introduced to restore natural color consistency in the generated shadow-free images. We evaluated our method on two public datasets: the adjusted image shadow triplet dataset (ISTD+) and the shadow removal dataset (SRD), which demonstrate strong generalization capabilities under diverse conditions. On ISTD+, our model achieved a root mean square error (RMSE) of 3.4627 and structural similarity index measure (SSIM) of 0.9382 across the entire image. On SRD, it recorded an RMSE of 4.3781 and an SSIM of 0.9341. These comprehensive results show that our approach performs competitively across both shadow and non-shadow regions while setting a promising direction for robust and perceptually natural shadow removal. Code is available at https://github.com/DongHyun99/SFRNet.
Figures & tables
Figure 1: Network architecture of the proposed shadow feature refinement network.
Figure 2: Architecture of FFC block and spectral transform.
Figure 3: Comparison of SFR-Net’s shadow-removed feature maps. (a) Without Lfr , (b) With Lfr .
Table 2: Quantitative evaluation results of SFR-Net and state-of-the-art shadow removal methods on the SRD dataset.
Method
ISTD+ Dataset
SRD Dataset
RMSE ↓
SSIM ↑
PSNR ↑
RMSE ↓
SSIM ↑
PSNR ↑
ShadowDiffusion [ 14 ]
2.495
0.9424
34.4869
2.7020
0.9755
30.3506
HomoFormer [ 43 ]
1.6334
0.9863
38.7659
3.0385
0.9457
28.8426
Inpaint4Shadow [ 27 ]
3.8885
0.8237
30.2485
2.5566
0.9617
30.0065
ShadowFormer [ 13 ]
1.2962
0.9909
41.0936
2.4080
0.9649
30.6063
BMNet [ 46 ]
1.4613
0.9928
40.0063
2.6261
0.9633
29.9894
Table 3: Quantitative evaluation results of SFR-Net and state-of-the-art shadow removal methods for the non-shadow region of the original shadow image IS on the ISTD+ and SRD datasets.
Figure 6: Division of layers into three stages for analyzing the effectiveness of knowledge distillation
Figure 7: Matte images when Lfr is applied to SFR-Net and when Lfr is not applied to SFR-Net. (a) Is , (b) With Lfr , (c) Without Lfr , (d) Igt
Method
Shadow Region
Non-Shadow Region
All
RMSE ↓
SSIM ↑
PSNR ↑
RMSE ↓
SSIM ↑
PSNR ↑
RMSE ↓
SSIM ↑
PSNR ↑
Config. #1 (Shallow Only)
6.0479
0.9852
35.9776
3.6449
0.9515
32.5604
4.1722
0.9332
30.2424
Config. #2 (Middle Only)
5.5084
0.9858
36.8616
3.4796
0.9542
33.7787
3.9229
0.9369
31.4653
Config. #3 (Deep Only)
5.2107
0.9862
37.4321
3.2502
0.9542
34.0711
3.6675
0.9372
31.8471
Config. #4 (Shallow & Middle)
5.6130
0.9849
36.5651
3.3170
0.9545
33.6416
3.8253
0.9360
31.3321
Config. #5 (Middle & Deep)
5.4761
0.9857
37.1288
3.3211
0.9518
34.0083
3.7688
0.9341
31.7260
Table 4: Evaluation of shadow removal performance when feature refinement loss is applied to each stage of SFR-Net.
Method
Shadow Region
Non-Shadow Region
All
LLaMa
Lmix
Lfr
RMSE ↓
SSIM ↑
PSNR ↑
RMSE ↓
SSIM ↑
PSNR ↑
RMSE ↓
SSIM ↑
PSNR ↑
✓
7.4366
0.9774
34.0190
5.1019
0.9249
29.7530
5.6161
0.8976
27.8767
✓
✓
6.0975
0.9849
35.4867
3.7747
0.9533
31.7673
4.2855
0.9347
29.6694
✓
✓
6.4203
0.9827
35.6343
4.1685
0.9472
31.5645
4.6311
0.9263
29.7403
✓
✓
✓
5.2107
0.9862
37.4321
3.2502
0.9542
34.0711
3.6675
0.9372
31.8471
Table 5: Evaluation of shadow removal performance of SFR-Net when LLaMa , Lmix , and Lfr are applied.
Method
Shadow Region
Non-Shadow Region
All
Processing Time (second)
RMSE ↓
SSIM ↑
PSNR ↑
RMSE ↓
SSIM ↑
PSNR ↑
RMSE ↓
SSIM ↑
PSNR ↑
SFR-Net (default)
5.2107
0.9862
37.4321
3.2502
0.9542
34.0711
3.6675
0.9372
31.8471
–
BGR Space
+ linear regression
4.9750
0.9865
37.8747
3.0631
0.9550
34.5793
3.4627
0.9382
32.3872
0.068
+ polynomial regression
5.2444
0.9860
37.6492
3.2386
0.9553
34.3712
3.6599
0.9380
32.1919
0.187
+ histogram matching
10.8398
0.9698
32.9278
3.2208
0.9541
34.2324
5.0366
0.9185
29.9032
0.050
Table 6: Quantitative comparison of different post-processing techniques on the ISTD+ dataset.
Method
Shadow Region
Non-Shadow Region
All
Processing Time (second)
RMSE ↓
SSIM ↑
PSNR ↑
RMSE ↓
SSIM ↑
PSNR ↑
RMSE ↓
SSIM ↑
PSNR ↑
SFR-Net (default)
3.7840
0.9719
32.9293
2.5129
0.9690
33.5797
5.2807
0.9266
29.4810
–
BGR Space
+ linear regression
3.5184
0.9744
33.1909
1.9665
0.9727
33.7399
4.3781
0.9341
29.8287
0.046
+ polynomial regression
3.8751
0.9719
32.8250
2.4020
0.9703
33.6360
5.1542
0.9277
29.4965
0.146
+ histogram matching
6.8171
0.9445
28.2236
1.9761
0.9743
33.7197
6.9081
0.9037
26.4355
0.030
Table 7: Quantitative comparison of different post-processing techniques on the SRD dataset.
Figure 8: Comparison of color histograms when color restoration post-processing is applied to SFR-Net. (a): Blue Space Pixel Histogram, (b): Green Space Pixel Histogram, (c): Red Space Pixel Histogram, (d): Igt , (e): Is , (f): Isf , (g): Isf−pp
Figure 9: Qualitative evaluation results of post-processing algorithm for color restoration. (a): Isf , (b): linear regression, (c): polynomial regression, (d): Reinhard et al. [ 33 ] , (e): gradient boosting, (f): gamma correction, (g): Is
Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world paired datasets, which are costly to collect and often limited in scene diversity, leading to limited generalization. To address these limitations, we propose One-step Deshadow Diffusion via Reward guidance (ODDR), a new framework that achieves efficient and high-fidelity shadow removal without relying on real-world paired supervision. Our method begins with One-step Deshadow Diffusion (ODD), a baseline model trained on synthetic shadow data for efficient one-step shadow-free reconstruction. We further adapt ODD into ODDR using ShadowReward. In contrast to traditional, annotation-heavy approaches, ShadowReward is the first reward model for shadow removal trained entirely without human annotation. It learns to mimic human perceptual judgments by ranking synthetically generated images with controlled degradations, such as texture distortion and boundary artifacts. This reward-guided fine-tuning enables ODDR to close the synthetic-to-real domain gap. Extensive experiments show that ODD achieves strong performance without relying on real-world paired supervision, and ODDR further improves the results, narrowing the gap to fully supervised methods trained on real-world paired data while maintaining higher computational efficiency as a single-step model.
Junseong Shin, Kijun Kim, Minseong Kim +2
Department of Artificial Intelligence, Hanyang University · AX Future Technology Institute, KT · Department of Computer Science, Hanyang University
Shadow removal is an important preprocessing step for many vision tasks, yet existing supervised methods require paired shadow and shadow-free images, while unsupervised approaches often still rely on shadow masks or shadow-free references. We propose ShadowCLR, an unsupervised framework that learns shadow removal directly from shadow images. Our key observation is that shadows vary across observations while the underlying scene content remains largely consistent. We therefore use consistency across shadow observations as regularization, encouraging the model to recover scene-consistent appearance while suppressing shadow-specific variations. Global and local consistency further enable us to explore visually related images, learn from imperfectly aligned observations, and focus the representation on shared scene information. Experiments on multiple benchmarks show that ShadowCLR achieves competitive and often superior performance over state-of-the-art unsupervised methods, demonstrating that consistency can provide regularization for shadow removal without shadow masks or shadow-free images.
Anh-Kiet Duong, Petra Gomez-Krämer, Jean-Michel Carozza
L3i Laboratory, La Rochelle University, 17042 La Rochelle Cedex 1, France · LIENSs Laboratory, La Rochelle University, 17042 La Rochelle Cedex 1, France
Existing supervised and unsupervised shadow removal methods often suffer from limited generalization due to the insufficient diversity of available training datasets, while zero-shot methods tend to produce artifacts and require time-consuming test-time optimization. To address these issues, we propose FreeShadow, a training-free shadow removal method built upon pretrained diffusion models, which exploits diffusion priors for shadow removal without any training or optimization. For illumination recovery, we propose an illumination transfer attention (ITA), which re-weights the self-attention maps in diffusion model to transfer illumination cues from non-shadow to shadow regions. For content preservation, we analyze the effects of illumination variations on self-attention maps and latent high-frequency features in diffusion model, and selectively preserve illumination-invariant components to maintain content fidelity while suppressing residual shadows. We further propose local texture-preserving relighting (LTPR) to mitigate local texture misalignment caused by VAE compression. Extensive experiments demonstrate that our method achieves strong generalization and produces realistic shadow-free images.
Yinan Wang, Yan Huang, Yong Xu +1
School of Computer Science and Engineering, South China University of Technology, Guangzhou 510006, China · Pazhou Lab, Guangzhou 510005, China · Polytech Nantes, Universit´e de Nantes, Nantes 44306, France