Adapting image restoration models to a stream of new tasks without revisiting past data remains challenging due to catastrophic forgetting. In this work, we propose Restoring without Forgetting (RwF), a filter-level continual adaptation framework for image restoration built upon a critical observation: task-specific knowledge is centered in a small subset of filters and can be separated from those reconstructing general content. RwF first performs parameter-space integrated gradients attribution to localize degradation-critical filters in a coarse-to-fine manner. It then adapts to new tasks by generating task-specific filters from a filter bank using compact factorized low-rank transformations, further augmented with cross-task attention and prototypical contrastive learning, and lastly assembles them back only at localized positions. Experiments on six restoration tasks show that RwF effectively avoids forgetting and achieves competitive restoration quality against all-in-one methods that have full data access, and outperforms LoRA-style adaptation with ∼10× fewer additional parameters. Code is available at https://github.com/funkdub/Restoring-without-Forgetting.
Figures & tables
Figure 1: (a) LoRA updates every filter of each adapted layer through low-rank matrices. (b) RwF localizes the few degradation-critical filters of each task and regenerates only them from a shared filter bank, leaving the rest of the backbone frozen.
Figure 2: Overview of RwF. Filter localization (left): parameter-space integrated gradients rank layers, then filters within them, and the selected set is expanded with high-IG, low-similarity filters. Filter learning (right): the localized filters are regenerated from a filter bank by factorized matrices with cross-task attention and assembled back into the frozen backbone.
Figure 3
Figure 5: Visualization of restored images on three randomly selected samples from test set. Bounding boxes emphasize remaining degradation artifacts resulting from catastrophic forgetting.
Table 5
Figure 6: Visual comparison on the four-task public benchmark after learning the last task (LLE).
Method
Deraining
Dehazing
Desnowing
LLE
Average
Param
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
increase
AutoDIR ( Jiang et al., 2024 )
34.49
0.961
29.31
0.970
27.56
0.868
20.54
0.799
27.98
0.900
-
AdaIR ( Cui et al., 2025 )
36.88
0.980
30.34
0.975
28.95
0.901
21.80
0.821
29.49
0.919
-
SFT-MPRNet ( Zamir et al., 2021 )
17.29
0.767
13.89
0.712
15.24
0.698
22.40
0.864
17.21
0.760
-
SFT-DiffUIR ( Zheng et al., 2024 )
23.45
0.819
27.39
0.940
21.20
0.746
21.37
0.860
23.35
0.841
-
SFT-HINT ( Zhou et al., 2025 )
14.84
0.736
10.33
0.674
11.75
0.574
20.04
0.855
14.24
0.711
-
Table 3: Quantitative results on the four-task public restoration benchmark with the order of deraining, dehazing, desnowing, and low-light enhancement (LLE). Best results are in bold and the second are underlined . Note that all-in-one methods are not ranked.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Restored images after resetting 1% of the backbone’s filters, localized for each task, to the shared baseline.
Filters reset
Dehazing
Desnowing
LLE
Localized
Random
Localized
Random
Localized
Random
None (intact)
27.52
27.85
20.23
0.10%
23.15
26.95
26.34
27.75
19.29
20.76
0.25%
21.15
24.01
24.08
27.47
19.71
21.31
0.50%
19.40
23.72
21.20
27.23
17.13
21.83
1.00%
16.10
18.79
20.19
23.73
13.16
18.01
Appendix
Table 4: PSNR (dB) after resetting a fraction of the backbone’s filters to the shared baseline. “Localized” resets filters localized for the task and “Random” resets the same number of randomly selected filters (evaluated on 250 / 150 / 100 test images).
Metric
Deraining
Dehazing
Desnowing
LLE
Average
Routing accuracy (%)
98.0
90.4
100.0
100.0
97.1
PSNR change (dB)
−0.10
−1.18
0.00
0.00
−0.32
Appendix
Table 5: Task identification at inference on the four-task public benchmark. “PSNR change” is the PSNR obtained with the task predicted by the prototype router minus the PSNR obtained with the ground-truth task identity.
Statistic
Number of learned tasks
2
4
6
8
10
Filters localized by the new task
2.05
2.44
2.87
2.95
2.67
Reused from previous tasks
0.0
54.1
71.9
81.1
88.3
Edited filters, random selection ↓
2.05
6.71
10.49
13.43
15.38
Edited filters, RwF (ours) ↓
2.05
5.25
6.83
7.77
8.35
Stored params, separate models ↓
100
300
500
700
900
Appendix
Table 6: Filter consumption and parameter overhead when the number of tasks grows to ten. All numbers are percentages.
Figure 8: t-SNE visualization of latent feature distributions across four restoration tasks, comparing models trained with and without prototypical contrastive learning.
Figure 9: Visualization of restored images on randomly selected samples from test set after expanding to 5 restoration tasks in preliminary experiments, with the training sequence of deraining, denoising, deblurring, dehazing and LLE.
Method
Deraining
Denoising
Deblurring
Dehazing
LLE
Average
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
PSNR
SSIM
SFT
13.64
0.633
14.85
0.521
15.60
0.627
11.70
0.713
28.75
0.939
16.91
0.687
PiggybackGAN ( Zhai et al., 2020 )
32.51
0.895
29.54
0.824
31.46
0.878
27.80
0.867
27.45
0.882
29.75
0.869
RwF-MPRNet (Ours)
38.52
0.942
33.69
0.874
34.49
0.913
32.01
0.920
28.08
0.916
33.36
0.913
Appendix
Table 7: Quantitative results after expanding to 5 restoration tasks, with the training sequence of deraining, denoising, deblurring, dehazing and LLE.
Method
Deraining
Dehazing
Desnowing
LLE
Average
LPIPS
DISTS
LPIPS
DISTS
LPIPS
DISTS
LPIPS
DISTS
LPIPS
DISTS
Degraded input
0.271
0.182
0.104
0.123
0.356
0.209
0.519
0.352
0.313
0.216
DiffUIR (zero-shot) ( Zheng et al., 2024 )
0.120
0.097
0.008
0.014
0.147
0.101
0.139
0.125
0.104
0.084
SFT-DiffUIR ( Zheng et al., 2024 )
0.272
0.183
0.076
0.078
0.270
0.170
0.165
0.139
0.196
0.143
RwF-DiffUIR
0.033
0.038
0.008
0.014
0.145
0.100
0.147
0.126
0.083
0.069
SFT-HINT ( Zhou et al., 2025 )
0.327
0.217
0.272
0.209
0.405
0.232
0.142
0.124
0.287
0.195
Appendix
Table 8: LPIPS and DISTS (lower is better) on the four-task public benchmark with the order of deraining, dehazing, desnowing, and low-light enhancement (LLE), evaluated on all tasks after learning the last task. Best results within each backbone are in bold.
Figure 10: Visualization of restored images on randomly selected samples from test set of four-task public benchmarks.
Figure 11: Visualization of restored images on randomly selected samples from test set of four-task public benchmarks.
Image restoration models are typically trained with a fixed set of capabilities. When new restoration requirements emerge, existing solutions usually train additional models or jointly retrain the original model with both new and historical data. Instead of designing another restoration backbone, we investigate how a trained restorer can continually acquire new capabilities without forgetting those learned previously. We propose RestoreMore, a continual capability-expansion framework that preserves the pretrained restoration model as a frozen capability anchor and learns residual expansion modules for newly arriving degradations. RestoreMore introduces a capability-oriented bi-level routing mechanism at multiple feature stages. The first routing level identifies restoration capabilities relevant to the current input, while the second selects and combines a sparse set of complementary degradation experts. This design enables newly introduced tasks to selectively reuse historical restoration knowledge and progressively enriches the expert bank available for subsequent restoration tasks. Extensive experiments on a wide range of restoration benchmarks demonstrate that RestoreMore consistently acquires new restoration abilities while preserving and improving previously learned capabilities.
Hu Gao, Yulong Chen, Lizhuang Ma
Department of Computer Science Shanghai Jiao Tong University Shanghai, China · Department of Architecture and Design Harbin Institute of Technology Heilongjiang, China
Image restoration seeks to recover high-quality images from degraded inputs but becomes highly ill-posed under complex, mixed degradations. While unified all-in-one models are common, their performance declines as degradation complexity increases. Recent works adopt Chain-of-Thought (CoT) reasoning for multi-round restoration using specialized modules. However, this approach faces two key limitations: (i) increased computational cost due to multi-step processing, and (ii) weak modeling of interactions between degradations during stepwise inference. We introduce CoTIR, a universal image restoration framework that internalizes CoT reasoning within a single model. Concretely, we view image restoration as a specialized subtask of image editing, which implies that a large-scale pre-trained editing model provides a more favorable optimization starting point. Building on this, we fine-tune the model for restoration and further encode structured CoT-style reasoning into the learning objective via a differentiable formulation inspired by Lagrangian optimization, enabling holistic restoration without chaining specialized restorers. To facilitate training and evaluation, we further present CoTIR-Bench, a large-scale benchmark comprising 5.2 million samples with CoT-style reasoning traces. Extensive experiments on CoTIR-Bench and broad real composite degradation scenes show that CoTIR achieves stronger perceptual quality and more competitive fidelity than both all-in-one models and multi-round restoration methods. The source code is available at https://github.com/gy65896/CoTIR.
Yu Guo, Zhengru Fang, Shengfeng He +4
Hong Kong JC Lab of Smart City and the Department of Computer Science, City University of Hong Kong · School of Computing and Information Systems, Singapore Management University · Computer Science and Information Engineering, National Taiwan University
Task-driven image restoration aims to improve both image quality and downstream task performance. However, existing methods predominantly focus on single degradation type and struggle to handle the diverse degradations encountered in real-world scenarios. Different degradations impose distinct restoration demands, and insufficient restoration may leave residual degradations and artifacts that impair object boundaries and semantic cues, thereby compromising downstream task performance. To address these challenges, we propose TaskIR, a two-stage task-driven unified image restoration framework that integrates degradation-adaptive restoration with task feedback refinement. In Stage I, a Degradation Representation Module (DRM) extracts degradation representations, enabling a Degradation-Guided Transformer Block (DGTB) to dynamically modulate feature transformations for adaptive restoration. In Stage II, a Task-to-Restoration Feedback Generation module (TRFG) transforms heterogeneous task features into restoration feedback by modeling task-representation discrepancies associated with the current restoration. Subsequently, a Selective Task Feedback Refinement module (STFR) assesses feedback relevance and selectively refines intermediate restoration features to mitigate interference with well-restored content. Extensive experiments demonstrate that TaskIR achieves competitive restoration quality and downstream task performance across diverse degradations and tasks.
Yanjie Tu, Qingsen Yan, Axi Niu +4
Northwestern Polytechnical University · Shenzhen Research Institute of Northwestern Polytechnical University · Xi’an University of Architecture and Technology