In this paper, we introduce a metric-guided dynamic loss controller (DLC) for multi-objective image restoration. Conventional image restoration pipelines usually train with a fixed weighted combination of multiple losses, without changing the relative importance of fidelity, perceptual similarity, and no-reference quality during optimization. DLC is an architecture- and loss-term-agnostic training-time controller: it does not modify the restoration architecture or introduce new differentiable loss terms, but dynamically reweights the existing training losses. During training, DLC periodically evaluates the current model on a small fixed feedback subset and uses the resulting quality metrics to update the loss-weight vector through an LLM-based controller. Because DLC operates on existing loss terms rather than task-specific architectures, the same controller formulation can be instantiated across diverse image restoration training pipelines. We evaluate DLC on three restoration domains: low-light image enhancement, deraining, and real-world super-resolution, using both reference-based and no-reference quality metrics. Across these settings, DLC considers metric-dependent trade-offs during optimization and guides training toward balanced operating points across fidelity and perceptual quality. The results show that DLC can move models toward more favorable operating points across different restoration domains, supporting its role as a practical plug-in controller for multi-objective image restoration.
Figures & tables
Figure 1 : Conceptual illustration of the fidelity–perceptual trade-off in image restoration. Fidelity-oriented optimization tends to preserve reference consistency, while perceptual-oriented optimization emphasizes visual sharpness and detail. DLC aims to guide training toward a balanced operating point through metric-guided loss reweighting.
Figure 2 : Overview of DLC. At intervention step k , the current restored output is evaluated by task-relevant metrics to construct a feedback state. The loss controller then produces the next loss-weight vector λk+1 for the following training interval.
Table 2 : Low-light image enhancement results. DLC values are bolded when they improve over the baseline.
Dataset
Method
PSNR ↑
SSIM ↑
LPIPS ↓
NIQE ↓
CLIP-IQA ↑
BRISQUE ↓
Rain100L
Base
33.2795
0.9537
0.0373
3.3761
0.7588
8.1748
Rain100L
DLC
33.5887
0.9526
0.0345
3.4173
0.7396
6.6552
RealRain1K-H
Base
29.6702
0.9280
0.1720
10.4063
0.2349
51.2430
RealRain1K-H
DLC
29.1054
0.9318
0.1513
10.6044
0.2746
51.6570
Table 3 : Deraining results. DLC values are bolded when they improve over the baseline.
Dataset
Method
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
CLIP-IQA ↑
NIQE ↓
MUSIQ ↑
MANIQA ↑
FID ↓
DRealSR
Base
28.3197
0.7804
0.2960
0.2169
0.6970
6.1792
66.1088
0.6161
130.4879
DRealSR
DLC
29.1599
0.7974
0.2806
0.2124
0.6847
6.3306
64.7712
0.6166
130.1198
RealSR
Base
25.5031
0.7418
0.2672
0.2044
0.6699
5.5009
70.1478
0.6551
124.1799
RealSR
DLC
26.1454
0.7490
0.2582
0.2003
0.6652
5.8489
68.4462
0.6627
124.5460
Table 4 : Real-world super-resolution results. DLC values are bolded when they improve over the baseline.
Figure 3 : Qualitative comparison between the fixed-loss baseline and DLC across restoration domains. Zoomed crops highlight texture, structure, and perceptual differences.
Method
PSNR ↑
SSIM ↑
LPIPS ↓
NIQE ↓
CLIP-IQA ↑
BRISQUE ↓
Fixed weight
23.11
0.8548
0.0926
4.4830
0.4580
21.0060
Rule-based
23.28
0.8485
0.0996
4.4940
0.4540
20.8500
GradNorm
22.81
0.8473
0.0934
4.3496
0.4720
19.6089
Uncertainty Weighting
23.02
0.8480
0.0946
4.3853
0.4548
19.9532
CoV Weighting
22.77
0.8482
0.0938
4.3791
0.4845
19.2890
DLC (Ours)
23.50
0.8519
0.0974
4.4457
0.4625
19.9560
Table 5 : Comparison of loss-weighting policies on LOL-v1. All methods use the same backbone, dataset, loss components, optimizer, and training budget, and differ only in the loss-weight decision policy. The best result for each metric is bolded.
Figure 4 : Qualitative comparison of loss-weighting policies on LOL-v1. The results illustrate the different visual characteristics induced by fixed, rule-based, and adaptive weighting strategies.
Controller
PSNR ↑
SSIM ↑
LPIPS ↓
NIQE ↓
CLIP-IQA ↑
BRISQUE ↓
Qwen2.5-7B
23.2304
0.8505
0.0978
4.3938
0.4600
19.4955
Mistral-7B
23.2334
0.8505
0.0979
4.4092
0.4599
19.5009
LLaMA-3 8B
23.5045
0.8519
0.0974
4.4457
0.4625
19.9560
Table 6 : Controller LLM sensitivity on LOL-v1. Best values are bolded.
Figure 5 : Loss-weight trajectory and controller decision example on LOL-v1. Left: representative trajectories of the pixel, SSIM, and perceptual loss weights over intervention steps. Right: an example controller decision at a single intervention step, where fidelity-oriented metrics improve while perceptual or no-reference feedback degrades. DLC adjusts the next loss weights according to the observed metric trade-off. † The loss-weight vector is ordered as (λpix,λssim,λperc) .
Domain
Per-call overhead
Total overhead
Relative overhead
Low-light enhancement
7.49–7.71 s
2.12–6.36 min
<0.2%
Deraining
6.94–9.29 s
1.49–2.17 min
0.16–0.18%
Real-world SR
7.07 s
0.68 min
1.69%
Table 7 : Training-time overhead of DLC measured from representative runs. The reported overhead is incurred only during training, and the final model has no additional test-time cost.
Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-smoothed results, whereas optimizing for perceptual realism often introduces structural deviations. Recent approaches attempt to balance this tradeoff via posterior sampling or multi-stage generative pipelines, yet remain computationally expensive and architecturally complex. To overcome these limitations, we propose PCFlow (Perceptually Consistent Flow Matching), a unified framework that directly parameterizes a continuous transport from degraded observations to clean targets, jointly optimizing distortion and perceptual quality. While its latent consistency flow objective drives stable and efficient few-step inference, a Latent Consistency Perceptual Loss (LCPL) imposes semantic constraints directly on the guiding velocity field, steering the dynamics toward visually sharp data manifolds. Furthermore, recognizing the inherent conflict between structural and perceptual consistencies, we integrate a conflict-free gradient projection strategy to stabilize the multi-objective optimization landscape. Combined with lightweight, convolution-only backbone, PCFlow achieves competitive performance across diverse restoration tasks at a fraction of traditional computational costs.
Sangwoo Jo, Donggeun Ko, Jayeon Kang +3
Korea University, Seoul, South Korea · Aim Future, Seoul, South Korea
All-in-one image restoration aims to handle diverse degradations within a single model. However, existing methods often suffer from three key limitations: 1) per-input computational overhead from dynamic degradation estimation; 2) optimization challenges due to task heterogeneity; and 3) inefficient, frequency-agnostic encoder designs. To overcome these, we introduce the Dynamic Reparameterization Network (DRNet), a novel framework operating on an initialization-stage reconfiguration paradigm that fundamentally eliminates per-input overhead. At its core, a Dynamic Reparameterization MLP (DRMLP) guided by a Task-Specific Modulator (TSM), which effectively mitigates task heterogeneity by orchestrating both specific restoration goals and a versatile general-purpose mode within a unified architecture. Furthermore, we incorporate a Continuous Wavelet Transform Encoder (CWTE) that explicitly leverages frequency characteristics via wavelet decomposition for a lightweight yet powerful design. Extensive experiments demonstrate that DRNet achieves state-of-the-art performance across five restoration tasks with superior parameter efficiency. Crucially, it showcases unique flexibility, excelling as both a highly competitive foundation model for blind restoration and a top-performing user-guided specialist.
Ao Li, Xiaoning Liu, Sheng Li +5
School of Information and Communication Engineering, University of Electronic Science and Technology of China, Chengdu 611731, Sichuan, China · School of Communications and Information Engineering, Chongqing University of Posts and Telecommunication, Chongqing 400065, China
Degraded images not only reduce visual quality but also impair downstream high-level vision tasks. Task-driven image restoration (TDIR) addresses this issue by jointly optimizing restoration quality and task performance. Recent works show that pretrained diffusion priors benefit TDIR, yet diffusion-based restoration is inherently stochastic, as the sampling process depends on a random noise term, which can undermine task consistency. In this paper, we show that a deterministic, noise-free one-step forward pass with pretrained diffusion priors can substantially improve TDIR, but the benefit critically depends on the adaptation module: LoRA yields consistent gains, whereas ControlNet-style conditioning does not. This enables one-step forwarding that surpasses conventional multi-step diffusion TDIR baselines. Furthermore, we introduce a task-preserving GAN training strategy that improves perceptual quality without sacrificing task performance. Extensive experiments on classification, segmentation, and detection demonstrate consistent gains over prior TDIR methods, and we further validate generalization on real-world degraded images and OCR.
Jaeha Kim, Kyoung Mu Lee
Dept. of ECE & ASRI, Seoul National University, Seoul, Korea · IPAI, Seoul National University, Seoul, Korea