Real-world low-quality images suffer from complex mixed degradations, including but not limited to noise, blur, atmospheric effects, etc. Recent agentic methods usually model real-world image restoration (Real-IR) as a sequential tool calling problem over task-specific single-degradation restoration models. This paradigm, however, is fundamentally limited because complex real-world degradations cannot be cleanly undone degradation by degradation, and the tool used for task-specific models caps the capability of the agent system. In this work, we present HarnessIR, an agentic framework for Real-IR by harnessing a multimodal foundation model (MFM) as the executor. HarnessIR consists of five stages: perception and diagnosis, on-demand tool invocation, prompt composition, execution, and verification-driven refinement. Unlike prior agentic Real-IR methods that rely on tool chains assembled from task-specific models, HarnessIR feeds the restoration requirements, the perceptual diagnosis, and the evidence into an MFM that performs restoration in a single pass, followed by verification stages to determine whether the result warrants further processing. Under our harness, off-the-shelf MFMs handle restoration tasks remarkably well, achieving state-of-the-art results on the widely used MiO100 synthetic benchmark. More importantly, by exploiting the strong generalization ability of MFMs, HarnessIR delivers compelling restoration quality on challenging real-world scenes where previous agentic IR systems often struggle. Codes is available at https://github.com/PolyU-VCLab/HarnessIR.
Figures & tables
Figure 1: Top and Bottom-left : Visual results of HarnessIR, which can restore image details without altering their content and texts. Bottom-right : Radar charts comparing different metrics on real-world test sets. HarnessIR achieves much better restoration performance than Nano Banana 2, GPT-Image-2.5, previous all-in-one and agentic IR methods.
Figure 2: Comparison between (a) prior agentic IR pipelines ( top ) and (b) HarnessIR ( bottom ). Prior methods sequentially apply task-specific restoration models planned by VLMs to produce the output. HarnessIR instead uses auxiliary tools for image-specific evidence, composes a restoration instruction for MFM, and invokes verification-driven re-execution when necessary.
Figure 3: Direct use of MFMs may yield higher NR-IQA scores than GT, even when the image content is largely altered. NR-IQA should not be used as an independent metric for IR performance.
Method
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
AirNet
20.30
0.65
0.46
0.25
0.22
0.28
29.00
0.32
0.27
25.1
89.7
29.78
PromptIR
20.66
0.66
0.45
0.26
0.21
0.28
28.98
0.33
0.26
26.4
91.7
30.41
MiOIR
20.76
0.66
0.43
0.25
0.22
0.29
27.29
0.33
0.26
28.5
90.6
37.09
DA-CLIP
20.33
0.64
0.46
0.26
0.22
0.29
28.37
0.33
0.25
19.4
87.7
25.47
InstructIR
20.84
0.60
0.54
0.28
0.24
0.31
31.84
0.37
0.32
21.9
90.6
29.42
Table 1: Group-averaged results on MiO100. The best and second-best results are marked in red and blue , respectively. Results in which HarnessIR improves on its baseline are in bold . Δ means the absolute deviation from the GT score; the smaller, the better. The GT and LQ rows report the raw metric values. D-Score and F-Score are given by the VLM evaluator for degradation removal and content preservation, both on a scale of 0 – 100 ; DF-Score is their per-image geometric mean.
Method
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
AirNet
22.40
0.724
0.440
0.254
0.108
0.107
13.27
0.109
0.133
16.8
85.7
25.2
PromptIR
24.08
0.743
0.416
0.241
0.098
0.099
12.78
0.105
0.114
16.7
91.3
25.3
MiOIR
24.87
0.735
0.418
0.251
0.096
0.086
12.20
0.104
0.124
10.5
93.0
18.9
DA-CLIP
23.87
0.739
0.406
0.243
0.099
0.106
11.83
0.093
0.111
17.8
88.2
26.7
InstructIR
22.92
0.747
0.411
0.236
0.108
0.116
12.34
0.104
0.124
25.4
90.2
36.9
Table 2: Results on Real-Paired-200, which contains 200 real-world LQ images with paired GT from public real-world datasets. Marking and metrics follow Tab. 1 .
Method
No-reference Metrics
VLM-based IQA
MANIQA ↑
CLIP-IQA ↑
MUSIQ ↑
TOPIQ ↑
AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
AirNet
0.525
0.406
42.94
0.307
-0.689
13.4
86.5
21.6
PromptIR
0.531
0.413
43.17
0.310
-0.703
11.2
90.4
21.0
MiOIR
0.537
0.399
44.44
0.321
-0.695
11.4
90.2
22.5
DA-CLIP
0.529
0.395
43.50
0.311
-0.700
12.6
88.3
20.3
InstructIR
0.521
0.383
43.39
0.311
-0.689
21.4
86.0
32.0
Table 3: Results on Real-NoGT-200. Marking and metrics follow Tab. 1 . Without GT, NR-IQA alone cannot correctly reflect the restoration performance, so these columns are grayed out .
Figure 4: Visual comparisons on three test sets. HarnessIR achieves better fidelity and visual quality.
Executor-input configuration
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
Direct MFM (Baseline)
23.94
0.729
0.309
0.181
0.058
0.077
11.91
0.118
0.074
67.9
47.8
52.9
Detailed User Request
24.75
0.733
0.298
0.177
0.054
0.069
10.88
0.105
0.065
71.4
50.9
54.5
Single-VLM Instruction
25.49
0.761
0.263
0.161
0.042
0.044
9.30
0.080
0.057
72.5
54.2
56.6
Two VLMs + Tools (Visual Evidence)
25.62
0.758
0.267
0.161
0.026
0.032
7.39
0.057
0.049
74.4
55.3
57.8
Two VLMs + Tools (Text Evidence)
26.01
0.763
0.261
0.158
0.032
0.024
7.72
0.062
0.045
73.0
58.5
59.6
Table 4: Ablation of executor inputs on Real-Paired-200. All variants use one execution round, without re-execution to isolate the effect of how information is supplied to the executor. The bold row uses the executor-input setting adopted by HarnessIR. Metrics follow Tab. 1 ; the best and second-best results in each column are marked in red and blue .
Refinement strategy
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
HarnessIR (Single-pass)
26.01
0.763
0.261
0.158
0.032
0.024
7.72
0.062
0.045
73.0
58.5
59.6
HarnessIR (In-place refinement)
25.64
0.760
0.265
0.159
0.020
0.016
6.79
0.052
0.032
71.8
55.6
57.2
HarnessIR (Redo from LQ)
26.63
0.775
0.246
0.149
0.020
0.001
5.52
0.028
0.033
75.2
63.6
62.0
GT
-
-
-
-
0.560
0.451
54.36
0.417
-0.901
-
-
-
Table 5: Ablation of refinement strategies on Real-Paired-200. We compare three settings: HarnessIR (Single-pass) , which performs a single MFM execution without re-execution; HarnessIR (In-place refinement) , which applies a revised instruction to the output of the previous round; and HarnessIR (Redo from LQ) which applies a revised instruction to original LQ input. Metrics follow Tab. 1 , the best and second-best results are marked in red and blue .
VLM
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
HarnessIR (Qwen3-VL-32B [ 56 ] )
25.68
0.760
0.267
0.161
0.034
0.024
7.82
0.067
0.050
71.2
56.1
57.5
HarnessIR (GPT-5.6-Sol [ 57 ] )
26.18
0.768
0.253
0.151
0.029
0.017
6.48
0.058
0.048
74.4
59.7
60.5
HarnessIR (GPT-5.6-Luna [ 58 ] )
26.05
0.762
0.259
0.155
0.034
0.027
7.78
0.066
0.059
72.7
58.7
59.1
HarnessIR (Gemini-3.7-Flash [ 44 ] )
26.01
0.763
0.261
0.158
0.032
0.024
7.72
0.062
0.045
73.0
58.5
59.6
GT
-
-
-
-
0.560
0.451
54.36
0.417
-0.901
-
-
-
Table 6: Effect of the VLMs used for diagnoser and composer on Real-Paired-200. Verification-driven re-execution is disabled in all variants. Metrics follow Tab. 1 ; the best and second-best results in each column are marked in red and blue .
Method
VLM
MFM
Tool / Perception
Time (s)
PSNR ↑
DF-Score ↑
Cost ($)
AgenticIR
0.70
–
5.94 / 9.87
88.8
22.41
38.8
0.137
4KAgent
0.53
–
7.47 / 10.08
125.3
22.18
37.4
0.190
Baseline (MFM only)
–
1.00
–
14.7
23.94
52.9
0.070
HarnessIR (One VLM call)
1.00
1.00
–
27.2
25.49
56.6
0.077
HarnessIR (Two VLM calls + Tools)
2.00
1.00
2.40 / –
41.8
26.01
59.6
0.085
HarnessIR (Two VLM calls + Tools, GPT-5.6-Sol)
2.00
1.00
2.40 / –
45.5
26.18
60.5
0.183
Table 7: Average per-image runtime, call counts, restoration performance, and cost on Real-Paired-200. VLM and MFM denote API calls. Tool / Perception: Tool reports local model used as restoration or auxiliary tools, perception reports local VLM model used for perception.
Figure 5: Human selection rates among our methods and baselines.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1: A complete case study illustrating the workflow of HarnessIR, including the input image, depth map and semantic map from auxiliary tools, and the round1 and round2 output images.
Table B.1: Source datasets of the two real-world test sets, with the number of images contributed by each in parentheses.
Figure C.2: Pairwise Align(m) with majority human preference on the real no-GT test set. Dashed line: chance ( 50% ).
Evaluator
PLCC vs. Gemini-3.7-Flash [ 44 ]
D-Score
F-Score
DF-Score
GPT-5.6-Sol [ 57 ]
0.85
0.87
0.94
Claude-Opus-5 [ 77 ]
0.87
0.89
0.92
Appendix
Table C.2: PLCC of D-Score, F-Score, and DF-Score between each alternative VLM evaluator and Gemini-3.7-Flash on restored outputs for the 400 images in Real-Paired-200 and Real-NoGT-200.
Datasets
Method
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
Group A
AirNet
20.70
0.66
0.43
0.24
0.19
0.27
26.37
0.31
0.25
20.6
88.3
22.42
PromptIR
21.29
0.67
0.43
0.25
0.19
0.26
26.14
0.31
0.24
21.6
92.9
24.57
MiOIR
21.74
0.69
0.39
0.23
0.19
0.27
23.93
0.30
0.23
27.5
89.1
35.04
DA-CLIP
20.87
0.65
0.42
0.24
0.19
0.26
24.94
0.30
0.24
20.4
92.9
25.87
InstructIR
21.52
0.64
0.47
0.25
0.21
0.31
27.86
0.34
0.27
24.4
87.9
31.96
Appendix
Table D.3: Per-group results on MiO100, for which Tab. 1 reports the average. Marking and metrics follow Tab. 1 .
Figure E.3: Additional visual comparisons of different methods. HarnessIR more effectively removes degradations while preserving image content. Zoom in for a better view.
Real-world image restoration (IR) is bottlenecked by the scarcity of high-quality paired training data. Synthetic datasets are abundant but often fail to model real-world degradations, while real-world paired datasets are expensive and difficult to capture. As a result, IR models trained on these datasets show limited generalization in real-world scenarios. In this work, we propose Generative Ground Truth (GGT) by using generative multimodal foundation models (MFMs) to produce high-quality (HQ) targets from real-world low-quality (LQ) images. We first conduct a systematic evaluation of nine state-of-the-art MFMs, including Nano-Banana-2 and GPT-Image-2, on images of various scenes and degradation types. The results demonstrate that Nano-Banana-2 with VLM-based adaptive prompting shows the highest capability to synthesize perceptually realistic and content-faithful HQ targets, which can serve as the GGT for the LQ input. We then employ Nano-Banana-2 to build a GGT synthesis pipeline, which involves multi-stage quality control to ensure data reliability, and construct GGT-100K, an LQ-HQ paired dataset comprising 103,707 training pairs and covering diverse scenes and complex real-world degradations. A test set of 500 image pairs is also established. Extensive experiments show that GGT-100K consistently improves the real-world generalization of a wide range of IR models, with particularly strong benefits for finetuning generative models for IR tasks. Our results suggest that MFMs can serve as practical tools for restoration-oriented data generation, and GGT-100K is a useful resource to expand the generalization boundaries of real-world IR models.
Xiangtao Kong, Jixin Zhao, Lingchen Sun +2
The Hong Kong Polytechnic University · OPPO Research Institute
Image restoration seeks to recover high-quality images from degraded inputs but becomes highly ill-posed under complex, mixed degradations. While unified all-in-one models are common, their performance declines as degradation complexity increases. Recent works adopt Chain-of-Thought (CoT) reasoning for multi-round restoration using specialized modules. However, this approach faces two key limitations: (i) increased computational cost due to multi-step processing, and (ii) weak modeling of interactions between degradations during stepwise inference. We introduce CoTIR, a universal image restoration framework that internalizes CoT reasoning within a single model. Concretely, we view image restoration as a specialized subtask of image editing, which implies that a large-scale pre-trained editing model provides a more favorable optimization starting point. Building on this, we fine-tune the model for restoration and further encode structured CoT-style reasoning into the learning objective via a differentiable formulation inspired by Lagrangian optimization, enabling holistic restoration without chaining specialized restorers. To facilitate training and evaluation, we further present CoTIR-Bench, a large-scale benchmark comprising 5.2 million samples with CoT-style reasoning traces. Extensive experiments on CoTIR-Bench and broad real composite degradation scenes show that CoTIR achieves stronger perceptual quality and more competitive fidelity than both all-in-one models and multi-round restoration methods. The source code is available at https://github.com/gy65896/CoTIR.
Yu Guo, Zhengru Fang, Shengfeng He +4
Hong Kong JC Lab of Smart City and the Department of Computer Science, City University of Hong Kong · School of Computing and Information Systems, Singapore Management University · Computer Science and Information Engineering, National Taiwan University
Vision-language agents that orchestrate specialized tools for image restoration (IR) have emerged as a promising method, yet most existing frameworks operate in a training-free manner. They rely on heuristic task scheduling and exhaustive tool traversal, resulting in sub-optimal restoration paths and prohibitive computational cost. We argue that the core bottleneck lies in the absence of a learned policy to make decision, as a vision-language model cannot efficiently handle degradation-aware task ordering and tool composition. To this end, we propose TIR-Agent, a trainable image restoration agent that performs a direct tool-calling policy through a two-stage training pipeline of supervised fine-tuning (SFT) followed by reinforcement learning (RL). Two key designs underpin effective RL training: (i) a random perturbation strategy applied to the SFT data, which broadens the policy's exploration over task schedules and tool compositions, and (ii) a multi-dimensional adaptive reward mechanism that dynamically re-weights heterogeneous image quality metrics to mitigate reward hacking. To support high-throughput, asynchronous GPU-based tool invocation during training, we further develop a globally shared model-call pool. Experiments on both in-domain and out-of-domain degradations show that TIR-Agent outperforms 12 baselines, including 6 all-in-one models, 3 training-free agents, and 3 proprietary models, and achieves over 2.5× inference speedup by eliminating redundant tool executions.
Guoli Jia, Yisheng Zhang, Haote Hu +11
Tsinghua University, Beijing, China · Data Science & Artificial Intelligence Research Institute, China Unicom, Beijing, China · Hunan University, Changsha, China +2