Real-world low-quality images suffer from complex mixed degradations, including but not limited to noise, blur, atmospheric effects, etc. Recent agentic methods usually model real-world image restoration (Real-IR) as a sequential tool calling problem over task-specific single-degradation restoration models. This paradigm, however, is fundamentally limited because complex real-world degradations cannot be cleanly undone degradation by degradation, and the tool used for task-specific models caps the capability of the agent system. In this work, we present HarnessIR, an agentic framework for Real-IR by harnessing a multimodal foundation model (MFM) as the executor. HarnessIR consists of five stages: perception and diagnosis, on-demand tool invocation, prompt composition, execution, and verification-driven refinement. Unlike prior agentic Real-IR methods that rely on tool chains assembled from task-specific models, HarnessIR feeds the restoration requirements, the perceptual diagnosis, and the evidence into an MFM that performs restoration in a single pass, followed by verification stages to determine whether the result warrants further processing. Under our harness, off-the-shelf MFMs handle restoration tasks remarkably well, achieving state-of-the-art results on the widely used MiO100 synthetic benchmark. More importantly, by exploiting the strong generalization ability of MFMs, HarnessIR delivers compelling restoration quality on challenging real-world scenes where previous agentic IR systems often struggle. Codes is available at https://github.com/PolyU-VCLab/HarnessIR.
Figures & tables
Figure 1: Top and Bottom-left : Visual results of HarnessIR, which can restore image details without altering their content and texts. Bottom-right : Radar charts comparing different metrics on real-world test sets. HarnessIR achieves much better restoration performance than Nano Banana 2, GPT-Image-2.5, previous all-in-one and agentic IR methods.
Figure 2: Comparison between (a) prior agentic IR pipelines ( top ) and (b) HarnessIR ( bottom ). Prior methods sequentially apply task-specific restoration models planned by VLMs to produce the output. HarnessIR instead uses auxiliary tools for image-specific evidence, composes a restoration instruction for MFM, and invokes verification-driven re-execution when necessary.
Figure 3: Direct use of MFMs may yield higher NR-IQA scores than GT, even when the image content is largely altered. NR-IQA should not be used as an independent metric for IR performance.
Method
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
AirNet
20.30
0.65
0.46
0.25
0.22
0.28
29.00
0.32
0.27
25.1
89.7
29.78
PromptIR
20.66
0.66
0.45
0.26
0.21
0.28
28.98
0.33
0.26
26.4
91.7
30.41
MiOIR
20.76
0.66
0.43
0.25
0.22
0.29
27.29
0.33
0.26
28.5
90.6
37.09
DA-CLIP
20.33
0.64
0.46
0.26
0.22
0.29
28.37
0.33
0.25
19.4
87.7
25.47
InstructIR
20.84
0.60
0.54
0.28
0.24
0.31
31.84
0.37
0.32
21.9
90.6
29.42
Table 1: Group-averaged results on MiO100. The best and second-best results are marked in red and blue , respectively. Results in which HarnessIR improves on its baseline are in bold . Δ means the absolute deviation from the GT score; the smaller, the better. The GT and LQ rows report the raw metric values. D-Score and F-Score are given by the VLM evaluator for degradation removal and content preservation, both on a scale of 0 – 100 ; DF-Score is their per-image geometric mean.
Method
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
AirNet
22.40
0.724
0.440
0.254
0.108
0.107
13.27
0.109
0.133
16.8
85.7
25.2
PromptIR
24.08
0.743
0.416
0.241
0.098
0.099
12.78
0.105
0.114
16.7
91.3
25.3
MiOIR
24.87
0.735
0.418
0.251
0.096
0.086
12.20
0.104
0.124
10.5
93.0
18.9
DA-CLIP
23.87
0.739
0.406
0.243
0.099
0.106
11.83
0.093
0.111
17.8
88.2
26.7
InstructIR
22.92
0.747
0.411
0.236
0.108
0.116
12.34
0.104
0.124
25.4
90.2
36.9
Table 2: Results on Real-Paired-200, which contains 200 real-world LQ images with paired GT from public real-world datasets. Marking and metrics follow Tab. 1 .
Method
No-reference Metrics
VLM-based IQA
MANIQA ↑
CLIP-IQA ↑
MUSIQ ↑
TOPIQ ↑
AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
AirNet
0.525
0.406
42.94
0.307
-0.689
13.4
86.5
21.6
PromptIR
0.531
0.413
43.17
0.310
-0.703
11.2
90.4
21.0
MiOIR
0.537
0.399
44.44
0.321
-0.695
11.4
90.2
22.5
DA-CLIP
0.529
0.395
43.50
0.311
-0.700
12.6
88.3
20.3
InstructIR
0.521
0.383
43.39
0.311
-0.689
21.4
86.0
32.0
Table 3: Results on Real-NoGT-200. Marking and metrics follow Tab. 1 . Without GT, NR-IQA alone cannot correctly reflect the restoration performance, so these columns are grayed out .
Figure 4: Visual comparisons on three test sets. HarnessIR achieves better fidelity and visual quality.
Executor-input configuration
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
Direct MFM (Baseline)
23.94
0.729
0.309
0.181
0.058
0.077
11.91
0.118
0.074
67.9
47.8
52.9
Detailed User Request
24.75
0.733
0.298
0.177
0.054
0.069
10.88
0.105
0.065
71.4
50.9
54.5
Single-VLM Instruction
25.49
0.761
0.263
0.161
0.042
0.044
9.30
0.080
0.057
72.5
54.2
56.6
Two VLMs + Tools (Visual Evidence)
25.62
0.758
0.267
0.161
0.026
0.032
7.39
0.057
0.049
74.4
55.3
57.8
Two VLMs + Tools (Text Evidence)
26.01
0.763
0.261
0.158
0.032
0.024
7.72
0.062
0.045
73.0
58.5
59.6
Table 4: Ablation of executor inputs on Real-Paired-200. All variants use one execution round, without re-execution to isolate the effect of how information is supplied to the executor. The bold row uses the executor-input setting adopted by HarnessIR. Metrics follow Tab. 1 ; the best and second-best results in each column are marked in red and blue .
Refinement strategy
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
HarnessIR (Single-pass)
26.01
0.763
0.261
0.158
0.032
0.024
7.72
0.062
0.045
73.0
58.5
59.6
HarnessIR (In-place refinement)
25.64
0.760
0.265
0.159
0.020
0.016
6.79
0.052
0.032
71.8
55.6
57.2
HarnessIR (Redo from LQ)
26.63
0.775
0.246
0.149
0.020
0.001
5.52
0.028
0.033
75.2
63.6
62.0
GT
-
-
-
-
0.560
0.451
54.36
0.417
-0.901
-
-
-
Table 5: Ablation of refinement strategies on Real-Paired-200. We compare three settings: HarnessIR (Single-pass) , which performs a single MFM execution without re-execution; HarnessIR (In-place refinement) , which applies a revised instruction to the output of the previous round; and HarnessIR (Redo from LQ) which applies a revised instruction to original LQ input. Metrics follow Tab. 1 , the best and second-best results are marked in red and blue .
VLM
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
HarnessIR (Qwen3-VL-32B [ 56 ] )
25.68
0.760
0.267
0.161
0.034
0.024
7.82
0.067
0.050
71.2
56.1
57.5
HarnessIR (GPT-5.6-Sol [ 57 ] )
26.18
0.768
0.253
0.151
0.029
0.017
6.48
0.058
0.048
74.4
59.7
60.5
HarnessIR (GPT-5.6-Luna [ 58 ] )
26.05
0.762
0.259
0.155
0.034
0.027
7.78
0.066
0.059
72.7
58.7
59.1
HarnessIR (Gemini-3.7-Flash [ 44 ] )
26.01
0.763
0.261
0.158
0.032
0.024
7.72
0.062
0.045
73.0
58.5
59.6
GT
-
-
-
-
0.560
0.451
54.36
0.417
-0.901
-
-
-
Table 6: Effect of the VLMs used for diagnoser and composer on Real-Paired-200. Verification-driven re-execution is disabled in all variants. Metrics follow Tab. 1 ; the best and second-best results in each column are marked in red and blue .
Method
VLM
MFM
Tool / Perception
Time (s)
PSNR ↑
DF-Score ↑
Cost ($)
AgenticIR
0.70
–
5.94 / 9.87
88.8
22.41
38.8
0.137
4KAgent
0.53
–
7.47 / 10.08
125.3
22.18
37.4
0.190
Baseline (MFM only)
–
1.00
–
14.7
23.94
52.9
0.070
HarnessIR (One VLM call)
1.00
1.00
–
27.2
25.49
56.6
0.077
HarnessIR (Two VLM calls + Tools)
2.00
1.00
2.40 / –
41.8
26.01
59.6
0.085
HarnessIR (Two VLM calls + Tools, GPT-5.6-Sol)
2.00
1.00
2.40 / –
45.5
26.18
60.5
0.183
Table 7: Average per-image runtime, call counts, restoration performance, and cost on Real-Paired-200. VLM and MFM denote API calls. Tool / Perception: Tool reports local model used as restoration or auxiliary tools, perception reports local VLM model used for perception.
Figure 5: Human selection rates among our methods and baselines.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A.1: A complete case study illustrating the workflow of HarnessIR, including the input image, depth map and semantic map from auxiliary tools, and the round1 and round2 output images.
Table B.1: Source datasets of the two real-world test sets, with the number of images contributed by each in parentheses.
Figure C.2: Pairwise Align(m) with majority human preference on the real no-GT test set. Dashed line: chance ( 50% ).
Evaluator
PLCC vs. Gemini-3.7-Flash [ 44 ]
D-Score
F-Score
DF-Score
GPT-5.6-Sol [ 57 ]
0.85
0.87
0.94
Claude-Opus-5 [ 77 ]
0.87
0.89
0.92
Appendix
Table C.2: PLCC of D-Score, F-Score, and DF-Score between each alternative VLM evaluator and Gemini-3.7-Flash on restored outputs for the 400 images in Real-Paired-200 and Real-NoGT-200.
Datasets
Method
Full-reference Metrics
No-reference Metrics
VLM-based IQA
PSNR ↑
SSIM ↑
LPIPS ↓
DISTS ↓
Δ MANIQA ↓
Δ CLIP-IQA ↓
Δ MUSIQ ↓
Δ TOPIQ ↓
Δ AFINE-NR ↓
D-Score ↑
F-Score ↑
DF-Score ↑
Group A
AirNet
20.70
0.66
0.43
0.24
0.19
0.27
26.37
0.31
0.25
20.6
88.3
22.42
PromptIR
21.29
0.67
0.43
0.25
0.19
0.26
26.14
0.31
0.24
21.6
92.9
24.57
MiOIR
21.74
0.69
0.39
0.23
0.19
0.27
23.93
0.30
0.23
27.5
89.1
35.04
DA-CLIP
20.87
0.65
0.42
0.24
0.19
0.26
24.94
0.30
0.24
20.4
92.9
25.87
InstructIR
21.52
0.64
0.47
0.25
0.21
0.31
27.86
0.34
0.27
24.4
87.9
31.96
Appendix
Table D.3: Per-group results on MiO100, for which Tab. 1 reports the average. Marking and metrics follow Tab. 1 .
Figure E.3: Additional visual comparisons of different methods. HarnessIR more effectively removes degradations while preserving image content. Zoom in for a better view.
Hong Kong JC Lab of Smart City and the Department of Computer Science, City University of Hong Kong · School of Computing and Information Systems, Singapore Management University · Computer Science and Information Engineering, National Taiwan University
Tsinghua University, Beijing, China · Data Science & Artificial Intelligence Research Institute, China Unicom, Beijing, China · Hunan University, Changsha, China +2