Organizations: Graduate School of Information, Production and Systems, Waseda University, Kitakyushu, Fukuoka 808-0135, Japan · RIKEN Center for Advanced Intelligence Project (AIP), RIKEN, Tokyo 103-0027, Japan · Graduate School of Frontier Sciences, The University of Tokyo, Kashiwa, Chiba 277-8561, Japan · Informatics Institute, University of Amsterdam, Amsterdam 1098 XH, The Netherlands · Wuhan University, Wuhan 430072, China
Remote sensing images often carry composite degradations, in which haze, cloud, noise, blur, low light, and low resolution coexist. Restoring them requires deciding which tool to apply, in what order, and when to stop, yet no clean reference is available at inference time to verify these decisions. All-in-one models trained on single degradations converge to a narrow PSNR band as degradations accumulate. To formulate real-world remote sensing restoration as a traceable trajectory, we present EORestore-Agent, which replaces this unmeasurable objective with reference-free, verifiable per-step decisions. A fine-tuned vision-language model reports all residual degradation types, whose tool pools are scored together, so the restoration order emerges from step-wise selection. A relative quality scorer, trained with full-reference supervision on synthetic degradation chains, predicts the changes in PSNR, SSIM, and LPIPS from the current image to each candidate. A step is accepted only when no predicted change is negative and the predicted PSNR gain is positive. Otherwise, the agent keeps the current image. On a synthetic Landsat-8 benchmark with six degradation types, EORestore-Agent improves PSNR by 2.3 to 3.2 dB over the strongest retrained all-in-one baseline on composites of two to six degradations, whereas zero-shot natural-image agents fall below the degraded input in PSNR in 17 of 18 settings. Replacing the learned scorer with no-reference quality differences costs 1.1 to 4.6 dB. The remaining harmful steps are small and cluster near the acceptance threshold. Sentinel-2 examples illustrate transfer to real atmospheric degradation without retraining.
Figures & tables
Fig. 1: Overview of EORestore-Agent. Top: the iterative restoration loop with perception, merged candidate pool, relative quality scoring, and three-metric gating. Middle: an example trajectory on a tile with four degradation types, showing the selected tool and PSNR at each step. Bottom: detailed view of the perception layer, execution layer, and toolbox.
TABLE II: Comparison on the synthetic benchmark. PSNR (dB, ↑ ), SSIM ( ↑ ), and LPIPS ( ↓ ) averaged over all degradation combinations at each N . Best and second best among restoration methods are in bold and underlined. The degraded input is not ranked. Ties share the mark.
Fig. 2: PSNR versus the number of composited degradations N , plotted from Table II . Each band spans the minimum and maximum of a baseline group, and the line with markers is the group median. The legend gives the number of models in each group.
Fig. 3: Qualitative comparison on the synthetic benchmark, one case per N from 2 to 6. Row labels give the degradation combination. The PSNR of each output is shown at the bottom left, with the best value of each row in bold.
N=2
N=3
N=4
Method
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Restormer
24.91
0.796
0.272
22.23
0.686
0.488
20.44
0.549
0.670
Ada4DIR
21.03
0.760
0.312
20.00
0.662
0.492
19.06
0.551
0.679
PhyDAE
19.03
0.746
0.254
19.93
0.687
0.419
19.10
0.593
0.592
CoRE-UIR
20.55
0.748
0.267
20.04
0.673
0.438
19.45
0.578
0.593
Ours
26.27
0.834
0.255
22.77
0.698
0.442
21.07
0.572
0.626
TABLE III: Comparison on combinations composed entirely of seen degradation types. Only the four types covered by AIO training (haze, noise, blur, low light) are included. Restormer is the strongest baseline at N≥2 in Table II . The N=4 column contains a single combination. Best in bold, second best underlined. Ties share the mark.
N=1
N=2
N=3
Perception
Selection signal
Fallback
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Single-label
Δ^P,S,L
✓
33.58
0.881
0.184
25.90
0.769
0.365
21.89
0.646
0.546
Multi-label
Δ^P
✓
33.66
0.874
0.194
26.05
0.762
0.376
22.06
0.640
0.556
Multi-label
Δ^P,S
✓
33.67
0.880
0.194
26.05
0.770
0.370
22.07
0.647
0.551
Multi-label
Δ^P
×
33.50
0.864
0.210
25.96
0.754
0.382
21.97
0.629
0.564
Multi-label
NIQE-diff
✓
29.42
0.844
0.215
23.42
0.730
0.369
18.94
0.596
0.556
TABLE IV: Ablation on the Full Benchmark (270 Tiles). PSNR (dB, ↑ ), SSIM ( ↑ ), and LPIPS ( ↓ ) are averaged over all combinations at each N . Selection signal: Δ^P,S,L gates on all three predicted changes, Δ^P and Δ^P,S drop scorer heads, and X-diff ranks candidates by the change in the no-reference score X. Fallback × removes the input image from the candidate set. Best in bold, second best underlined; ties share the mark.
Fig. 4: Real-world cases on Sentinel-2 L2A tiles. The reference is the L2A acquisition of the same area from a nearby cloud-free date, for visual comparison only (no full-reference metric). Row labels list the degradation set perceived at step 1. The red box marks the region enlarged at the bottom right of each cell.
Fig. 5: Trajectory statistics. (a) Mean true PSNR gain at each step, one curve per N . Terminated trajectories hold their final value. (b) Share of accepted steps with two or more active types. (c) Share of trajectories that end by abstention, by exhausting the repair budget, or because no degradation is perceived.
Fig. 6: Restoration trajectories, one per N from 2 to 6. Rows 2 to 4 use the tiles of Fig. 3 , rows 1 and 5 a different tile of the same combination. Each solid arrow is an accepted step, labeled with its true PSNR change in dB. Under each frame, the strip marks the perceived types and the type of the applied tool, followed by that type, the tool, and its predicted PSNR change. Each trajectory ends with an abstention (dashed box). No tool is predicted to raise PSNR without worsening SSIM or LPIPS, so the agent keeps the current image and stops. Below the box is the declined tool, the one with the highest predicted PSNR gain. The three-metric gate rejected it because its predicted SSIM or LPIPS change was unfavorable.
Harmful rate (%) ↓
Gain (dB) ↑
N
Random
Ranking
Ours
Random
Ranking
Ours
Oracle
% of oracle
1
46.7
40.0
27.9
3.97
7.68
7.69
8.28
93
2
43.2
30.4
24.7
1.04
3.63
3.57
4.40
81
3
48.9
35.5
27.1
0.07
1.87
1.84
2.57
72
4
48.2
33.8
26.6
0.29
1.65
1.72
2.32
74
5
41.4
28.2
21.4
0.29
1.30
1.22
2.16
56
TABLE V: Single-step comparison of selection rules on the same candidate pools. Harmful rate: share of steps whose chosen candidate lowers true PSNR. Gain: mean true PSNR change of the chosen candidate, over the steps where the full rule accepts a tool. Oracle picks the candidate with the highest true gain. The last column is the gain of Ours as a share of the oracle gain.
TABLE VI: Per-combination results, N=1 . Column labels: Cl cloud, Hz haze, Ns noise, Bl blur, LL low light, LR low resolution. A check in the Seen row marks combinations composed entirely of degradation types covered by AIO training. Method references are given in Table II .
Bl+LL
LL+Hz
Hz+Ns
Bl+LR
Cl+Ns
Cl+LR
Method
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Seen
✓
✓
✓
Degraded input
15.94
0.487
0.535
18.45
0.768
0.386
13.96
0.360
0.745
32.01
0.749
0.517
16.04
0.385
0.695
16.63
0.624
0.640
Natural-image all-in-one
PromptIR
19.70
0.687
0.296
20.59
0.792
0.212
15.38
0.696
0.403
28.22
0.747
0.516
16.95
0.719
0.401
15.56
0.613
0.640
Restormer
26.08
0.790
0.306
26.02
0.871
0.189
22.62
0.727
0.322
29.87
0.753
0.512
16.87
0.716
0.388
15.38
0.609
0.635
Appendix
TABLE VII: Per-combination results, N=2 . Column labels: Cl cloud, Hz haze, Ns noise, Bl blur, LL low light, LR low resolution. A check in the Seen row marks combinations composed entirely of degradation types covered by AIO training. Method references are given in Table II .
Bl+LL+Hz
LL+Hz+Ns
Bl+Cl+LL
Bl+Cl+LR
Cl+Ns+LR
Hz+Ns+LR
Method
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Seen
✓
✓
Degraded input
19.03
0.660
0.584
17.17
0.308
0.954
13.82
0.394
0.683
16.50
0.604
0.673
16.18
0.460
0.857
13.85
0.427
0.868
Natural-image all-in-one
PromptIR
20.25
0.689
0.419
19.53
0.632
0.544
14.72
0.507
0.549
16.05
0.600
0.670
16.45
0.576
0.689
16.76
0.580
0.712
Restormer
22.63
0.749
0.402
21.82
0.622
0.573
16.21
0.625
0.519
15.76
0.599
0.670
15.78
0.569
0.684
22.05
0.608
0.660
Appendix
TABLE VIII: Per-combination results, N=3 . Column labels: Cl cloud, Hz haze, Ns noise, Bl blur, LL low light, LR low resolution. A check in the Seen row marks combinations composed entirely of degradation types covered by AIO training. Method references are given in Table II .
Bl+LL+Hz+Ns
Bl+Cl+LL+LR
Bl+Cl+Hz+LR
Bl+LL+Ns+LR
Cl+LL+Hz+Ns
Cl+Hz+Ns+LR
Method
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Seen
✓
Degraded input
17.33
0.254
1.033
14.21
0.389
0.741
11.94
0.514
0.758
15.51
0.313
0.874
14.52
0.295
0.937
11.95
0.385
0.904
Natural-image all-in-one
PromptIR
19.29
0.576
0.663
16.39
0.588
0.676
17.35
0.595
0.682
15.65
0.391
0.738
15.59
0.563
0.644
13.47
0.512
0.774
Restormer
20.44
0.549
0.670
16.22
0.596
0.674
16.72
0.588
0.683
22.79
0.564
0.695
17.07
0.600
0.542
15.55
0.545
0.703
Appendix
TABLE IX: Per-combination results, N=4 . Column labels: Cl cloud, Hz haze, Ns noise, Bl blur, LL low light, LR low resolution. A check in the Seen row marks combinations composed entirely of degradation types covered by AIO training. Method references are given in Table II .
Bl+Cl+LL+Hz+Ns
Bl+Cl+LL+Hz+LR
Bl+Cl+LL+Ns+LR
Bl+Cl+Hz+Ns+LR
Bl+LL+Hz+Ns+LR
Cl+LL+Hz+Ns+LR
Method
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Seen
Degraded input
14.74
0.226
1.030
14.91
0.529
0.791
14.09
0.282
0.897
11.79
0.355
0.917
17.48
0.401
0.922
14.92
0.381
0.931
Natural-image all-in-one
PromptIR
15.93
0.518
0.725
16.07
0.504
0.736
14.04
0.342
0.806
13.45
0.491
0.799
19.45
0.533
0.781
16.00
0.501
0.798
Restormer
17.27
0.543
0.632
17.02
0.554
0.719
16.24
0.509
0.754
15.20
0.520
0.739
20.10
0.530
0.787
16.90
0.525
0.749
Appendix
TABLE X: Per-combination results, N=5 . Column labels: Cl cloud, Hz haze, Ns noise, Bl blur, LL low light, LR low resolution. A check in the Seen row marks combinations composed entirely of degradation types covered by AIO training. Method references are given in Table II .
Order 1
Order 2
Order 3
Order 4
Method
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Seen
Degraded input
14.37
0.362
0.938
14.97
0.519
0.818
14.14
0.509
0.840
14.32
0.257
0.912
Natural-image all-in-one
PromptIR
15.75
0.484
0.821
16.08
0.494
0.765
15.45
0.474
0.783
15.18
0.461
0.868
Restormer
16.66
0.508
0.774
16.89
0.542
0.736
16.11
0.529
0.754
15.76
0.472
0.838
Appendix
TABLE XI: Per-combination results, N=6 . The four columns apply all six types in different orders: Order 1: Bl → Cl → LL → Hz → Ns → LR, Order 2: Ns → Bl → Cl → LR → LL → Hz, Order 3: Ns → LL → Cl → Bl → LR → Hz, Order 4: LR → Hz → Bl → LL → Cl → Ns. Column labels: Cl cloud, Hz haze, Ns noise, Bl blur, LL low light, LR low resolution. Method references are given in Table II .
Real-world image restoration is challenging due to complex and interacting mixed degradations. Recent agent-based approaches address this problem by composing multiple task-specific restoration tools. However, empirical analysis reveals that their performance is fundamentally limited by implicitly constrained planning spaces and the lack of coordination among independently pretrained tools. To address these issues, we propose OPERA (Optimized Planning-Execution Restoration Agent), a framework that jointly optimizes restoration planning and tool execution in an end-to-end manner. On the planning side, OPERA uses reinforcement learning to directly optimize tool composition over a combinatorial plan space, with the final restoration quality as the reward. On the execution side, OPERA introduces agent-guided co-training of restoration tools, enabling them to learn cooperative behaviors under sequential composition. Extensive experiments on multi-degradation benchmarks and real-world datasets demonstrate that OPERA consistently outperforms both all-in-one restoration models and existing agent-based methods across diverse and complex degradation scenarios.
Feng Zhu, Shuyang Xie, Yihan Zeng +2
Harbin Institute of Technology · Huawei Noah’s Ark Lab
Vision-language agents that orchestrate specialized tools for image restoration (IR) have emerged as a promising method, yet most existing frameworks operate in a training-free manner. They rely on heuristic task scheduling and exhaustive tool traversal, resulting in sub-optimal restoration paths and prohibitive computational cost. We argue that the core bottleneck lies in the absence of a learned policy to make decision, as a vision-language model cannot efficiently handle degradation-aware task ordering and tool composition. To this end, we propose TIR-Agent, a trainable image restoration agent that performs a direct tool-calling policy through a two-stage training pipeline of supervised fine-tuning (SFT) followed by reinforcement learning (RL). Two key designs underpin effective RL training: (i) a random perturbation strategy applied to the SFT data, which broadens the policy's exploration over task schedules and tool compositions, and (ii) a multi-dimensional adaptive reward mechanism that dynamically re-weights heterogeneous image quality metrics to mitigate reward hacking. To support high-throughput, asynchronous GPU-based tool invocation during training, we further develop a globally shared model-call pool. Experiments on both in-domain and out-of-domain degradations show that TIR-Agent outperforms 12 baselines, including 6 all-in-one models, 3 training-free agents, and 3 proprietary models, and achieves over 2.5× inference speedup by eliminating redundant tool executions.
Guoli Jia, Yisheng Zhang, Haote Hu +11
Tsinghua University, Beijing, China · Data Science & Artificial Intelligence Research Institute, China Unicom, Beijing, China · Hunan University, Changsha, China +2
Real-world images rarely suffer from a single degradation, and the order in which degradations are removed substantially affects the final restoration quality, motivating agent-based image restoration (IR), where a vision-language model schedules a pool of pre-built restoration-experts. However, existing training-based agents require O((ND)2) restoration-expert calls per image to construct the Optimal Restoration-action Trajectory Dataset (ORTD), where ND denotes the number of degradation types in the universe D, and couple agent training to a fixed restoration-expert pool, preventing extension to newly introduced restoration-experts without full retraining. To overcome these efficiency and extensibility bottlenecks, we propose \textbf{DiTTo}, a novel order-aware image restoration agent framework consisting of the DiTTo Simulator and the DiTTo Agent. The DiTTo Simulator combines ∪S-IR for single-step restoration-action simulation and AiO-IQA for per-action quality prediction, reducing ORTD construction to O(ND) simulator calls per image; the DiTTo Agent is trained by SFT on the simulator-generated ORTD, followed by \textbf{Order-aware Restoration Alignment (ORA)} that aligns degradation identification, restoration-action-ordering, and output format along independent axes. This enables \textbf{plug-and-play scalable extensibility}: adding a new restoration-expert requires updating only the lightweight ORA stage. On the MiO-100 evaluation set with up to five concurrent degradations, our DiTTo Agent achieves state-of-the-art multi-degradation restoration quality among previous agent-based IR methods.