Organizations: School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China · College of Design, Construction, and Planning, University of Florida, Gainesville, FL, USA · School of Japanese Studies, Dalian University of Foreign Languages, Dalian, China
Full-page blind restoration of historical Manchu manuscripts is challenging due to scarce annotations, unknown degradation regions, and fragile connected strokes. Generic restoration models may improve visual quality but often modify intact content, leading to over-restoration. We propose SAGE-Restore (Stroke-Aware Gated rEstoration), a selective restoration framework that first assesses where restoration is needed and then uses this assessment to guide restoration candidate generation and pixel-level selection. Its encoder predicts patch-level repair probabilities from complementary appearance and stroke-structural cues to condition restoration candidate generation, while the corresponding repair logits are refined into a pixel-level soft gate that selectively controls where the restoration candidate is applied. We further introduce a fidelity-aware evaluation protocol that jointly measures degraded-region recovery, intact-content preservation, and their balance. SAGE-Restore achieves the highest R-Recovery (0.463) and RFS (0.626), while maintaining high U-Fidelity (0.968), demonstrating an effective balance between restoration and content preservation.
Figures & tables
Full-reference (FR)
No-reference (NR)
Method
PSNR ↑
SSIM ↑
MAE ↓
MSE ↓
RMSE ↓
MS-S ↑
FSIM ↑
GMSD ↓
VIF ↑
NCC ↑
NIQE ↓
BRIS ↓
IL-N ↓
MAN ↑
MUSIQ ↑
CLIP ↑
Blind-Omni [ 10 ]
23.68
0.98
0.01
0.01
0.07
0.96
0.95
0.10
0.89
0.94
3.93
25.61
34.79
0.54
60.19
0.65
DiffBIR [ 5 ]
17.26
0.58
0.06
0.02
0.14
0.81
0.89
0.20
0.09
0.77
6.80
17.71
40.01
0.59
64.97
0.77
GSDM [ 26 ]
18.89
0.91
0.04
0.02
0.12
0.87
0.84
0.17
0.44
0.78
4.25
30.17
30.85
0.47
61.08
0.68
Table 1: Comparison of full-reference (FR) and no-reference (NR) image quality metrics.
Figure 1: Restoration performance across different methods, with red boxes indicating fidelity preservation and blue boxes indicating restoration quality.
Figure 2: Overview of the proposed SAGE-Restore framework. The SGA encoder predicts patch-level repair logits Lp , whose probabilities P condition restoration candidate generation. The logits are further refined into a pixel-level soft gate Gt , which selectively combines the degraded input Idt and restoration candidate Irt . Restored overlapping tiles are combined using Hann-weighted blending.
Figure 3: Qualitative comparison on two Manchu manuscripts. GSDM recovers degraded strokes but introduces substantial changes to intact content, while Blind-Omni largely preserves the degraded input without recovering missing structures. DiffBIR and HYPIR improve overall appearance but fail to reconstruct key missing strokes. In contrast, SAGE-Restore selectively restores damaged structures while preserving intact handwriting and paper texture, achieving a better balance between recovery and fidelity.
Fidelity-aware
Full-reference (FR)
No-reference (NR)
Method
R-Rec. ↑
U-Fid. ↑
RFS ↑
PSNR ↑
SSIM ↑
MAE ↓
FSIM ↑
GMSD ↓
VIF ↑
NIQE ↓
MUSIQ ↑
CLIP ↑
Blind-Omni
0.033
0.993
0.063
23.677
0.977
0.010
0.954
0.101
0.889
3.932
60.188
0.653
DiffBIR
0.062
0.848
0.116
17.264
0.576
0.057
0.886
0.204
0.088
6.795
64.967
0.774
GSDM
0.420
0.819
0.555
18.888
0.913
0.040
0.837
0.171
0.442
4.250
61.080
0.683
HYPIR
0.047
0.889
0.089
19.399
0.807
0.038
0.923
0.156
0.170
5.867
60.762
0.819
SAGE-Restore
0.463
0.968
0.626
26.077
0.975
0.006
0.971
0.094
0.856
4.056
60.478
0.688
Table 2: Quantitative comparison with state-of-the-art restoration methods. Best results are in bold and second-best results are underlined.
SGA
Gate
R-Rec.
U-Fid.
RFS
PSNR
SSIM
MAE
–
–
0.322
0.881
0.472
21.270
0.873
0.028
✓
–
0.448
0.925
0.596
22.980
0.914
0.014
–
✓
0.405
0.965
0.571
24.780
0.944
0.009
✓
✓
0.463
0.968
0.626
26.077
0.975
0.006
Table 3: Ablation results for SGA and pixel-level gating.
Ancient inscriptions frequently suffer missing or corrupted regions from fragmentation, erosion, or other damage, hindering reading, and analysis. We review prior image restoration methods and their applicability to inscription image recovery, then introduce MESA (Multi-Exemplar, Style-Aware) -an image-level restoration method that uses well-preserved exemplar inscriptions (from the same epigraphic monument, material, or similar letterforms) to guide reconstruction of damaged text. MESA encodes VGG19 convolutional features as Gram matrices to capture exemplar texture, style, and stroke structure; for each neural network layer it selects the exemplar minimizing Mean-Squared Displacement (MSD) to the damaged input. Layer-wise contribution weights are derived from Optical Character Recognition-estimated character widths in the exemplar set to bias filters toward scales matching letter geometry, and a training mask preserves intact regions so synthesis is restricted to damaged areas. We also summarize prior network architectures and exemplar and single-image synthesis, inpainting, and Generative Adversarial Network (GAN) approaches, highlighting limitations that MESA addresses. Comparative experiments demonstrate the advantages of MESA. Finally, we provide a practical roadmap for choosing restoration strategies given available exemplars and metadata.
Vasileios Toulatzis, Sofia Theodoridou, Ioannis Fudos
Digital restoration of historical manuscript images aims to improve readability while preserving the authenticity of cultural heritage documents. However, evaluating quality of restored manuscripts remains challenging, where readability is often subjective and expert annotations are scarce. This study investigates the suitability of contrast-based image quality measures to assess quality and legibility of reconstructed manuscript images from multi-spectral imaging. Two experiments were conducted with publicly-available data sets, facilitating manual quality scores by experts and full-reference image quality measures as reference evaluations. The results show that potential contrast achieves the highest correlation with expert ratings, while contrast-to-noise ratio demonstrates the strongest agreement with full-reference quality measures. Overall, contrast-based measures consistently outperform general image quality measures, demonstrating their potential as objective indicators of manuscript legibility and reconstruction quality.
Anna Breger
Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, UK
Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external historical knowledge. To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG). By combining the implicit knowledge of pre-trained LLMs with explicitly retrieved external context, our model ARI effectively mitigates the challenge of inferring context-dependent proper nouns. Extensive experiments on Korean historical documents demonstrate that our approach significantly outperforms baselines, achieving substantial gains in restoring both general characters and named entities. Furthermore, comprehensive evaluations including expert assessments confirm that ARI serves as a practical tool for domain experts, promising to accelerate the analysis of historical records.
Gabeen Kim, Kyeongpil Kang
Department of AI Convergence, Kangwon National University · Department of Computer Science and Engineering, Kangwon National University