DISRQAD: Diffusion Image Super-Resolution Quality Assessment Dataset and Benchmark
Organizations: Lomonosov Moscow State University Moscow, Russia · MSU Institute for AI, Lomonosov Moscow State University Moscow, Russia
Abstract
Diffusion-based image super-resolution (SR) can create visually plausible detail that is not supported by the low-resolution input. We introduce DISRQAD, a subjective-quality dataset and diagnostic benchmark for this setting. It contains mean opinion scores (MOS) for 14,000 SR outputs from ten diffusion and four non-diffusion methods, spanning four low-resolution degradation conditions and x2/x4 upscaling. We evaluate 51 standard full-reference and no-reference metric configurations and 11 adapted variants. Agreement with MOS is substantially weaker on diffusion outputs: the strongest standard no-reference baseline reaches 0.431 SRCC on diffusion SR versus 0.813 on non-diffusion SR. As a case study in benchmark use, a pruned and distilled Q-ReAlign-mini student reaches 0.496 SRCC on diffusion SR. DISRQAD measures perceived output quality, not faithfulness to the input; it enables analysis of metric behavior across generator families and input conditions. Our findings reveal a substantial gap in the assessment of diffusion-based SR and provide a basis for developing quality models sensitive to diffusion-specific artifacts.
Figures & tables
| Dataset | Sources | SR outputs | Non-diff. SR | Diffusion SR | Human input |
|---|---|---|---|---|---|
| CVIU-2017 ( Ma et al., 2017 ) | 30 | 1,620 | 9 | 0 | Ratings |
| QADS ( Zhou et al., 2019 ) | 20 | 980 | 21 | 0 | Pairs |
| SISRSet ( Shi et al., 2019 ) | 15 | 360 | 8 | 0 | Pairs |
| RealSRQ ( Jiang et al., 2022 ) | 60 | 1,620 | 10 | 0 | Pairs |
| SISAR ( Zhao et al., 2022 ) | 100 | 12,600 ∗ | 10 ∗ | 0 | Ratings |
| SRIQA-Bench ( Chen et al., 2025b ) | 100 | 1,000 | 5 | 5 | Pairs |
| Ratings per image | SROCC | 95% confidence interval | |
|---|---|---|---|
| Lower bound | Upper bound | ||
| 10 | 0.982547 | 0.969290 | 0.990110 |
| 11 | 0.978173 | 0.961658 | 0.987619 |
| 12 | 0.973870 | 0.954175 | 0.985165 |
| 13 | 0.983084 | 0.970229 | 0.990415 |
| 14 | 0.983596 | 0.971124 | 0.990706 |
| Non-diffusion SR | Diffusion SR | ||||
| Metric | Type | SRCC | PLCC | SRCC | PLCC |
| TOPIQ(FR) ( Chen et al., 2024 ) | FR | ||||
| AFINE(FR) ( Chen et al., 2025b ) | FR | ||||
| VSI ( Zhang et al., 2014 ) | FR | ||||
| GMSD ( Xue et al., 2014 ) | FR | ||||
| LPIPS-VGG ( Zhang et al., 2018 ) | FR | ||||
| SRCC with MOS | |||||
| Factor | Subset | PSNR (FR) | TOPIQ-FR | Q-Align (NR) | Q-ReAlign-mini (distilled, NR) |
| Diffusion SR | |||||
| Overall | All outputs | ||||
| SR method | ResShift | ||||
| SinSR | |||||
| SeeSR | |||||
| Non-diffusion SR | Diffusion SR | ||||
| Metric | Type | SRCC | PLCC | SRCC | PLCC |
| LPIPS-VGG(FT) ( Zhang et al., 2018 ) | FR | ||||
| LPIPS-AlexNet(FT) ( Zhang et al., 2018 ) | FR | ||||
| AHIQ(FT) ( Lao et al., 2022 ) | FR | ||||
| TOPIQ(FR)(FT) ( Chen et al., 2024 ) | FR | ||||
| PieAPP(FT) ( Prashnani et al., 2018 ) | FR | ||||
| Model | Params | VRAM | FPS | SRIQA-Bench | Diffusion | Non-diff. |
|---|---|---|---|---|---|---|
| Q-ReAlign-mini | 0.8B | 2.3 GB | 3.76 | 0.479/0.504 | 0.341/0.393 | 0.683/0.597 |
| Q-ReAlign-lite | 4.0B | 10.7 GB | 2.69 | 0.475/0.519 | 0.405/0.458 | 0.772/0.719 |
| Q-ReAlign-pro | 9.0B | 19.2 GB | 2.65 | 0.428/0.488 | 0.431/0.469 | 0.760/0.716 |
| Q-Align | 8.2B | 16.0 GB | 5.39 | 0.624/0.677 | 0.415/0.487 | 0.813/0.758 |
| Ours | 0.34B | 1.3 GB | 6.37 | 0.755/0.794 | 0.496/0.542 | 0.729/0.747 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Feature | Type | Description |
|---|---|---|
| SI (Spatial Information) | Numeric | Characterizes the amount of detail and texture in the scene. |
| Blockiness Wu and Yuen (1997) | Numeric | Evaluates how pronounced the artificial boundaries between adjacent image blocks are. |
| SAM Kirillov et al. (2023) | Numeric | Number of objects detected by the model in the image. |
| Q-Align | Numeric | Expresses not only the visual quality of the image but also its aesthetic component, as it was trained on both tasks. |
| ARNIQA Agnolucci et al. (2024a) | Embedding | Reflects information about image quality and visual distortions. |
| CLIP Radford et al. (2021) | Embedding | Expresses semantic information about the image. |