cs.CVOct 6, 2026

DISRQAD: Diffusion Image Super-Resolution Quality Assessment Dataset and Benchmark

Authors: Nikita Kukuzei, Artem Borisov, Evgeney Bogatyrev, Khaled Abud, Egor Chistov, Dmitriy Vatolin

Organizations: Lomonosov Moscow State University Moscow, Russia · MSU Institute for AI, Lomonosov Moscow State University Moscow, Russia

Abstract

Diffusion-based image super-resolution (SR) can create visually plausible detail that is not supported by the low-resolution input. We introduce DISRQAD, a subjective-quality dataset and diagnostic benchmark for this setting. It contains mean opinion scores (MOS) for 14,000 SR outputs from ten diffusion and four non-diffusion methods, spanning four low-resolution degradation conditions and x2/x4 upscaling. We evaluate 51 standard full-reference and no-reference metric configurations and 11 adapted variants. Agreement with MOS is substantially weaker on diffusion outputs: the strongest standard no-reference baseline reaches 0.431 SRCC on diffusion SR versus 0.813 on non-diffusion SR. As a case study in benchmark use, a pruned and distilled Q-ReAlign-mini student reaches 0.496 SRCC on diffusion SR. DISRQAD measures perceived output quality, not faithfulness to the input; it enables analysis of metric behavior across generator families and input conditions. Our findings reveal a substantial gap in the assessment of diffusion-based SR and provide a basis for developing quality models sensitive to diffusion-specific artifacts.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 8, 2026cs.CV

TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolution

Diffusion-based generative models have achieved remarkable success in real-world image super-resolution (SR). With tiled diffusion techniques, these models can produce high-resolution images that exceed their native-supported resolution. However, the quality of such high-resolution (e.g 204822048^2) outputs often remains extremely poor, primarily due to two factors we consider: the image upsampling ratio (e.g ×8\times8) exceeding the model's native-supported upsampling ratio (e.g ×4\times4), and the model's native-supported resolution. In practice, training a native high-resolution model requires larger architectures, which incur significant computational overhead and GPU memory costs, making it hard on limited-resource equipment. Thus, we present TUDSR, a Twice Upsampling-Diffusion framework for higher SR. The TUDSR framework mainly consists of two stages: the first involves training at RR-resolution, and the second introduces a looped chunk-based training strategy at NRNR-resolution. Each stage adapts a one-step GAN architecture comprising a generator and a discriminator. Based on SD2.1-base, we develop TUDSR-S, which achieves state-of-the-art performance across multiple benchmarks. Extensive experiments further demonstrate that TUDSR-S generates high-quality images at the resolutions of 102421024^2 and even 204822048^2, significantly outperforming existing approaches. Code is available at https://github.com/wuer5/TUDSR.
May 25, 2026eess.IV

How Accurate are Video Quality Models for Diffusion-Based Video Super-Resolution?

Recent video super-resolution (VSR) approaches use deep neural networks to enhance low-quality input videos and recover visual detail, with diffusion-based methods in particular showing promising results. In this paper, we investigate whether existing video quality models can be used to assess the performance of these diffusion-based VSR methods, by comparing model predictions with results from a subjective test. The study compares six upscaling methods (Lanczos, Rhea, SCST, DOVE, SeedVR2, Starlight Mini) applied to both compressed (AV1 and DCVC-RT) and uncompressed low-resolution videos considering the play-out on a UHD-1/4K screen. A range of full- and no-reference quality models are used to assess their applicability to this new type of quality degradation, focusing on within-sequence performance. The results highlight that CNN-based full-reference models, such as LPIPS, DISTS, and CVQA-FR show significantly higher correlation coefficients than both conventional full- as well as the tested no-reference models. Most overestimate the overly sharp results of SCST, with VMAF mainly failing due to spatial inconsistencies introduced by Starlight Mini. None of the tested video quality models reach sufficient accuracy so as to replace complementary subjective testing. The reference, degraded and upscaled videos, as well as the user ratings and model scores are made available with the paper at https://github.com/Telecommunication-Telemedia-Assessment/AVT-VQDB-UHD-1-VSR as open data.
May 20, 2026cs.CV

SR-Ground: Image Quality Grounding for Super-Resolved Content

Super-Resolution (SR) has advanced rapidly in recent years, with diffusion-based models achieving unprecedented fidelity at the cost of introducing new types of visual artifacts. While existing Image Quality Assessment (IQA) methods provide holistic quality scores, they lack interpretability and fail to distinguish between different artifact types arising from modern SR approaches. To address this gap, we introduce SR-Ground, a large-scale dataset specifically designed for fine-grained artifact segmentation in super-resolved images. The dataset comprises images processed by a diverse set of state-of-the-art SR models, with pixel-level annotations for multiple artifact categories. We conduct a large-scale crowdsourcing study involving 1,062 participants to validate and refine automatically generated segmentations, resulting in a highquality dataset of 63,000 images spanning 6 distinct artifact types. We demonstrate that training IQA models with grounding capabilities on SR-Ground significantly improves performance on downstream tasks. Furthermore, we introduce a fine-tuning pipeline that leverages our grounding model to reduce perceptible artifacts in SR outputs, showcasing the practical utility of our dataset. On a separate benchmark of 1,000 outputs from five unseen SR methods, an Low-Resolution-referenced grounding model significantly improves prominence alignment over its no-reference counterpart and performs best on native real-world datasets of Low-Resolution images without High-Resolution Ground-Truth.