Recent real-world image super-resolution (SR) methods often adapt text-to-image (T2I) backbones with ControlNet-style branches or spatial conditioning tokens, which increases memory and computes with resolution and often constrains training to a fixed scale. We propose Fill2SR, which repurposes a masked-inpainting Diffusion Transformer for SR without extra spatial branches. Our Inpainting-Interface Evidence Adapter (IIEA) writes the low-quality (LQ) observation into the native masked-image slot under a full-image mask, turning inpainting into a reverse-degradation conditional rectified flow trained with LoRA-only tuning. We further introduce RCDT, an offline pipeline that distills degradation descriptors from unpaired real images and transfers them onto clean targets using frozen open-source models. Fill2SR supports mixed-resolution training up to QHD and yields stable performance across 512/1024/2048 outputs. On synthetic benchmarks, our base model with IIEA achieves the best LPIPS on DIV2K and LSDIR; adding RCDT trades a small LPIPS drop for consistently stronger no-reference quality on RealLQ250 and RealPhoto60. Fill2SR remains memory-predictable, running 15362 inference on a single 32GB GPU and extending to multi-megapixel outputs via tiled restoration.
Figures & tables
Figure 1 : Visual highlights on RealLQ250 at 8× (real-world SR). Fill2SR produces artifact-suppressed, visually coherent high-resolution results across diverse content: non-rigid textures (left, Wukong), rigid urban structures (middle, Cityscapes), and fine biological details (right, Panda). Bottom: zoom-ins from the red boxes.
Figure 2 : Overview of the Fill2SR architecture. (a) The Main Restoration Pipeline utilizes the proposed IIEA to inject VAE-encoded LQ evidence into the native FLUX-Fill inpainting interface via a full-image mask ( m=1 ). Sampling follows a conditional rectified flow trajectory from pure noise z1 to the restored HQ latent z0 . (b) The Offline RCDT Pipeline leverages multi-modal agents (e.g., Qwen3-VL) to perceive real-world degradations and generate refined specifications, driving an instruction-based editor to synthesize realistic training pairs {x,y^} with zero test-time overhead.
Figure 3 : Qualitative comparisons on real-world scenarios at up to 8x degradation. Fill2SR produces visually plausible restorations across diverse and challenging domains. Row 1 (Industrial Text): Fill2SR can improve text legibility and boundary sharpness, while some baselines introduce distracting structured artifacts. Row 2 (Natural Textures): Fill2SR tends to recover fine textures (e.g., frost-like micro-structures) without excessive over-smoothing. Row 3 (Human Faces): Fill2SR often preserves facial structure and eye highlights, while some baselines exhibit local geometric drift. The left split-screen intuitively highlights our global visual enhancement. More results with additional baselines are provided in the Supplementary Material.
Figure 4 : Qualitative ablation on representative real-world inputs (RealLQ250). Columns show the LQ input, latent refinement ( β=0.5 ), w/o RCDT, w/o Redux, weakened evidence injection ( α=0.5 ), and the default Fill2SR output. Top: industrial text. Bottom: natural texture.