Recent real-world image super-resolution (SR) methods often adapt text-to-image (T2I) backbones with ControlNet-style branches or spatial conditioning tokens, which increases memory and computes with resolution and often constrains training to a fixed scale. We propose Fill2SR, which repurposes a masked-inpainting Diffusion Transformer for SR without extra spatial branches. Our Inpainting-Interface Evidence Adapter (IIEA) writes the low-quality (LQ) observation into the native masked-image slot under a full-image mask, turning inpainting into a reverse-degradation conditional rectified flow trained with LoRA-only tuning. We further introduce RCDT, an offline pipeline that distills degradation descriptors from unpaired real images and transfers them onto clean targets using frozen open-source models. Fill2SR supports mixed-resolution training up to QHD and yields stable performance across 512/1024/2048 outputs. On synthetic benchmarks, our base model with IIEA achieves the best LPIPS on DIV2K and LSDIR; adding RCDT trades a small LPIPS drop for consistently stronger no-reference quality on RealLQ250 and RealPhoto60. Fill2SR remains memory-predictable, running 15362 inference on a single 32GB GPU and extending to multi-megapixel outputs via tiled restoration.
Figures & tables
Figure 1 : Visual highlights on RealLQ250 at 8× (real-world SR). Fill2SR produces artifact-suppressed, visually coherent high-resolution results across diverse content: non-rigid textures (left, Wukong), rigid urban structures (middle, Cityscapes), and fine biological details (right, Panda). Bottom: zoom-ins from the red boxes.
Figure 2 : Overview of the Fill2SR architecture. (a) The Main Restoration Pipeline utilizes the proposed IIEA to inject VAE-encoded LQ evidence into the native FLUX-Fill inpainting interface via a full-image mask ( m=1 ). Sampling follows a conditional rectified flow trajectory from pure noise z1 to the restored HQ latent z0 . (b) The Offline RCDT Pipeline leverages multi-modal agents (e.g., Qwen3-VL) to perceive real-world degradations and generate refined specifications, driving an instruction-based editor to synthesize realistic training pairs {x,y^} with zero test-time overhead.
Figure 3 : Qualitative comparisons on real-world scenarios at up to 8x degradation. Fill2SR produces visually plausible restorations across diverse and challenging domains. Row 1 (Industrial Text): Fill2SR can improve text legibility and boundary sharpness, while some baselines introduce distracting structured artifacts. Row 2 (Natural Textures): Fill2SR tends to recover fine textures (e.g., frost-like micro-structures) without excessive over-smoothing. Row 3 (Human Faces): Fill2SR often preserves facial structure and eye highlights, while some baselines exhibit local geometric drift. The left split-screen intuitively highlights our global visual enhancement. More results with additional baselines are provided in the Supplementary Material.
Figure 4 : Qualitative ablation on representative real-world inputs (RealLQ250). Columns show the LQ input, latent refinement ( β=0.5 ), w/o RCDT, w/o Redux, weakened evidence injection ( α=0.5 ), and the default Fill2SR output. Top: industrial text. Bottom: natural texture.
Real-world image super-resolution (Real-ISR) requires balancing structural fidelity to degraded observations with realistic detail synthesis. However, existing generative Real-ISR methods often rely on entangled conditioning mechanisms, leading to structural drift or semantically inconsistent details. To address this issue, we propose Visual In-Context Restoration (VICR), a Diffusion Transformer (DiT)-based framework that formulates Real-ISR as image completion. Specifically, we introduce a decoupled visual prior injection mechanism that derives local and global cues from the low-quality (LQ) image: local cues help recover image structures and support high-frequency detail synthesis, while global cues guide overall generation and promote semantic consistency. For ambiguous regions under severe degradation, VICR employs an inference-time agent to refine semantic prompts using visual evidence from the LQ input while keeping model parameters fixed. Experiments show that VICR achieves state-of-the-art performance across multiple Real-ISR benchmarks with only 127M trainable parameters.
Qichang Zhang, Hailong Wang, Baiang Li +4
Faculty of Science and Technology, University of Macau · 2Nullmax · 3Hefei University of Technology +1
We present In-Token Learning, an image restoration framework that adapts a pretrained diffusion transformer using conditional rectified flow matching. Clean targets paired with degraded inputs supervise transport from Gaussian noise to restored images. Spatially aligned degraded-image tokens are fused with evolving latent tokens along the channel dimension, preserving the image-token count at a given resolution. Direct Low-Quality Guidance (DLG) combines frozen degraded-image embeddings with a fixed task prompt through the native conditioning pathway, without a trainable ControlNet-style branch or image captioning. We evaluate super-resolution and denoising on DIV2K, LSDIR, FFHQ, RealLQ250, and RealPhoto60, and automatic colorization on DIV2K and LSDIR. The tasks use separately trained checkpoints under the same framework. Results show competitive fidelity and perceptual quality under the evaluated protocols, with weaker generalization on RealLQ250. We report full-image QHD (2560×1440) inference and a tiled 12K restoration demonstration of Along the River During the Qingming Festival. Attention cost still increases with resolution. This technical report preserves the early broader study underlying Fill2SR, which subsequently developed the real-world super-resolution direction.
Xingfu Yi, Xiaoxue Yu
Independent Researcher, Hangzhou, China · Zhejiang University, Hangzhou, China
Pre-trained text-to-image (T2I) diffusion models have shown strong potential for real-world image super-resolution (Real-ISR), owing to their noise-started generation process that enables realistic texture synthesis and captures the one-to-many nature of super-resolution. However, diffusion-based Real-ISR methods still face a fundamental efficiency-quality trade-off. Multi-step methods generate high-quality results by iteratively denoising random Gaussian noise under LR conditioning, but suffer from slow sampling. Recent one-step methods greatly improve efficiency, yet they typically replace noise-started generation with direct LR-to-HR restoration, which weakens stochasticity and limits realistic detail synthesis. To address this issue, we propose SMFSR, a noise-started one-step Real-ISR framework via LR-conditioned SplitMeanFlow and GAN refinement. SMFSR preserves the random-noise starting point of diffusion models and learns a direct noise-to-HR mapping conditioned on the LR image. To this end, Interval Splitting Consistency distills the multi-step generative trajectory into a single average-velocity prediction, enabling efficient one-step generation. To compensate for the reduced opportunity for progressive refinement, we further introduce a GAN refinement stage, where a DINOv3-based discriminator enhances realistic texture synthesis and variational score distillation aligns the generated outputs with the natural image distribution under a frozen diffusion teacher. Extensive experiments demonstrate that SMFSR achieves state-of-the-art perceptual quality among one-step diffusion-based Real-ISR methods while retaining fast single-step inference.
Wei Zhu, Kai Zhang, Yu Zheng +3
Nanjing University of Science and Technology · Nanjing University · Huawei