Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to reinforce each other and progressively amplify an early recognition error into a sharp yet semantically incorrect character. To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion. TOLA consists of two key modules. First, a confidence-weighted text conditioning module constructs the semantic condition only once and suppresses unreliable OCR predictions before they contaminate image reconstruction. Second, a lightweight latent residual correction module explicitly estimates and corrects the structured residual errors to recover missing or distorted stroke details. Extensive experiments demonstrate our state-of-the-art performance across all evaluation metrics on both CTR-TSR-Test (×4) and RealCE-200 benchmarks. It is worth noting that our TOLA consistently surpasses existing diffusion-based TSR methods by at least 2.72 dB in PSNR on CTR-TSR-Test.
Figures & tables
Figure 1: Comparison of three different diffusion-based TSR inference paradigms. (a) DiffTSR performs 200-step iterative text-prior and image restoration. (b) PRISM uses one-step image reconstruction with 16-step text-prior recovery. (c) TOLA constructs the text condition once, evaluates IDM once, and applies a Latent Residual Correction (LRC) for clean-latent correction. Prior iterative text-image refinement may repeatedly propagate inaccurate semantic cues through multiple denoising stages, leading to accumulated recognition errors. In contrast, TOLA avoids error amplification by designing a one-step text-conditioned IDM.
Figure 2: TOLA architecture. The LR image is encoded and q-sampled at fixed t=999 , while confidence-weighted TransOCR tokens and latent features form one MoM condition. A LoRA-adapted IDM predicts noise once, and LRC corrects the clean latent before frozen VAE decoding. Snowflakes and flames denote frozen and trainable components, respectively.
Figure 3: Comparison on five selected CTR-TSR-Test ×4 images covering Chinese, English, and numeric text. Each column presents an aligned test sample, and the rows show the LR input, displayed baselines, Ours, and the HR reference.
Figure 4: RT50 comparison on Chinese, English, and numeric text under varied real degradations. Rows show LR inputs, displayed baselines, and Ours; examples cover blur, illumination, viewpoint, resolution, and background variations.
Figure 5: Quality–latency comparison on CTR-TSR-Test ×4 . Panels (a) and (b) plot PSNR and TransOCR ACC against synchronized batch-1 inference time measured on 100 fixed CTR inputs using one NVIDIA RTX PRO 6000.
Figure 6: Core component visualization. (a) LoRA enables one-step adaptation and LRC corrects local structure over frozen and single-module controls. (b) No MoM and direct-token conditioning cause character errors; LR and HR provide references.
Figure 7: Recognition diagnostics on CTR-TSR-Test ×4 : (a) ACC across confidence bins, (b) ACC across text categories, and (c) character-level edit counts.
Text image super-resolution (Text-SR) requires more than visually plausible detail synthesis: slight errors in stroke topology may alter character identity and break readability. Existing methods improve text fidelity with stronger recognition-based or generative priors, yet they still face two unresolved challenges under severe degradation: the text condition extracted from low-quality inputs can itself be unreliable, and a plausible global prior does not fully determine fine-grained stroke boundaries. We present PRISM, a single-step diffusion-based Text-SR framework that addresses these two challenges through Flow-Matching Prior Rectification (FMPR) and a Structure-guided Uncertainty-aware Residual Encoder (SURE). FMPR constructs a privileged training-time prior from paired low-quality/high-quality latents and learns a flow matching that transports degraded embeddings toward this restoration-oriented prior space, yielding more accurate and reliable global text guidance. SURE further predicts uncertainty-aware structural residuals to selectively absorb reliable local boundary evidence while suppressing ambiguous stroke cues. Together, these components enable explicit global prior rectification and local structure refinement within a single diffusion restoration pass. Experiments on both synthetic and real-world benchmarks show that PRISM achieves state-of-the-art performance with millisecond-level inference. Our dataset and code will be available at https://github.com/faithxuz/PRISM.
Scene text image super-resolution (STISR) aims to recover visually plausible appearance while preserving character semantics from degraded inputs. Existing STISR systems often rely on externally generated priors or separate image and text models, resulting in error propagation and costly multi-stage inference. We present DualTSR, a unified framework that formulates STISR as coupled continuous-discrete generation. Conditional flow matching restores continuous image latents, while absorbing-state discrete diffusion reconstructs text tokens. Both processes share a multimodal transformer backbone, allowing the evolving image and text states to interact throughout generation without an external OCR prior at inference. On CTR-TSR, DualTSR achieves the best FID, LPIPS, ACC, and NED among the compared methods at both X2 and X4. On an aligned RealCE subset, it obtains the best FID, ACC, and NED with competitive LPIPS. Compared with DiffTSR at X4, DualTSR improves ACC by 12.78 percentage points while reducing the parameter count from 1.23B to 203M and end-to-end latency from 13.3s to 132ms. These results establish DualTSR as an accurate and efficient method for STISR.
Axi Niu, Knag Zhang, Qingsen Yan +3
Northwestern Polytechnical University, Xi’an, China · Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea
Despite recent advances, single-image super-resolution (SR) remains challenging, especially in real-world scenarios with complex degradations. Diffusion-based SR methods, particularly those built on Stable Diffusion, leverage strong generative priors but commonly rely on text conditioning derived from semantic captioning. Such textual descriptions provide only high-level semantics and lack the spatially aligned visual information required for faithful restoration, leading to a representation gap between abstract semantics and spatially aligned visual details. To address this limitation, we propose GramSR, a one-step diffusion-based SR framework that replaces text conditioning with dense visual features extracted from the low-resolution input using a pre-trained DINOv3 encoder. GramSR adopts a three-stage LoRA architecture, where pixel-level, semantic-level, and texture-level LoRA modules are trained sequentially. The pixel-level module focuses on degradation removal using ℓ2 loss, the semantic-level module enhances perceptual details via LPIPS and CSD losses, and the texture-level module enforces feature correlation consistency through a Gram matrix loss computed from DINOv3 features. At inference, independent guidance scales enable flexible control over degradation removal, semantic enhancement, and texture preservation. Extensive experiments on standard SR benchmarks demonstrate that GramSR consistently outperforms existing one-step diffusion-based methods, achieving superior structural fidelity and texture realism. The code for this work is available at: https://github.com/aimagelab/GramSR.