Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to reinforce each other and progressively amplify an early recognition error into a sharp yet semantically incorrect character. To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion. TOLA consists of two key modules. First, a confidence-weighted text conditioning module constructs the semantic condition only once and suppresses unreliable OCR predictions before they contaminate image reconstruction. Second, a lightweight latent residual correction module explicitly estimates and corrects the structured residual errors to recover missing or distorted stroke details. Extensive experiments demonstrate our state-of-the-art performance across all evaluation metrics on both CTR-TSR-Test (×4) and RealCE-200 benchmarks. It is worth noting that our TOLA consistently surpasses existing diffusion-based TSR methods by at least 2.72 dB in PSNR on CTR-TSR-Test.
Figures & tables
Figure 1: Comparison of three different diffusion-based TSR inference paradigms. (a) DiffTSR performs 200-step iterative text-prior and image restoration. (b) PRISM uses one-step image reconstruction with 16-step text-prior recovery. (c) TOLA constructs the text condition once, evaluates IDM once, and applies a Latent Residual Correction (LRC) for clean-latent correction. Prior iterative text-image refinement may repeatedly propagate inaccurate semantic cues through multiple denoising stages, leading to accumulated recognition errors. In contrast, TOLA avoids error amplification by designing a one-step text-conditioned IDM.
Figure 2: TOLA architecture. The LR image is encoded and q-sampled at fixed t=999 , while confidence-weighted TransOCR tokens and latent features form one MoM condition. A LoRA-adapted IDM predicts noise once, and LRC corrects the clean latent before frozen VAE decoding. Snowflakes and flames denote frozen and trainable components, respectively.
Figure 3: Comparison on five selected CTR-TSR-Test ×4 images covering Chinese, English, and numeric text. Each column presents an aligned test sample, and the rows show the LR input, displayed baselines, Ours, and the HR reference.
Figure 4: RT50 comparison on Chinese, English, and numeric text under varied real degradations. Rows show LR inputs, displayed baselines, and Ours; examples cover blur, illumination, viewpoint, resolution, and background variations.
Figure 5: Quality–latency comparison on CTR-TSR-Test ×4 . Panels (a) and (b) plot PSNR and TransOCR ACC against synchronized batch-1 inference time measured on 100 fixed CTR inputs using one NVIDIA RTX PRO 6000.
Figure 6: Core component visualization. (a) LoRA enables one-step adaptation and LRC corrects local structure over frozen and single-module controls. (b) No MoM and direct-token conditioning cause character errors; LR and HR provide references.
Figure 7: Recognition diagnostics on CTR-TSR-Test ×4 : (a) ACC across confidence bins, (b) ACC across text categories, and (c) character-level edit counts.