We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation, background, etc. STEPS advances the state of the art in STE through directed focus on improved style preservation. We introduce a style encoder for visual text that captures style independently of textual content, and a model architecture that combines the style encoder with multiple semantic conditions (target text characters encoding and rendered glyphs). STEPS achieves superior results to previous STE methods in style preservation, output readability, and subjective quality.
Figures & tables
Figure 1: Examples of Scene Text Editing with STEPS applied to natural language translation of visual text in images with arbitrary backgrounds, fonts, and styles. (Top) Input images. (Bottom) translations of visual text into Indonesian (first three) and Italian (last three). STEPS is able to preserve the original text style, plausibly inpaint the disoccluded background, and to accommodate target texts with different lengths than the original.
Figure 2: STEPS multi-conditions cross-attention module. From the source text patch and desired rendered text on the left, we compute four conditioning input embeddings (Style, OCR, Char, and Prompt). Each cross-attention layer of the U-Net is replaced by the sum of four corresponding cross-attention modules. The cross-attention modules share the same query vector Q and their outputs are simply summed (details in the main text).
Figure 3: Style Encoder Training. To learn style embeddings (font, color, orientation, background) without encoding the textual content, we observe that stylistic similarity is correlated with spatial proximity in images that contain visual text. We then sample visual text patches from either different images (style similarity of zero) or the same image (style similarity of one minus bounding boxes distance). Finally, we train a bi-encoder on those stylistic similarity scores. A single copy of the resulting style encoder provides useful embeddings that we use as conditioning input for the main diffusion model.
Figure 4: Samples comparison (best viewed zoomed in). In some of those examples, MOSTEL fails to properly erase the original text, while TextDiffuser only partially matches the style of the input text. TextCtrl preserves style best for some examples, but sometimes suffers from blurriness or fails to produce the correct letters. Our new STEPS method offers the best readability and style-preservation properties.
Accuracy @ 1
Accuracy @ 10
SigLIP2-base
0.1793
0.5354
STEPS Style Encoder
0.5880
0.8882
Table 1: Style encoder accuracy on the ScenePair retrieval task. To validate the training method of our style encoder, we treat pair retrieval within all 1,280 examples as a classification task. We perform nearest neighbor classification in the embedding space of the SigLIP2-base image encoder (86M parameters) and our style encoder (22M parameters). While being smaller, our model offers significantly higher accuracy, confirming that it captures style similarity more accurately than a generic image encoder such as SigLIP2.
Figure 5: Style Encoder example for illustration purposes. The leftmost image is a sample text patch (top: original, bottom: preprocessed input to the encoding model), for which we computed the cosine similarity with all text patches in a sample of 1,000 images. The top row contains the most similar patches within this sample according to our style encoder, while the bottom row shows the least similar images. Text patches with similar colors, fonts, and backgrounds have aligned embeddings, while low embedding alignment correlates with text style discrepancies.
Original
Indonesian Translation
Italian Translation
LPIPS ↓
FID ↓
OCR ↑
LPIPS ↓
FID ↓
OCR ↑
LPIPS ↓
FID ↓
OCR ↑
MOSTEL
0.063
16.119
0.681
0.072
18.445
0.506
0.071
18.179
0.536
TextDiffuser2
0.080
16.095
0.690
0.099
19.226
0.651
0.097
18.879
0.653
TextCtrl
0.097
22.975
0.623
0.104
23.336
0.619
0.104
23.786
0.636
STEPS (new)
0.045
9.116
0.708
0.070
12.392
0.686
0.069
11.984
0.691
Table 2: Quantitative results on the AnyWord-3M-LAION benchmark dataset. Original refers to repainting the original image after masking text boxes, while Indonesian and Italian translations were selected because their alphabet is a subset of the English alphabet on which the models considered were trained. The translation task notably measures the ability to edit text that deviates from the original in length. Higher OCR indicates better readability, lower LPIPS better style preservation, and lower FID better output realism. Both MOSTEL and TextCtrl operate at the patch level. On the translation task, TextCtrl struggles with style consistency, while MOSTEL achieves strong style preservation—though often at the expense of readability, as it sometimes fails to fully erase the original text.
Indonesian Translation
Italian Translation
Visual Quality (0–5) ↑
Legibility (0–5) ↑
Visual Quality (0–5) ↑
Legibility (0–5) ↑
MOSTEL
1.52
1.38
1.84
1.68
TextDiffuser2
2.46
2.83
2.56
3.01
TextCtrl
1.496
1.536
1.424
1.472
STEPS (new)
3.224
3.624
3.544
3.936
Table 3: Vision-Language Model evaluation. In this setting of scene text translation with multiple text boxes per image, methods that operate at the text patch-level struggle with visual quality and legibility, from the perspective of the VLM evaluator. This is partly due to the blending post-processing step that is necessary to edit the entire input image. Inpainting methods like TextDiffuser2 and STEPS do not suffer from this limitation, and STEPS outperforms competing methods in both visual quality and legibility.
Original
Indonesian Translation
Italian Translation
LPIPS ↓
FID ↓
OCR ↑
LPIPS ↓
FID ↓
OCR ↑
LPIPS ↓
FID ↓
OCR ↑
STEPS (all conditions)
0.045
9.116
0.708
0.070
12.392
0.686
0.069
11.984
0.691
W/o style encoder
0.056
11.501
0.701
0.083
15.416
0.676
0.080
14.848
0.672
Table 4: Quantitative results from the ablation study. Higher OCR values indicate improved readability, while lower LPIPS scores reflect better style preservation, and lower FID scores suggest higher output realism. Removing the style encoder from the conditioning inputs negatively affects style preservation and realism, but has a more limited impact on readability.
Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely solely on image background information while neglecting the visual details of target regions, which discards stylistic features in the original text and essentially degrades the task to text rendering. Moreover, the conditions imposed by pre-trained glyph encoder limit the scope of editable text. To address these issues, this paper proposes a self-prompting scene text editing method that constructs style and glyph prompts directly from the original image, without introducing additional style or glyph encoders. We employ a two-stage training strategy: the diffusion transformer is first trained on large-scale self-supervised data and then refined using a small set of paired images. By leveraging the in-context learning capability of the Multi-Modal Diffusion Transformer (MM-DiT), it achieves open-vocabulary and style-consistent text editing. Experimental results on various languages demonstrate that our method achieves the state-of-the-art performance in both text accuracy and style consistency. Our project page: hongxiii.github.io/mstedit.
Hongxi Li, Tong Wang, Chengjing Wu +6
1MT Lab, Meitu Inc., Beijing, China · School of Computer Science & Technology, Beijing Institute of Technology, Beijing, China
We present StyleText, a large-scale dataset and benchmark for localized scene-text inpainting with style preservation. StyleText contains 28,518 image-mask-prompt triplets grouped into 9,932 scene families, enabling controlled evaluation of text legibility and visual consistency under shared scene context. We construct the dataset with an automated pipeline that combines LLM prompt templating, Flux-based source generation with key-value (KV) cache injection, OCR-based semantic filtering, polygon mask extraction, and mask-conditioned FluxFill augmentation. We define a reproducible evaluation protocol using normalized OCR metrics (word accuracy and character error rate) and CLIP image-image similarity with explicit preprocessing. A FluxFill+LoRA baseline trained on StyleText improves OCR accuracy substantially over initialization while maintaining scene style consistency, establishing a strong reference point for future comparisons.
Text-guided diffusion image editing aims to modify semantic attributes of an image while preserving its identity, layout, and background. However, naïvely switching the text condition during sampling often causes global drift, as denoising dynamics propagate changes across tokens and can disrupt unedited regions. To address this issue, we propose \textbf{A}synchronous \textbf{T}oken \textbf{D}ecoding \textbf{Edit} (ATDEdit), an inference-time framework that views each sampler step as a parallel update of a globally coupled token matrix and enables token-indexed condition switching with differentiated update policies. Instead of applying synchronous target-conditioned updates to all tokens, ATDEdit estimates editable locations using token-wise conditional surprisal and applies target-conditioned corrections to the selected token set. It supplies source key/value memory at keep-token positions and projects selected keep-token latent rows back to their source values; these operations promote background preservation but do not constitute a pixel-level invariance guarantee. This approach combines local editing and background preservation without external or user-provided spatial masks and without model fine-tuning. On PIE-Bench, ATDEdit achieves the strongest reported preservation metrics, including 27.44~dB PSNR and 0.055 LPIPS, while retaining competitive semantic alignment.
Yang Shi, Liangsi Lu, Minzhe Guo +4
Guangdong University of Technology Guangzhou, China · Hong Kong Baptist University Hong Kong, China · Peking University Beijing, China +1