We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation, background, etc. STEPS advances the state of the art in STE through directed focus on improved style preservation. We introduce a style encoder for visual text that captures style independently of textual content, and a model architecture that combines the style encoder with multiple semantic conditions (target text characters encoding and rendered glyphs). STEPS achieves superior results to previous STE methods in style preservation, output readability, and subjective quality.
Figures & tables
Figure 1: Examples of Scene Text Editing with STEPS applied to natural language translation of visual text in images with arbitrary backgrounds, fonts, and styles. (Top) Input images. (Bottom) translations of visual text into Indonesian (first three) and Italian (last three). STEPS is able to preserve the original text style, plausibly inpaint the disoccluded background, and to accommodate target texts with different lengths than the original.
Figure 2: STEPS multi-conditions cross-attention module. From the source text patch and desired rendered text on the left, we compute four conditioning input embeddings (Style, OCR, Char, and Prompt). Each cross-attention layer of the U-Net is replaced by the sum of four corresponding cross-attention modules. The cross-attention modules share the same query vector Q and their outputs are simply summed (details in the main text).
Figure 3: Style Encoder Training. To learn style embeddings (font, color, orientation, background) without encoding the textual content, we observe that stylistic similarity is correlated with spatial proximity in images that contain visual text. We then sample visual text patches from either different images (style similarity of zero) or the same image (style similarity of one minus bounding boxes distance). Finally, we train a bi-encoder on those stylistic similarity scores. A single copy of the resulting style encoder provides useful embeddings that we use as conditioning input for the main diffusion model.
Figure 4: Samples comparison (best viewed zoomed in). In some of those examples, MOSTEL fails to properly erase the original text, while TextDiffuser only partially matches the style of the input text. TextCtrl preserves style best for some examples, but sometimes suffers from blurriness or fails to produce the correct letters. Our new STEPS method offers the best readability and style-preservation properties.
Accuracy @ 1
Accuracy @ 10
SigLIP2-base
0.1793
0.5354
STEPS Style Encoder
0.5880
0.8882
Table 1: Style encoder accuracy on the ScenePair retrieval task. To validate the training method of our style encoder, we treat pair retrieval within all 1,280 examples as a classification task. We perform nearest neighbor classification in the embedding space of the SigLIP2-base image encoder (86M parameters) and our style encoder (22M parameters). While being smaller, our model offers significantly higher accuracy, confirming that it captures style similarity more accurately than a generic image encoder such as SigLIP2.
Figure 5: Style Encoder example for illustration purposes. The leftmost image is a sample text patch (top: original, bottom: preprocessed input to the encoding model), for which we computed the cosine similarity with all text patches in a sample of 1,000 images. The top row contains the most similar patches within this sample according to our style encoder, while the bottom row shows the least similar images. Text patches with similar colors, fonts, and backgrounds have aligned embeddings, while low embedding alignment correlates with text style discrepancies.
Original
Indonesian Translation
Italian Translation
LPIPS ↓
FID ↓
OCR ↑
LPIPS ↓
FID ↓
OCR ↑
LPIPS ↓
FID ↓
OCR ↑
MOSTEL
0.063
16.119
0.681
0.072
18.445
0.506
0.071
18.179
0.536
TextDiffuser2
0.080
16.095
0.690
0.099
19.226
0.651
0.097
18.879
0.653
TextCtrl
0.097
22.975
0.623
0.104
23.336
0.619
0.104
23.786
0.636
STEPS (new)
0.045
9.116
0.708
0.070
12.392
0.686
0.069
11.984
0.691
Table 2: Quantitative results on the AnyWord-3M-LAION benchmark dataset. Original refers to repainting the original image after masking text boxes, while Indonesian and Italian translations were selected because their alphabet is a subset of the English alphabet on which the models considered were trained. The translation task notably measures the ability to edit text that deviates from the original in length. Higher OCR indicates better readability, lower LPIPS better style preservation, and lower FID better output realism. Both MOSTEL and TextCtrl operate at the patch level. On the translation task, TextCtrl struggles with style consistency, while MOSTEL achieves strong style preservation—though often at the expense of readability, as it sometimes fails to fully erase the original text.
Indonesian Translation
Italian Translation
Visual Quality (0–5) ↑
Legibility (0–5) ↑
Visual Quality (0–5) ↑
Legibility (0–5) ↑
MOSTEL
1.52
1.38
1.84
1.68
TextDiffuser2
2.46
2.83
2.56
3.01
TextCtrl
1.496
1.536
1.424
1.472
STEPS (new)
3.224
3.624
3.544
3.936
Table 3: Vision-Language Model evaluation. In this setting of scene text translation with multiple text boxes per image, methods that operate at the text patch-level struggle with visual quality and legibility, from the perspective of the VLM evaluator. This is partly due to the blending post-processing step that is necessary to edit the entire input image. Inpainting methods like TextDiffuser2 and STEPS do not suffer from this limitation, and STEPS outperforms competing methods in both visual quality and legibility.
Original
Indonesian Translation
Italian Translation
LPIPS ↓
FID ↓
OCR ↑
LPIPS ↓
FID ↓
OCR ↑
LPIPS ↓
FID ↓
OCR ↑
STEPS (all conditions)
0.045
9.116
0.708
0.070
12.392
0.686
0.069
11.984
0.691
W/o style encoder
0.056
11.501
0.701
0.083
15.416
0.676
0.080
14.848
0.672
Table 4: Quantitative results from the ablation study. Higher OCR values indicate improved readability, while lower LPIPS scores reflect better style preservation, and lower FID scores suggest higher output realism. Removing the style encoder from the conditioning inputs negatively affects style preservation and realism, but has a more limited impact on readability.