cs.CVSep 29, 2026

STEPS: Scene Text Editing with Preserved Style Using Diffusion and Contrastive Style Encoding

Authors: Nicolas Thiebaut, Nameer Hirschkind, Xiao Yu, Kyle Spence

Organizations: Roblox San Mateo, CA, USA

Abstract

We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation, background, etc. STEPS advances the state of the art in STE through directed focus on improved style preservation. We introduce a style encoder for visual text that captures style independently of textual content, and a model architecture that combines the style encoder with multiple semantic conditions (target text characters encoding and rendered glyphs). STEPS achieves superior results to previous STE methods in style preservation, output readability, and subjective quality.

Figures & tables

Explore similar work

CardsList
  1. Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning

    May 15, 2026Hongxi Li, Tong Wang, Chengjing Wu +6Diffusion-Based Image EditingText-To-Image Diffusion Models

  2. Diffusion Image Editing via Asynchronous Token Decoding

    Aug 10, 2026Yang Shi, Liangsi Lu, Minzhe Guo +4Diffusion-Based Image EditingText-To-Image Diffusion Models