Adapting step-distilled text-to-image (T2I) models through post-training incurs additional computational costs and affects native few-step generation behavior. This motivates a complementary route beyond style-specific adaptation: drawing on the visual knowledge already encoded in step-distilled T2I models to elicit stylistic capabilities through language. Pursuing this direction requires textual guidance that captures how visual attributes jointly define a style and remain applicable as the depicted content changes. To explore this approach, we introduce StyleForge, a fully automatic, training-free framework that expresses reference styles as reusable rendering instructions. By integrating overall rendering characteristics with local color and lighting behavior, StyleForge organizes visual evidence from reference images into a coherent specification of how the target style should be expressed. The specification is then compiled into textual guidance that can be reused across content prompts, enabling frozen step-distilled T2I models to render different subjects and scenes in the reference style while retaining native few-step generation. Extensive experiments show relative gains of up to 29.47% in generation quality scores over the strongest baseline, while Pareto analysis indicates that improved stylization is accompanied by strong adherence to the requested content.
Figures & tables
Figure 1: Precise textual control of rendering attributes.
Figure 2: Overview of StyleForge. (1) The agent extracts structured style evidence from reference images. (2) Color and luminance statistics ground the extracted evidence. (3) The agent compiles a reusable style prompt and combines it with new content prompts to guide a step-distilled T2I model.
Category
Description
Transfer Policy
Rendering MR
Medium, Surface, Line, Shape, Texture, Detail Density, etc.
Preserve ✓
Color and Light AR
Carrier Surfaces, Extent, Distribution of Colors, Light, etc.
Preserve ✓
Composition TR
Layout, Depth, Negative Space, etc.
Weaken ↓
Reference Content LR
Subjects, Props, Scenes, Enclosing Structures, etc.
Exclude ✗
Table 1: Transfer policies for the four types of information in the style state SR .
Figure 3: Qualitative comparison with baselines on FLUX.2-Klein-4B.
Figure 4: Qualitative comparison with baselines on FLUX.2-Klein-9B.
Method
Qwen3-VL-8B-Instruct
Qwen3-VL-32B-Instruct
Qwen3.5-9B
Avg. Rank
Score
Win Rate (%)
Score
Win Rate (%)
Score
Win Rate (%)
FLUX.2-Klein-4B Backbone
DSE
3.226
52.67%
3.187
46.75%
2.077
46.06%
3.667
RCG
2.689
41.84%
2.950
47.53%
1.698
44.12%
4.500
Vanilla SFT
3.769
40.54%
3.300
29.43%
1.506
34.86%
4.167
PSO
3.757
37.82%
3.255
27.49%
1.496
32.65%
5.167
Table 2: VLM evaluation of content and style transfer on FLUX.2-Klein backbones ( Best , Second Best ). Gain is relative to the strongest baseline.
Figure 5: Pareto comparison of content and style across FLUX.2-Klein backbones.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Additional qualitative comparison with baselines on FLUX.2-Klein-4B.
Figure 7: Additional qualitative comparison with baselines on FLUX.2-Klein-9B.
Figure 8: Supplementary qualitative comparison with i2L on the FLUX.2-Klein-4B series. Each group shows style references, StyleForge using the step-distilled model, and i2L using the base model with CFG scales of 1 and 4.
Figure 9: Qualitative comparison with baselines on Z-Image-Turbo 6B.
Figure 10: Qualitative style-transfer results with GPT-Image-2.5-Flare. The first column shows style references, and the remaining columns show generated images for different content prompts.
Figure 12: Qualitative illustration of subject driven generation with a step-distilled model. Three reference subjects, a can, a toy car, and a berry bowl, are placed in a new cinematic scene. The reference-conditioned result preserves their colors, graphics, shapes, and lettering, while the text-only sample uses the same prompt and seed without access to the reference images.
Figure 13: Additional Pareto comparisons using Qwen3-VL-8B-Instruct and Qwen3.5-9B as judges. The two left panels show the FLUX.2-Klein-4B and 9B results for Qwen3-VL-8B-Instruct, while the two right panels show the corresponding results for Qwen3.5-9B.
We introduce Scene Text Editing with Preserved Style (STEPS), a novel diffusion model architecture for quality text replacement in images. Scene Text Editing (STE), also known as Visual Text Editing, consists of changing the textual content in an image while conserving the original style, e.g. font, colors, orientation, background, etc. STEPS advances the state of the art in STE through directed focus on improved style preservation. We introduce a style encoder for visual text that captures style independently of textual content, and a model architecture that combines the style encoder with multiple semantic conditions (target text characters encoding and rendered glyphs). STEPS achieves superior results to previous STE methods in style preservation, output readability, and subjective quality.
Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate states provide distillation supervision. Based on this insight, we propose REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage co-training framework in which a decoupled student learns from the evolving RL teacher's trajectories without changing teacher optimization. Advantage-Modulated Distillation (AMD) transforms rollout advantages into signed weights, strengthening imitation of preferred trajectories and aligning distillation priorities with task value. The resulting framework is general and lightweight, requires no extra image rollouts, no separate distillation dataset, and no adversarial training. Experiments on compositional generation, visual text rendering, and human-preference alignment demonstrate competitive few-step, CFG-free generation with RAM or DiffusionNFT teachers. With only four sampling steps, REST-RAM achieves a DrawBench PickScore of 23.97, outperforming both the 40-step RAM teacher (23.95) and RTDMD (23.71).
Yuhan Li, Fangao Zeng, Sicong Kang +7
Shanghai Jiao Tong University · Taobao & Tmall group of Alibaba
Few-step distillation has emerged as a critical component in the development of advanced visual generative foundation models, substantially reducing inference overhead while enabling real-time generation and cost-efficient deployment across a broad range of practical scenarios. However, prior work has predominantly focused on advancing training objectives, while comparatively overlooking the training recipe, which has become increasingly critical in the era of large-scale foundation models. In this work, we systematically revisit the training recipe under the well-established distribution matching distillation (DMD) framework for both text-to-image generation and image editing, focusing on three key dimensions: training data composition, teacher guidance within DMD, and task mixture. Our empirical analysis reveals several non-obvious and counterintuitive phenomena, ultimately motivating the development of Qwen-Image-Flash. These findings highlight that effective few-step distillation depends not only on carefully designed objectives, but also on a principled training recipe.