Salvation Lies Within: Eliciting Inherent Style Transfer in Step-Distilled Diffusion Models
Organizations: Huawei Technologies · Hong Kong Polytechnic University · City University of Hong Kong
Abstract
Adapting step-distilled text-to-image (T2I) models through post-training incurs additional computational costs and affects native few-step generation behavior. This motivates a complementary route beyond style-specific adaptation: drawing on the visual knowledge already encoded in step-distilled T2I models to elicit stylistic capabilities through language. Pursuing this direction requires textual guidance that captures how visual attributes jointly define a style and remain applicable as the depicted content changes. To explore this approach, we introduce StyleForge, a fully automatic, training-free framework that expresses reference styles as reusable rendering instructions. By integrating overall rendering characteristics with local color and lighting behavior, StyleForge organizes visual evidence from reference images into a coherent specification of how the target style should be expressed. The specification is then compiled into textual guidance that can be reused across content prompts, enabling frozen step-distilled T2I models to render different subjects and scenes in the reference style while retaining native few-step generation. Extensive experiments show relative gains of up to 29.47% in generation quality scores over the strongest baseline, while Pareto analysis indicates that improved stylization is accompanied by strong adherence to the requested content.
Figures & tables
| Category | Description | Transfer Policy |
| Rendering | Medium, Surface, Line, Shape, Texture, Detail Density, etc. | Preserve ✓ |
| Color and Light | Carrier Surfaces, Extent, Distribution of Colors, Light, etc. | Preserve ✓ |
| Composition | Layout, Depth, Negative Space, etc. | Weaken |
| Reference Content | Subjects, Props, Scenes, Enclosing Structures, etc. | Exclude ✗ |
| Method | Qwen3-VL-8B-Instruct | Qwen3-VL-32B-Instruct | Qwen3.5-9B | Avg. Rank | |||
| Score | Win Rate (%) | Score | Win Rate (%) | Score | Win Rate (%) | ||
| FLUX.2-Klein-4B Backbone | |||||||
| DSE | 3.226 | 52.67% | 3.187 | 46.75% | 2.077 | 46.06% | 3.667 |
| RCG | 2.689 | 41.84% | 2.950 | 47.53% | 1.698 | 44.12% | 4.500 |
| Vanilla SFT | 3.769 | 40.54% | 3.300 | 29.43% | 1.506 | 34.86% | 4.167 |
| PSO | 3.757 | 37.82% | 3.255 | 27.49% | 1.496 | 32.65% | 5.167 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.