Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During training, a raw-prompt student follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field targets at the states visited by the student. This on-policy supervision distills the behavior induced by enhanced prompts into the raw-prompt student without requiring additional text--image pairs. At inference, both the PE and teacher are removed, and the student generates directly from raw prompts. Across multiple model families, PEs, and benchmarks, PE-OPSD achieves the strongest aggregate prompt fidelity among the evaluated baselines, yields positive aggregate visual appeal gains, and retains the base-model inference efficiency.
Figures & tables
Figure 1: Comparison of generation quality, throughput, and training dynamics. PE-OPSD achieves higher GenEval scores than both Vanilla and PE while preserving the inference efficiency. Moreover, PE-OPSD achieves faster convergence and greater performance gains than SFT and off-policy distillation. Training dynamics is reported in Z-Image-Turbo training.
Figure 2: Three different uses of prompt conditioning. (a) Direct generation from raw prompts results in limited prompt alignment; (b) Inference-time PE improves prompt alignment by rewriting raw prompts, but introduces additional computational overhead; (c) Our PE-OPSD internalizes the PE knowledge into the generator, improving alignment without additional inference cost.
Figure 3
Figure 4: Overview of PE-OPSD. At training, the student generates rollouts from raw prompts, while an EMA teacher provides enhanced-prompt supervision on the same rollout states. At inference, the trained student generates directly from raw prompts without inference-time PE.
Method
GenEval (GE) Task
GenEval2 (GE2) Task
GE
PickScore
CLIP
Aes.
GE2 GM
GE2 AM
PickScore
CLIP
Aes.
IPF
IVA
SD3.5-L
0.604
0.862
0.288
5.286
0.213
0.660
0.869
0.315
5.521
-
-
FLUX.1-dev
0.618
0.899
0.280
5.624
0.179
0.643
0.895
0.298
5.909
-
-
SD3.5-M (2.5B)
Base
0.628
0.878
0.289
5.319
0.176
0.633
0.882
0.313
5.579
0.00%
0.00%
Base+PE
0.743
0.892
0.291
5.439
0.220
0.680
0.887
0.312
5.678
10.22%
1.55%
Table 2: Main results across different models. We use GPT-5.6 Sol as PE. PickScore is normalized by 26; Aes. denotes aesthetics; Bold: best; Underlined: second-best.
Figure 5: Qualitative comparison and human preference study on Z-Image-Turbo. We use GPT-5.6 Sol as PE. Left: visual comparisons among Base, PE, SFT, Off-policy distillation, and our method. Right: pairwise human preferences against Base in prompt fidelity and visual appeal.
Method
GE
GE2 GM
GE2 AM
Lat. (s)
PromptEnhancer-7B as PE
Base
0.650
0.306
0.761
9.69
Base+PE
0.684
0.375
0.801
19.00
Base+Ours
0.782
0.422
0.819
9.69
Δ (vs +PE)
+0.098
+0.047
+0.018
1.96 ×
PromptEnhancer-32B as PE
Table 3: Performance and latency.
Method
GenEval (GE) Task
GenEval2 (GE2) Task
GE
PickScore
CLIP
Aes.
GE2 GM
GE2 AM
PickScore
CLIP
Aes.
IPF
IVA
GPT-5.6 Sol as PE
Z-Image
0.650
0.875
0.288
5.288
0.306
0.761
0.871
0.319
5.421
0.00%
0.00%
Z-Image+PE
0.810
0.899
0.299
5.433
0.404
0.827
0.893
0.326
5.638
14.27%
3.00%
SFT
0.821
0.893
0.302
5.426
0.403
0.834
0.887
0.326
5.466
14.93%
1.83%
Off-Policy Distillation
0.824
0.901
0.300
5.432
0.395
0.828
0.897
0.327
5.629
14.27%
3.12%
Table 4: Result across different PEs. Bold: best; Underlined: second-best.
Method
DPG-bench
T2I-CompBench++
EvalMuse
Overall ⋆
Glo.
Ent.
Attr.
Rela.
Other
Overall ⋆
Color
Shape
Tex.
Num.
Comp.
Spa.
3D Spa.
Non-spa.
Overall ⋆
SD3.5-M
83.9
84.6
89.6
88.1
93.0
80.9
0.51
0.80
0.54
0.74
0.59
0.37
0.32
0.36
0.31
3.199
+Ours
83.8
83.9
89.3
88.4
93.0
82.1
0.55
0.83
0.60
0.73
0.63
0.39
0.46
0.42
0.31
3.368
Z-Image
86.3
83.5
91.7
90.1
94.6
87.8
0.53
0.85
0.59
0.79
0.63
0.40
0.31
0.37
0.31
3.340
+Ours
87.0
82.6
92.0
90.2
94.5
89.4
0.58
0.87
0.61
0.81
0.71
0.42
0.36
0.41
0.32
3.561
Z-Image-Turbo
84.5
78.6
91.2
88.3
93.2
88.4
0.53
0.81
0.56
0.75
0.69
0.40
0.36
0.41
0.31
3.515
Table 5: Results on out-of-domain benchmarks. For DPG-bench, we report Global (Glo.), Entity (Ent.), Attribute (Attr.), Relation (Rela.), Other and Overall ⋆ . For T2I-CompBench++, we report Color, Shape, Textual (Tex.), Numeracy (Num.), Complex (Comp.), Spatial (Spa.), 3D Spatial (3D Spa.), Non-spatial (Non-spa.) and Overall ⋆ . Bold: best in Overall ⋆ .
Figure 6: Scaling to larger models. PickScore, CLIP, and Aesthetics are averaged over GenEval and GenEval2. For visualization, each metric is independently normalized to [0.3,1.0] .
Figure 7: Ablation study results on different training datasets. We use MixDataset by default.
Table 12
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Setting
Optimizer
AdamW
Learning rate
1×10−4 , constant without warm-up
Optimizer momentum
(β1,β2)=(0.9,0.999)
Adam ϵ / weight decay
10−8 / 0
Gradient clipping
1.0
Global batch size
64
Appendix
Table 9: Default training hyperparameters for PE-OPSD. All models follow the same recipe unless otherwise specified.
Model
Checkpoint
SD3.5-Medium
stabilityai/stable-diffusion-3.5-medium
SD3.5-Large
stabilityai/stable-diffusion-3.5-large
Z-Image
Tongyi-MAI/Z-Image
Z-Image-Turbo
Tongyi-MAI/Z-Image-Turbo
FLUX.1-dev
black-forest-labs/FLUX.1-dev
FLUX.2-klein-base-9B
black-forest-labs/FLUX.2-klein-base-9B
Appendix
Table 10: Models and checkpoints used in our experiments.
Model
SFT
Off-policy
PE-OPSD
One-step
Full †
One-step
Full
One-step
Full
SD3.5-M
1.47 s
0.4 + 6.6 h
11.02 s
3.1 h
11.02 s
3.1 h
Z-Image
3.14 s
0.9 + 27.5 h
49.15 s
13.6 h
49.15 s
13.6 h
Z-Image-Turbo
3.14 s
0.9 + 2.5 h
14.31 s
4.0 h
14.31 s
4.0 h
Appendix
Table 11: Training efficiency in our experiments. One-step denotes the wall-clock time per optimization step under the default configuration, and Full denotes the cost of 1,000 optimization steps. For SFT, Full † is reported as training time + offline pseudo-target generation time.
Model
Latency (s)
Base
+PE
+Ours
Speedup
PromptEnhancer-7B as PE
SD3.5-M
2.70
10.36
2.70
3.84 ×
Z-Image
9.69
19.00
9.69
1.96 ×
Z-Image-Turbo
0.88
8.53
0.88
9.69 ×
FLUX.2-klein
0.67
8.18
0.67
12.21 ×
Appendix
Table 12: Inference latency and speedup across different settings. Latency is the average generation time per image in seconds. +PE includes both prompt rewriting and image generation, whereas +Ours generates directly from the raw prompt without invoking PE.
Figure 8: Training dynamics and GenEval performance with different EMA settings. Left: the loss curves where faint lines denote raw losses and bold lines show a the moving average; Right: the corresponding GenEval performance curves.
Figure 9: Visual examples. We visualize the samples generated by the student and teacher across different EMA settings at 500 training steps.
Method
GenEval (GE) Task
GenEval2 (GE2) Task
GE
PickScore
CLIP
Aes.
GE2 GM
GE2 AM
PickScore
CLIP
Aes.
IPF
IVA
FLUX.2-klein-base (9B)
0.775
0.892
0.303
5.140
0.359
0.773
0.880
0.328
5.310
0.00%
0.00%
+PE
0.862
0.910
0.302
5.467
0.442
0.838
0.901
0.330
5.658
8.61%
4.33%
+Ours
0.873
0.913
0.305
5.421
0.444
0.831
0.905
0.333
5.628
9.20%
4.16%
Δ (vs Base)
+0.098
+0.021
+0.002
+0.281
+0.085
+0.058
+0.025
+0.005
+0.318
+9.20%
+4.16%
FLUX.2-klein (9B)
0.856
0.911
0.299
5.288
0.348
0.797
0.901
0.329
5.430
0.00%
0.00%
Appendix
Table 13: Results on scaling up to large models. We use GPT-5.6 Sol as PE. PickScore is normalized by 26; Aes. denotes aesthetics; Bold: best; Underlined: second-best.
Figure 10: Training dynamics with different PEs. We report the PE-OPSD loss over 1,000 training steps. Faint lines denote raw losses, while bold lines show a the moving average.
Figure 11: Comparison between the raw prompt and the enhanced prompt. Here, we use GPT-5.6 Sol as PE.
Figure 12: Comparison between the raw prompt and the enhanced prompt. Here, we use PromptEnhancer-7B as PE.
Figure 13: Comparison between the raw prompt and the enhanced prompt. Here, we use PromptEnhancer-32B as PE.
Method
Single Obj
Two Obj
Counting
Color
Position
Attr Binding
Overall
SD3.5-M (2.5B)
Base
0.975
0.778
0.613
0.787
0.223
0.472
0.628
Base+PE
0.959
0.838
0.634
0.838
0.603
0.635
0.743
SFT
0.972
0.856
0.691
0.832
0.512
0.500
0.718
Off-Policy Distillation
0.981
0.846
0.628
0.838
0.698
0.650
0.770
PE-OPSD (Ours)
0.988
0.886
0.631
0.878
0.698
0.710
0.797
Appendix
Table 14: Detailed GenEval performance breakdown. We report the fine-grained performance in Single Object, Two Object, Counting, Colors, Position, and Attribute Binding. We also report the Overall score. Here, we use GPT-5.6 Sol as PE.
Method
Single Obj
Two Obj
Counting
Color
Position
Attr Binding
Overall
PromptEnhancer-7B as PE
Base
0.959
0.801
0.569
0.809
0.338
0.480
0.650
Base+PE
0.947
0.801
0.588
0.750
0.505
0.555
0.684
SFT
0.966
0.871
0.762
0.835
0.627
0.560
0.763
Off-Policy Distillation
0.978
0.856
0.716
0.856
0.613
0.623
0.767
PE-OPSD (Ours)
0.978
0.864
0.647
0.872
0.610
0.740
0.782
Appendix
Table 15: Detailed GenEval performance breakdown on Z-image with other PEs. We report the fine-grained performance in Single Object, Two Object, Counting, Colors, Position, and Attribute Binding. We also report the Overall score.
Method
Object
Attribute
Count
Position
Verb
GenEval2 AM
GenEval2 GM
SD3.5-M (2.5B)
Base
0.859
0.672
0.431
0.397
0.161
0.633
0.176
Base+PE
0.868
0.751
0.476
0.521
0.207
0.680
0.220
SFT
0.835
0.703
0.478
0.438
0.124
0.656
0.207
Off-Policy Distillation
0.850
0.715
0.461
0.483
0.195
0.671
0.223
PE-OPSD (Ours)
0.868
0.758
0.477
0.489
0.179
0.682
0.226
Appendix
Table 16: Detailed GenEval2 performance breakdown. We report the fine-grained performance in Object, Attribute, Count, Position, and Verb. We also report the overall soft-tifa score in GenEval2 AM and GenEval2 GM . Here, we use GPT-5.6 Sol as PE.
Method
Object
Attribute
Count
Position
Verb
GenEval2 AM
GenEval2 GM
PromptEnhancer-7B as PE
Base
0.927
0.844
0.599
0.607
0.294
0.761
0.306
Base+PE
0.962
0.818
0.653
0.723
0.471
0.801
0.375
SFT
0.975
0.865
0.660
0.765
0.340
0.819
0.400
Off-Policy Distillation
0.975
0.832
0.649
0.720
0.274
0.804
0.380
PE-OPSD (Ours)
0.979
0.875
0.636
0.784
0.388
0.819
0.422
Appendix
Table 17: Detailed GenEval2 performance breakdown on Z-Image with other PEs. We report the fine-grained performance in Object, Attribute, Count, Position, and Verb. We also report the overall soft-tifa score in GenEval2 AM and GenEval2 GM .
Method
Attribute
Location
Color
Object
Material
A./H.
Food
Shape
Activity
Spatial
Counting
Overall
SD3.5-M
0.801
0.724
0.570
0.675
0.542
0.524
0.654
0.826
0.539
0.459
0.234
3.199
+Ours
0.820
0.763
0.634
0.710
0.587
0.562
0.697
0.832
0.551
0.588
0.283
3.368
Z-Image
0.797
0.723
0.618
0.709
0.620
0.608
0.656
0.806
0.582
0.655
0.314
3.340
+Ours
0.809
0.768
0.694
0.738
0.681
0.654
0.716
0.811
0.596
0.643
0.384
3.561
Z-Image-Turbo
0.800
0.762
0.711
0.727
0.707
0.620
0.699
0.798
0.589
0.641
0.340
3.515
+Ours
0.807
0.753
0.626
0.729
0.639
0.632
0.735
0.806
0.620
0.584
0.337
3.534
Appendix
Table 18: Detailed EvalMuse performance breakdown. We report the fine-grained performance in Attribute, Location, Color, Object, Material, A./H. (Animal/Human), Food, Shape, Activity, Spatial, and Counting. We also report the overall score. Bold: best in Overall.
Figure 14: Visual examples in SD3.5-M. We use the GPT-5.6 Sol as PE.
Figure 15: Visual examples in Z-Image. We use the GPT-5.6 Sol as PE.
Figure 16: Visual examples in Z-Image-Turbo. We use the GPT-5.6 Sol as PE.
Existing Flow Matching (FM) text-to-image models suffer from two critical bottlenecks under multi-task alignment: the reward sparsity induced by scalar-valued rewards, and the gradient interference arising from jointly optimizing heterogeneous objectives, which together give rise to a 'seesaw effect' of competing metrics and pervasive reward hacking. Inspired by the success of On-Policy Distillation (OPD) in the large language model community, we propose Flow-OPD, the first unified post-training framework that integrates on-policy distillation into Flow Matching models. Flow-OPD adopts a two-stage alignment strategy: it first cultivates domain-specialized teacher models via single-reward GRPO fine-tuning, allowing each expert to reach its performance ceiling in isolation; it then establishes a robust initial policy through a Flow-based Cold-Start scheme and seamlessly consolidates heterogeneous expertise into a single student via a three-step orchestration of on-policy sampling, task-routing labeling, and dense trajectory-level supervision. We further introduce Manifold Anchor Regularization (MAR), which leverages a task-agnostic teacher to provide full-data supervision that anchors generation to a high-quality manifold, effectively mitigating the aesthetic degradation commonly observed in purely RL-driven alignment. Built upon Stable Diffusion 3.5 Medium, Flow-OPD raises the GenEval score from 63 to 92 and the OCR accuracy from 59 to 94, yielding an overall improvement of roughly 10 points over vanilla GRPO, while preserving image fidelity and human-preference alignment and exhibiting an emergent 'teacher-surpassing' effect. These results establish Flow-OPD as a scalable alignment paradigm for building generalist text-to-image models. The codes and weights will be released in: https://github.com/CostaliyA/Flow-OPD .
Zhen Fang, Wenxuan Huang, Yu Zeng +8
University of Science and Technology of China · University of California, Los Angeles · 3The Chinese University of Hong Kong +1
Leading open text-to-image models often carry complementary strengths: one may lead on preference-aligned aesthetics while another follows compositional instructions more faithfully. However, differences in their autoencoders and noise schedules make it difficult to transfer these strengths across models. In this paper, we present Poly-OPD, a framework that can consolidate complementary strengths of heterogeneous teachers into a single compact flow-matching student. To bridge the incompatible latent spaces of different teachers, Poly-OPD performs on-policy distillation through a pixel bridge. Each student-generated image is re-encoded by a selected teacher's encoder and refined from a noise level matched by magnitude under the teacher's noise schedule. The resulting target is further matched to the student in frozen DINOv2 space, enabling supervision across incompatible latent spaces. To retain complementary capabilities without cross-teacher interference, Poly-OPD uses a gradient compatibility diagnostic to organize its adapters: attention LoRA modules are shared across teachers, whereas feed-forward adapters remain teacher-specific. During distillation, a gap-aware curriculum devotes more training to compositional categories where the student still falls short of the teacher. As each gap narrows, training shifts toward categories with larger remaining gaps. By distilling FLUX.1-dev and Z-Image into a 2.5B SD3.5-Medium student, Poly-OPD improves GenEval from 67.3 to 73.3, surpassing both larger teachers, and raises DrawBench HPSv3 from 9.34 to 11.35, consolidating both strengths within a switchable model.
Siming Fu, Haojun Xu, Ruizhe He +9
1Joy Future Academy · 2Zhejiang University · 3Beihang University
The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and FLUX.2-klein). However, these models present significant challenges for direct continuous supervised fine-tuning. For example, applying the commonly used fine-tuning technique would compromise their inherent few-step inference capability. To address this, we propose D-OPSD, a novel training paradigm for step-distilled diffusion models that enables on-policy learning during supervised fine-tuning. We first find that the modern diffusion models, where the LLM/VLM serves as the encoder, can inherit its encoder's in-context capabilities. This enables us to formulate the training as an on-policy self-distillation process. Specifically, during training, we make the model act as both the teacher and the student with different contexts, where the student is conditioned only on the text feature, while the teacher is conditioned on the multimodal feature of both the text prompt and the target image. Training minimizes the two predicted distributions over the student's own roll-outs. By optimizing on the model's own trajectory and under its own supervision, D-OPSD enables the model to learn new concepts, styles, etc., without sacrificing the original few-step capacity.
Dengyang Jiang, Xin Jin, Dongyang Liu +9
1The Hong Kong University of Science and Technology · 2Z-Image Team, Alibaba Group · 4The Chinese University of Hong Kong +1