Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During training, a raw-prompt student follows its own generation trajectory, while an enhanced-prompt teacher provides vector-field targets at the states visited by the student. This on-policy supervision distills the behavior induced by enhanced prompts into the raw-prompt student without requiring additional text--image pairs. At inference, both the PE and teacher are removed, and the student generates directly from raw prompts. Across multiple model families, PEs, and benchmarks, PE-OPSD achieves the strongest aggregate prompt fidelity among the evaluated baselines, yields positive aggregate visual appeal gains, and retains the base-model inference efficiency.
Figures & tables
Figure 1: Comparison of generation quality, throughput, and training dynamics. PE-OPSD achieves higher GenEval scores than both Vanilla and PE while preserving the inference efficiency. Moreover, PE-OPSD achieves faster convergence and greater performance gains than SFT and off-policy distillation. Training dynamics is reported in Z-Image-Turbo training.
Figure 2: Three different uses of prompt conditioning. (a) Direct generation from raw prompts results in limited prompt alignment; (b) Inference-time PE improves prompt alignment by rewriting raw prompts, but introduces additional computational overhead; (c) Our PE-OPSD internalizes the PE knowledge into the generator, improving alignment without additional inference cost.
Figure 3
Figure 4: Overview of PE-OPSD. At training, the student generates rollouts from raw prompts, while an EMA teacher provides enhanced-prompt supervision on the same rollout states. At inference, the trained student generates directly from raw prompts without inference-time PE.
Method
GenEval (GE) Task
GenEval2 (GE2) Task
GE
PickScore
CLIP
Aes.
GE2 GM
GE2 AM
PickScore
CLIP
Aes.
IPF
IVA
SD3.5-L
0.604
0.862
0.288
5.286
0.213
0.660
0.869
0.315
5.521
-
-
FLUX.1-dev
0.618
0.899
0.280
5.624
0.179
0.643
0.895
0.298
5.909
-
-
SD3.5-M (2.5B)
Base
0.628
0.878
0.289
5.319
0.176
0.633
0.882
0.313
5.579
0.00%
0.00%
Base+PE
0.743
0.892
0.291
5.439
0.220
0.680
0.887
0.312
5.678
10.22%
1.55%
Table 2: Main results across different models. We use GPT-5.6 Sol as PE. PickScore is normalized by 26; Aes. denotes aesthetics; Bold: best; Underlined: second-best.
Figure 5: Qualitative comparison and human preference study on Z-Image-Turbo. We use GPT-5.6 Sol as PE. Left: visual comparisons among Base, PE, SFT, Off-policy distillation, and our method. Right: pairwise human preferences against Base in prompt fidelity and visual appeal.
Method
GE
GE2 GM
GE2 AM
Lat. (s)
PromptEnhancer-7B as PE
Base
0.650
0.306
0.761
9.69
Base+PE
0.684
0.375
0.801
19.00
Base+Ours
0.782
0.422
0.819
9.69
Δ (vs +PE)
+0.098
+0.047
+0.018
1.96 ×
PromptEnhancer-32B as PE
Table 3: Performance and latency.
Method
GenEval (GE) Task
GenEval2 (GE2) Task
GE
PickScore
CLIP
Aes.
GE2 GM
GE2 AM
PickScore
CLIP
Aes.
IPF
IVA
GPT-5.6 Sol as PE
Z-Image
0.650
0.875
0.288
5.288
0.306
0.761
0.871
0.319
5.421
0.00%
0.00%
Z-Image+PE
0.810
0.899
0.299
5.433
0.404
0.827
0.893
0.326
5.638
14.27%
3.00%
SFT
0.821
0.893
0.302
5.426
0.403
0.834
0.887
0.326
5.466
14.93%
1.83%
Off-Policy Distillation
0.824
0.901
0.300
5.432
0.395
0.828
0.897
0.327
5.629
14.27%
3.12%
Table 4: Result across different PEs. Bold: best; Underlined: second-best.
Method
DPG-bench
T2I-CompBench++
EvalMuse
Overall ⋆
Glo.
Ent.
Attr.
Rela.
Other
Overall ⋆
Color
Shape
Tex.
Num.
Comp.
Spa.
3D Spa.
Non-spa.
Overall ⋆
SD3.5-M
83.9
84.6
89.6
88.1
93.0
80.9
0.51
0.80
0.54
0.74
0.59
0.37
0.32
0.36
0.31
3.199
+Ours
83.8
83.9
89.3
88.4
93.0
82.1
0.55
0.83
0.60
0.73
0.63
0.39
0.46
0.42
0.31
3.368
Z-Image
86.3
83.5
91.7
90.1
94.6
87.8
0.53
0.85
0.59
0.79
0.63
0.40
0.31
0.37
0.31
3.340
+Ours
87.0
82.6
92.0
90.2
94.5
89.4
0.58
0.87
0.61
0.81
0.71
0.42
0.36
0.41
0.32
3.561
Z-Image-Turbo
84.5
78.6
91.2
88.3
93.2
88.4
0.53
0.81
0.56
0.75
0.69
0.40
0.36
0.41
0.31
3.515
Table 5: Results on out-of-domain benchmarks. For DPG-bench, we report Global (Glo.), Entity (Ent.), Attribute (Attr.), Relation (Rela.), Other and Overall ⋆ . For T2I-CompBench++, we report Color, Shape, Textual (Tex.), Numeracy (Num.), Complex (Comp.), Spatial (Spa.), 3D Spatial (3D Spa.), Non-spatial (Non-spa.) and Overall ⋆ . Bold: best in Overall ⋆ .
Figure 6: Scaling to larger models. PickScore, CLIP, and Aesthetics are averaged over GenEval and GenEval2. For visualization, each metric is independently normalized to [0.3,1.0] .
Figure 7: Ablation study results on different training datasets. We use MixDataset by default.
Table 12
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Setting
Optimizer
AdamW
Learning rate
1×10−4 , constant without warm-up
Optimizer momentum
(β1,β2)=(0.9,0.999)
Adam ϵ / weight decay
10−8 / 0
Gradient clipping
1.0
Global batch size
64
Appendix
Table 9: Default training hyperparameters for PE-OPSD. All models follow the same recipe unless otherwise specified.
Model
Checkpoint
SD3.5-Medium
stabilityai/stable-diffusion-3.5-medium
SD3.5-Large
stabilityai/stable-diffusion-3.5-large
Z-Image
Tongyi-MAI/Z-Image
Z-Image-Turbo
Tongyi-MAI/Z-Image-Turbo
FLUX.1-dev
black-forest-labs/FLUX.1-dev
FLUX.2-klein-base-9B
black-forest-labs/FLUX.2-klein-base-9B
Appendix
Table 10: Models and checkpoints used in our experiments.
Model
SFT
Off-policy
PE-OPSD
One-step
Full †
One-step
Full
One-step
Full
SD3.5-M
1.47 s
0.4 + 6.6 h
11.02 s
3.1 h
11.02 s
3.1 h
Z-Image
3.14 s
0.9 + 27.5 h
49.15 s
13.6 h
49.15 s
13.6 h
Z-Image-Turbo
3.14 s
0.9 + 2.5 h
14.31 s
4.0 h
14.31 s
4.0 h
Appendix
Table 11: Training efficiency in our experiments. One-step denotes the wall-clock time per optimization step under the default configuration, and Full denotes the cost of 1,000 optimization steps. For SFT, Full † is reported as training time + offline pseudo-target generation time.
Model
Latency (s)
Base
+PE
+Ours
Speedup
PromptEnhancer-7B as PE
SD3.5-M
2.70
10.36
2.70
3.84 ×
Z-Image
9.69
19.00
9.69
1.96 ×
Z-Image-Turbo
0.88
8.53
0.88
9.69 ×
FLUX.2-klein
0.67
8.18
0.67
12.21 ×
Appendix
Table 12: Inference latency and speedup across different settings. Latency is the average generation time per image in seconds. +PE includes both prompt rewriting and image generation, whereas +Ours generates directly from the raw prompt without invoking PE.
Figure 8: Training dynamics and GenEval performance with different EMA settings. Left: the loss curves where faint lines denote raw losses and bold lines show a the moving average; Right: the corresponding GenEval performance curves.
Figure 9: Visual examples. We visualize the samples generated by the student and teacher across different EMA settings at 500 training steps.
Method
GenEval (GE) Task
GenEval2 (GE2) Task
GE
PickScore
CLIP
Aes.
GE2 GM
GE2 AM
PickScore
CLIP
Aes.
IPF
IVA
FLUX.2-klein-base (9B)
0.775
0.892
0.303
5.140
0.359
0.773
0.880
0.328
5.310
0.00%
0.00%
+PE
0.862
0.910
0.302
5.467
0.442
0.838
0.901
0.330
5.658
8.61%
4.33%
+Ours
0.873
0.913
0.305
5.421
0.444
0.831
0.905
0.333
5.628
9.20%
4.16%
Δ (vs Base)
+0.098
+0.021
+0.002
+0.281
+0.085
+0.058
+0.025
+0.005
+0.318
+9.20%
+4.16%
FLUX.2-klein (9B)
0.856
0.911
0.299
5.288
0.348
0.797
0.901
0.329
5.430
0.00%
0.00%
Appendix
Table 13: Results on scaling up to large models. We use GPT-5.6 Sol as PE. PickScore is normalized by 26; Aes. denotes aesthetics; Bold: best; Underlined: second-best.
Figure 10: Training dynamics with different PEs. We report the PE-OPSD loss over 1,000 training steps. Faint lines denote raw losses, while bold lines show a the moving average.
Figure 11: Comparison between the raw prompt and the enhanced prompt. Here, we use GPT-5.6 Sol as PE.
Figure 12: Comparison between the raw prompt and the enhanced prompt. Here, we use PromptEnhancer-7B as PE.
Figure 13: Comparison between the raw prompt and the enhanced prompt. Here, we use PromptEnhancer-32B as PE.
Method
Single Obj
Two Obj
Counting
Color
Position
Attr Binding
Overall
SD3.5-M (2.5B)
Base
0.975
0.778
0.613
0.787
0.223
0.472
0.628
Base+PE
0.959
0.838
0.634
0.838
0.603
0.635
0.743
SFT
0.972
0.856
0.691
0.832
0.512
0.500
0.718
Off-Policy Distillation
0.981
0.846
0.628
0.838
0.698
0.650
0.770
PE-OPSD (Ours)
0.988
0.886
0.631
0.878
0.698
0.710
0.797
Appendix
Table 14: Detailed GenEval performance breakdown. We report the fine-grained performance in Single Object, Two Object, Counting, Colors, Position, and Attribute Binding. We also report the Overall score. Here, we use GPT-5.6 Sol as PE.
Method
Single Obj
Two Obj
Counting
Color
Position
Attr Binding
Overall
PromptEnhancer-7B as PE
Base
0.959
0.801
0.569
0.809
0.338
0.480
0.650
Base+PE
0.947
0.801
0.588
0.750
0.505
0.555
0.684
SFT
0.966
0.871
0.762
0.835
0.627
0.560
0.763
Off-Policy Distillation
0.978
0.856
0.716
0.856
0.613
0.623
0.767
PE-OPSD (Ours)
0.978
0.864
0.647
0.872
0.610
0.740
0.782
Appendix
Table 15: Detailed GenEval performance breakdown on Z-image with other PEs. We report the fine-grained performance in Single Object, Two Object, Counting, Colors, Position, and Attribute Binding. We also report the Overall score.
Method
Object
Attribute
Count
Position
Verb
GenEval2 AM
GenEval2 GM
SD3.5-M (2.5B)
Base
0.859
0.672
0.431
0.397
0.161
0.633
0.176
Base+PE
0.868
0.751
0.476
0.521
0.207
0.680
0.220
SFT
0.835
0.703
0.478
0.438
0.124
0.656
0.207
Off-Policy Distillation
0.850
0.715
0.461
0.483
0.195
0.671
0.223
PE-OPSD (Ours)
0.868
0.758
0.477
0.489
0.179
0.682
0.226
Appendix
Table 16: Detailed GenEval2 performance breakdown. We report the fine-grained performance in Object, Attribute, Count, Position, and Verb. We also report the overall soft-tifa score in GenEval2 AM and GenEval2 GM . Here, we use GPT-5.6 Sol as PE.
Method
Object
Attribute
Count
Position
Verb
GenEval2 AM
GenEval2 GM
PromptEnhancer-7B as PE
Base
0.927
0.844
0.599
0.607
0.294
0.761
0.306
Base+PE
0.962
0.818
0.653
0.723
0.471
0.801
0.375
SFT
0.975
0.865
0.660
0.765
0.340
0.819
0.400
Off-Policy Distillation
0.975
0.832
0.649
0.720
0.274
0.804
0.380
PE-OPSD (Ours)
0.979
0.875
0.636
0.784
0.388
0.819
0.422
Appendix
Table 17: Detailed GenEval2 performance breakdown on Z-Image with other PEs. We report the fine-grained performance in Object, Attribute, Count, Position, and Verb. We also report the overall soft-tifa score in GenEval2 AM and GenEval2 GM .
Method
Attribute
Location
Color
Object
Material
A./H.
Food
Shape
Activity
Spatial
Counting
Overall
SD3.5-M
0.801
0.724
0.570
0.675
0.542
0.524
0.654
0.826
0.539
0.459
0.234
3.199
+Ours
0.820
0.763
0.634
0.710
0.587
0.562
0.697
0.832
0.551
0.588
0.283
3.368
Z-Image
0.797
0.723
0.618
0.709
0.620
0.608
0.656
0.806
0.582
0.655
0.314
3.340
+Ours
0.809
0.768
0.694
0.738
0.681
0.654
0.716
0.811
0.596
0.643
0.384
3.561
Z-Image-Turbo
0.800
0.762
0.711
0.727
0.707
0.620
0.699
0.798
0.589
0.641
0.340
3.515
+Ours
0.807
0.753
0.626
0.729
0.639
0.632
0.735
0.806
0.620
0.584
0.337
3.534
Appendix
Table 18: Detailed EvalMuse performance breakdown. We report the fine-grained performance in Attribute, Location, Color, Object, Material, A./H. (Animal/Human), Food, Shape, Activity, Spatial, and Counting. We also report the overall score. Bold: best in Overall.
Figure 14: Visual examples in SD3.5-M. We use the GPT-5.6 Sol as PE.
Figure 15: Visual examples in Z-Image. We use the GPT-5.6 Sol as PE.
Figure 16: Visual examples in Z-Image-Turbo. We use the GPT-5.6 Sol as PE.