Recent diffusion-based image generation backbones have grown substantially in scale, making the network inference cost increase rapidly. While diffusion distillation techniques can reduce the number of inference steps, high-quality image generation within a single full-backbone-forward compute budget remains challenging. Existing one-step methods typically allocate this budget to a single evaluation of a monolithic student. However, approximating the heterogeneous coarse-to-fine transport with a single monolithic mapping is difficult and often leads to over-smoothed outputs. To address this issue, we propose Phase-wise Velocity Distillation (PVD), which partitions the generation timeline into a coarse and a fine phase, and models the transition within each phase via the average velocity. A dedicated half-sized expert is assigned to each phase, decoupling structural composition from detail refinement while keeping the cumulative computation equivalent to one full-backbone forward pass. We show that the use of two half-sized phase-specific experts outperforms a single full-size monolithic student. On class-conditional image generation, PVD achieves an FID of 1.48 on ImageNet 256 x 256. On more complex text-to-image (T2I) tasks, PVD-distilled models (Stable Diffusion 3.5-Medium, FLUX.1-dev, Qwen-Image) produce results competitive with their multi-step teachers, significantly outperforming prior distillation methods. Moreover, across the evaluated T2I backbones, PVD reduces active parameters by 49.10-50.89% and peak VRAM by 45.76-48.36% compared to the corresponding teachers. Source code and distilled models are available at https://github.com/PolyU-VCLab/PVD.
Figures & tables
Figure 1 : Text-to-image results. Top : Teacher model (Stable Diffusion3.5-Medium, FLUX.1 dev and Qwen-Image under 50×2 , 50 , and 50×2 steps, respectively). Bottom : With two half-sized phase-wise experts, PVD achieves fast and high-quality image generation using cumulative computation equivalent to one full-backbone forward pass of the teacher model.
Figure 2 : Trajectory-based distillation (left) and distribution-based distillation (middle) employ a uniform student model to approximate the full coarse-to-fine mapping, while our proposed phase-wise velocity distillation (right) decomposes the objective into phase-wise velocity targets and sequentially evaluates two half-sized experts under a single full-backbone-forward compute budget.
Figure 3 : Overview of Phase-wise Velocity Distillation (PVD). (a) A half-sized backbone is extracted from the teacher and fine-tuned via flow matching to initialize the local students. (b) Phase-specific LoRA experts are distilled by mean velocity learning across partitioned time intervals. (c) For T2I tasks, phase-wise discriminators are used to provide adversarial refinement.
Table 4
Model
GenEval
DPG
WISE
QI
TIIF-Bench
Aesthetic
Pick
Image
Nflops
Active Params.
Peak VRAM
Bench
short
long
Score
Score
Reward
SD3.5 Medium [ 4 ]
0.6829
84.49
0.45
43.19
0.7166
0.7190
5.7212
21.45
0.6127
100.00
2.24
4.99
LADD [ 26 ]
0.0941
26.46
0.05
6.55
0.2232
0.2019
4.3533
18.40
-2.0724
1.00
2.27
4.86
DMD2 [ 28 ]
0.4563
46.15
0.31
29.27
0.5410
0.4931
4.7885
20.09
-0.1003
1.00
2.24
4.81
TwinFlow [ 58 ]
0.6171
72.81
0.37
30.55
0.5856
0.6094
5.0111
20.26
-0.0154
1.00
2.25
4.83
SenseFlow [ 30 ]
0.5222
81.98
0.31
29.43
0.6273
0.5957
4.9707
20.69
0.4834
1.00
2.24
4.81
Table 3 : Quality and efficiency of distilled text-to-image generation. The right two columns report active parameters (B) and peak VRAM (GiB), together with PVD’s reductions relative to its teacher model. Higher is better for quality metrics; the best and second-best results among distilled methods within each backbone group are highlighted in bold and underlined , respectively.
Figure 4 : Visual comparisons of T2I generation. Compared with existing distillation methods, PVD preserves much clearer structure and finer details under a single-forward budget, demonstrating close T2I generation quality to its multi-step teachers.
Figure 5 : Visual comparisons after further training on Unsplash data. The models further trained on Unsplash images are marked with ∗ , which can generate more authentic photorealistic images with improved lighting, shadows, composition, and fine-grained details.
Backbone
Weights
Aesthetic score ↑
DreamSim diversity ↑
GenEval ↑
SD3.5-Medium
PVD
6.4753
0.1473
0.6948
PVD ∗
6.5993
0.1832
0.6697
FLUX.1-dev
PVD
6.4471
0.1457
0.6532
PVD ∗
6.7941
0.2238
0.6334
Qwen-Image
PVD
6.6463
0.1079
0.8846
PVD ∗
6.7289
0.1432
0.8594
Table 4 : Results of further training on Unsplash data. The models further trained on Unsplash images are marked with ∗ . Better results per backbone are highlighted in bold .
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Transport factorization
Inference efficiency
Method
Phase-restricted training
Phase-specific parameters
Average-velocity objective
Phase-factorized inference
Reduced active params. & VRAM
Few-step distillation
MeanFlow
×
×
✓
×
×
✓
Hierarchical Distill.
×
×
Init.
×
×
✓
TimeStep Master
✓
✓
×
×
×
×
Glance
✓
✓
×
×
×
×
Phased DMD
✓
✓
×
✓
×
✓
Appendix
Table 5 : Comparison with representative related methods in terms of temporal localization of the training objective, parameter specialization, average-velocity modeling, and explicit phase-factorized inference. “Init.” indicates that average-velocity learning is used only for student initialization rather than retained as the final phase-specific transport objective.
Table 6 : Training configurations for PVD. Backbone initialization and phase-expert training are reported separately. Weight decay is 0 and gradient clipping is 1.0 .
Figure 6 : More class-conditional samples generated by PVD.
Model
Single Object
Two Object
Counting
Colors
Position
Attribute Binding
Overall ↑
SD3.5 Medium
0.9875
0.8131
0.6469
0.8324
0.2400
0.5775
0.6829
LADD
0.2812
0.0025
0.0625
0.2181
0.0000
0.0000
0.0941
DMD2
0.9000
0.3636
0.4750
0.6489
0.1200
0.2300
0.4563
TwinFlow
0.9594
0.6490
0.5688
0.7527
0.4250
0.3475
0.6171
SenseFlow
0.9344
0.6212
0.4562
0.6888
0.1625
0.2700
0.5222
SWD
0.9656
0.7601
0.5844
0.7793
0.2625
0.4200
0.6286
Appendix
Table 7 : Quantitative comparison on GenEval. Higher is better ( ↑ ) for all metrics. The best and second-best accelerated results within each backbone group are in bold and underlined , respectively.
Model
Global
Entity
Attribute
Relation
Other
Overall
Whole
State
Part
Color
Shape
Size
Tex.
Oth.
Spat.
N-Spat.
Cnt.
Txt.
SD3.5 Medium
91.04
90.44
90.07
88.36
90.36
89.01
89.99
90.17
89.28
91.15
89.46
91.03
87.40
84.49
LADD
35.63
40.89
49.58
60.31
58.88
51.69
38.16
47.68
47.23
54.96
57.50
60.65
51.86
26.46
DMD2
62.69
49.92
61.15
42.04
54.09
32.37
54.49
54.74
56.69
39.39
57.86
60.65
47.09
46.15
TwinFlow
82.16
83.64
82.10
83.57
86.04
89.33
82.89
80.87
82.32
85.47
83.95
84.87
81.22
72.81
SenseFlow
88.56
87.29
87.86
85.11
89.30
91.55
90.30
84.80
86.92
89.23
89.12
87.65
88.66
81.98
Appendix
Table 8 : Quantitative evaluation on DPG-Bench. Higher is better ( ↑ ) for all metrics. The best and second-best accelerated results within each backbone group are in bold and underlined , respectively.
Model
Cultural
Time
Space
Biology
Physics
Chemistry
Overall
SD3.5 Medium
0.43
0.50
0.52
0.41
0.53
0.33
0.45
LADD
0.04
0.03
0.07
0.04
0.05
0.07
0.05
DMD2
0.28
0.28
0.40
0.25
0.44
0.30
0.31
TwinFlow
0.27
0.40
0.53
0.40
0.50
0.38
0.37
SenseFlow
0.27
0.32
0.43
0.33
0.39
0.22
0.31
SWD
0.45
0.27
0.45
0.30
0.35
0.10
0.36
Appendix
Table 9 : Quantitative comparison on WISE. Higher is better ( ↑ ) for all metrics. The best and second-best accelerated results within each backbone group are in bold and underlined , respectively.
Model
TIIF Overall
Spatial (2D)
Spatial (3D)
Action
Color
Texture
S ↑
L ↑
(L/S)
(L/S)
2D
3D
Col
Tex
2D
3D
Tex
2D
3D
Col
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
SD3.5 Medium
71.66
71.90
83/83
79/70
91/86
82/82
73/65
68/75
95/90
70/50
92/88
90/85
83/77
92/100
LADD
22.32
20.19
0/41
29/41
8/13
17/12
26/34
20/27
40/40
35/20
16/8
28/19
33/50
36/56
DMD2
54.10
49.31
75/ 87
62 / 79
72/80
76/ 87
61/ 73
72 / 68
95 / 95
70/70
52/72
76/ 100
72/66
80/80
TwinFlow
58.56
60.94
79/ 91
79 /75
75 / 88
84 / 84
57/ 73
68/ 75
90 / 95
86 / 75
72/84
85 / 95
77 / 77
84/88
Appendix
Table 10 : Quantitative comparison on TIIF Benchmark mini (Part I: Visual & Relational). L/S stands for Long/Short. Scores are multiplied by 100. Higher is better ( ↑ ) for all metrics. The best and second-best accelerated results within each backbone group are in bold and underlined , respectively.
Model
Comparison
Differentiation
Negation
Base
2D
3D
Col
Tex
Base
2D
3D
Col
Tex
Base
2D
3D
Col
Tex
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
SD3.5 Medium
79/75
68/75
62/41
66/66
30/20
84/76
56/62
82/70
80/61
75/85
70/62
68/81
64/64
61/72
72/66
LADD
16/29
0/0
6/6
4/38
15/30
52/48
0/0
0/5
47/52
10/25
50/50
56/50
35/47
72/44
44/33
DMD2
66/ 70
75 /50
25/62
61/ 90
33/20
88 /80
43/50
84 /64
71/61
25/ 70
66 /54
56/ 75
64/ 76
72/ 77
61 / 55
TwinFlow
79/ 83
50/ 75
31/ 68
90 /57
70 / 40
92 / 84
87 / 93
70/70
71/ 90
55/ 75
62 /58
75 / 75
58/58
77 /61
55 / 55
Appendix
Table 11 : Quantitative comparison on TIIF Benchmark mini (Part II: Logic & Reasoning). L/S stands for Long/Short. Scores are multiplied by 100. Higher is better ( ↑ ) for all metrics. The best and second-best accelerated results within each backbone group are in bold and underlined , respectively.
Model
Numeracy
Shape
Real World
Style
Text
Base
2D
3D
Col
Tex
2D
3D
Col
Tex
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
(L/S)
SD3.5 Medium
72/70
77/48
78/87
65/77
70/59
45/40
65/70
78/86
72/68
75/73
70/66
48/67
LADD
22/12
7/22
25/40
22/30
31/9
25/15
20/20
26/36
24/16
27/23
3/3
0/0
DMD2
50/64
55/ 70
46/71
47/60
38/54
45/30
55/60
66/ 88
58/66
46/42
16/13
0/5
TwinFlow
68/ 70
62/55
65/71
60 /57
47/56
60 / 55
80 / 80
72/ 90
64/ 78
57/55
33/26
3/6
Appendix
Table 12 : Quantitative comparison on TIIF Benchmark mini (Part III: Concepts & Domains). L/S stands for Long/Short. Scores are multiplied by 100. Higher is better ( ↑ ) for all metrics.The best and second-best accelerated results within each backbone group are in bold and underlined , respectively.
Model
Quality
Aesthetics
Alignment
Real-world
Creative
Overall
Fidelity
Generation
SD3.5-Medium
47.77
45.88
41.83
42.08
33.05
43.19
LADD
1.22
7.19
4.41
30.44
0.54
6.55
DMD2
22.68
35.84
30.18
36.90
19.89
29.27
TwinFlow
29.21
37.13
28.83
36.03
18.54
30.55
SenseFlow
30.26
33.04
27.60
36.35
17.83
29.43
Appendix
Table 13 : Quantitative comparison on Qwen-Image-Bench. Higher is better ( ↑ ) for all metrics. Best results within each distillation group are in bold , and the second-best results are underlined .
Model
Compute
Model runtime
Model footprint
Change vs. teacher
Nflops
TFLOPs
Latency (s)
Throughput (images / s)
Active / Total Params. (B)
Peak VRAM (GiB)
Active Params.
Peak VRAM
SD3.5 Medium
100.00
701.93
15.49
0.07
2.24 / 2.24
4.99
Ref.
Ref.
LADD
1.00
7.02
0.16
6.69
2.27 / 2.27
4.86
↑ 1.34%
↓ 2.61%
DMD2
1.00
7.02
0.15
6.89
2.24 / 2.24
4.81
0.00%
↓ 3.61%
TwinFlow
1.00
7.02
0.16
6.67
2.25 / 2.25
4.83
↑ 0.45%
↓ 3.21%
SenseFlow
1.00
7.02
0.16
6.67
2.24 / 2.24
4.81
0.00%
↓ 3.61%
Appendix
Table 14 : Detailed hardware-cost comparison. The final two columns report changes relative to the corresponding teacher; ↓ and ↑ denote reductions and increases, respectively.
Figure 7 : Qualitative comparisons on SD3.5-Medium, FLUX.1-dev, and Qwen-Image. Rows are grouped by backbone and labeled with the corresponding teacher or distillation method.
Figure 8 : Additional visualizations after further training on Unsplash data. The models further trained on Unsplash images are marked with ∗ .
Figure 9 : Coarse-to-fine progression of PVD distilled from Qwen-Image. Within each pair, the left image visualizes the state after the first phase and the right image shows the output after the second phase. The first phase reveals coarse composition, approximate subject silhouettes, and color of regions despite residual noise. The second phase refines this structure with sharper edges, surface textures, and finer letter strokes.
Method
Nflops
Active Params. (B)
Total Params. (B)
TFLOPs
Peak VRAM (GiB)
GenEval ↑
Qwen-Image
100.00
20.43
20.43
5585.14
38.42
0.87
MeanFlow
4.00
20.44
20.44
223.40
38.48
0.44
MeanFlow
8.00
20.44
20.44
446.80
38.48
0.49
PVD (Ours)
1.02
10.40
10.55
56.81
19.84
0.88
Appendix
Table 15 : Reported quality–compute comparison on Qwen-Image. GenEval results for MeanFlow are sourced from [ 58 ] .
Configuration
Nflops
Active Params. (B)
Total Params. (B)
TFLOPs
Peak VRAM (GiB)
GenEval ↑
w/o discriminator
0.99
1.10
2.22
6.93
2.67
0.5599
Full-size student
1.00
2.24
2.24
7.02
4.99
0.4142
Full-size student
2.00
2.24
2.24
14.04
4.99
0.5695
PVD
0.99
1.10
2.22
6.93
2.67
0.6948
Appendix
Table 16 : Component and compute-control ablations on SD3.5-Medium.
Phase 1 → Phase 2
GenEval ↑
ImageReward ↑
Eearly→Elate
0.6948
0.4947
Eearly→Eearly
0.6754
0.4598
Elate→Elate
0.1820
-1.7190
Elate→Eearly
0.2346
-1.5811
Appendix
Table 17 : Expert specialization and execution-order ablations on SD3.5-Medium. Rows specify the expert used in the first and second trajectory phases, respectively.
Figure 10 : Prompt: A top-down studio photograph of exactly thirty-five small pink macarons arranged in a precise rectangular grid of five rows and seven columns on a dark slate tray. Every row contains exactly seven macarons and every column exactly five. All macarons are identical, separate, fully visible, and not touching. The entire tray is inside the frame. Both methods violate the requested count and grid dimensions despite producing recognizable, regularly arranged objects.
Figure 11 : Prompt: A clear studio photograph of one upright white ceramic mug with its handle on the right. A long red pencil passes diagonally through the empty hole of the mug handle, from upper left to lower right. The pencil is outside the cup, never entering the opening at the top. Both ends of the pencil extend visibly beyond the handle. Three-quarter front view showing the handle hole and the empty cup opening, on a plain gray tabletop. Both methods place the pencil through the cup opening and mug wall instead of the existing handle hole.