Flow matching has emerged as a scalable paradigm for training high-quality generative models, but sampling from the learned probability flow requires many network evaluations. Distillation can reduce this cost to one or a few evaluations; however, one-step generation often sacrifices quality, making few-step generation the practical operating regime. Existing few-step methods perform their iterative computation along the probability flow and therefore require a fixed, manually chosen timestep discretization. This discretization is often chosen heuristically and is expensive to tune; it may also be restrictive when refinement difficulty differs across samples or spatial locations. We introduce data-space iteration, a few-step generation framework that removes flow discretization altogether. Starting from noise, a shared generator directly refines its prediction in data space, with every iteration trained to produce the best sample permitted by its capacity. Our formulation integrates with distribution matching distillation (DMD) with minimal changes, enabling a controlled comparison between iteration methods under matched training settings. On class-conditional ImageNet 256x256, data-space iteration outperforms standard discretization baselines and matches or improves upon variants selected through schedule search, without requiring schedule-specific training. These results show that data-space iteration provides a simple and effective alternative to discretized flow-space iteration for fast generation.
Figures & tables
Figure 1: (a) Existing few-step generators perform iterative sampling in probability-flow space, requiring a predefined timestep discretization that can be suboptimal. (b) Our method carries out all refinements in data space, removing the discretization bottleneck and yielding better performance.
Table 1: Ablation studies of the data-space iteration generator. FID lower is better.
Table 2: Ablation studies across sampling methods. FID lower is better.
Method
Paradigm
Discretization
CFG
250NFE
1NFE
2NFE
3NFE
4NFE
Flow Matching
flow-ODE
uniform
No
8.26
DMD
flow-renoise
uniform
4.45
3.22
2.67
2.75
oracle (best of 3)
4.20
3.44
2.88
2.72
Phased DMD
flow-map
uniform
-
-
-
2.55
oracle (best of 3)
-
-
-
2.48
Data-space DMD
data-space
none
3.85
2.68
2.47
2.46
Table 3: Main results. FID lower is better.
Figure 2: Visualization of iterative generation trajectories. Each row uses the same initial noise and class condition. Flow-based methods use the uniform discretization.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Value
Model architecture (DiT-XL/2)
Parameters
674M
FLOPs
119G
Depth
28
Hidden dimension
1152
Attention heads
16
Appendix
Table 4: Final architecture and training settings used to train the data-space iteration generator.
Table 5: Full metrics.
Figure 3: Visualization of iterative generation trajectories and their stepwise differences. For each sample, the first row shows the generated images and the second row shows the signed difference between consecutive steps, aligned beneath the later step. Zero difference is mapped to mid-gray using the same fixed normalization for all images.
Method
r=1
r=2
r=3
Flow-renoise
82.72
51.18
29.91
Data-space
59.65
44.25
35.98
Appendix
Table 6: Average L2 difference between rollout steps measured over 50k samples.
Table 7: Probability density functions of the three Beta distributions.