Flow matching has emerged as a scalable paradigm for training high-quality generative models, but sampling from the learned probability flow requires many network evaluations. Distillation can reduce this cost to one or a few evaluations; however, one-step generation often sacrifices quality, making few-step generation the practical operating regime. Existing few-step methods perform their iterative computation along the probability flow and therefore require a fixed, manually chosen timestep discretization. This discretization is often chosen heuristically and is expensive to tune; it may also be restrictive when refinement difficulty differs across samples or spatial locations. We introduce data-space iteration, a few-step generation framework that removes flow discretization altogether. Starting from noise, a shared generator directly refines its prediction in data space, with every iteration trained to produce the best sample permitted by its capacity. Our formulation integrates with distribution matching distillation (DMD) with minimal changes, enabling a controlled comparison between iteration methods under matched training settings. On class-conditional ImageNet 256x256, data-space iteration outperforms standard discretization baselines and matches or improves upon variants selected through schedule search, without requiring schedule-specific training. These results show that data-space iteration provides a simple and effective alternative to discretized flow-space iteration for fast generation.
Figures & tables
Figure 1: (a) Existing few-step generators perform iterative sampling in probability-flow space, requiring a predefined timestep discretization that can be suboptimal. (b) Our method carries out all refinements in data space, removing the discretization bottleneck and yielding better performance.
Table 1: Ablation studies of the data-space iteration generator. FID lower is better.
Table 2: Ablation studies across sampling methods. FID lower is better.
Method
Paradigm
Discretization
CFG
250NFE
1NFE
2NFE
3NFE
4NFE
Flow Matching
flow-ODE
uniform
No
8.26
DMD
flow-renoise
uniform
4.45
3.22
2.67
2.75
oracle (best of 3)
4.20
3.44
2.88
2.72
Phased DMD
flow-map
uniform
-
-
-
2.55
oracle (best of 3)
-
-
-
2.48
Data-space DMD
data-space
none
3.85
2.68
2.47
2.46
Table 3: Main results. FID lower is better.
Figure 2: Visualization of iterative generation trajectories. Each row uses the same initial noise and class condition. Flow-based methods use the uniform discretization.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Value
Model architecture (DiT-XL/2)
Parameters
674M
FLOPs
119G
Depth
28
Hidden dimension
1152
Attention heads
16
Appendix
Table 4: Final architecture and training settings used to train the data-space iteration generator.
Table 5: Full metrics.
Figure 3: Visualization of iterative generation trajectories and their stepwise differences. For each sample, the first row shows the generated images and the second row shows the signed difference between consecutive steps, aligned beneath the later step. Zero difference is mapped to mid-gray using the same fixed normalization for all images.
Method
r=1
r=2
r=3
Flow-renoise
82.72
51.18
29.91
Data-space
59.65
44.25
35.98
Appendix
Table 6: Average L2 difference between rollout steps measured over 50k samples.
Table 7: Probability density functions of the three Beta distributions.
We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching models. Rather than performing velocity regression in the conventional VAE latent space, PFM supervises flow matching in a perceptual feature space using pretrained perceptual models. This simple change substantially improves the few-step generation capability of flow-matching models, reducing the number of sampling steps from 35-50 to 4-8 while preserving generation quality. Unlike existing acceleration and distillation approaches, PFM requires neither teacher models nor auxiliary score networks and can be integrated into standard flow-matching training pipelines with minimal modifications. Extensive experiments on image generation, video generation, and image editing tasks demonstrate that PFM consistently produces high-quality results while producing fewer artifacts than existing distillation-based methods. We further show that perceptual supervision shifts the regression minimizer from mean-seeking to mode-seeking, biasing predictions toward on-manifold modes that remain accurate under coarse few-step integration. Our results reveal that standard flow-matching training can naturally yield high-quality few-step generators when supervised in an appropriate representation space. We hope this insight inspires future research into representation-aware objectives for efficient generative modeling.
Chuyang Zhao, Yifei Song, Hongfa Wang +7
1Joy Future Academy · 2Fudan University · 3Tsinghua University +1
Diffusion models and flow-based methods have shown impressive generative capability, especially for images, but their sampling is expensive because it requires many iterative updates. We introduce W-Flow, a framework for training a generator that transforms samples from a simple reference distribution into samples from a target data distribution in a single step. This is achieved in two steps: we first define an evolution from the reference distribution to the target distribution through a Wasserstein gradient flow that minimizes an energy functional; second, we train a static neural generator to compress this evolution into one-step generation. We instantiate the energy functional with the Sinkhorn divergence, which yields an efficient optimal-transport-based update rule that captures global distributional discrepancy and improves coverage of the target distribution. We further prove that the finite-sample training dynamics converge to the continuous-time distributional dynamics under suitable assumptions. Empirically, W-Flow sets a new state of the art for one-step ImageNet 256×256 generation, achieving 1.29 FID, with improved mode coverage and domain transfer. Compared to multi-step diffusion models with similar FID scores, our method yields approximately 100× faster sampling. These results show that Wasserstein gradient flows provide a principled and effective foundation for fast and high-fidelity generative modeling.
Flow Matching (FM) is a simulation-free method for learning a continuous, invertible flow that interpolates between two distributions, and in particular generates data from noise. Inspired by the variational nature of the diffusion process as a gradient flow, we introduce a stepwise FM model, Local Flow Matching (LFM), which sequentially learns a sequence of FM submodels, each matching a diffusion process up to the time-step size in the data-to-noise direction. In each step, the two distributions to be interpolated by the sub-flow model are closer than those in the full-flow matching model, which interpolates data to noise distributions, enabling smaller models with more efficient training. This variational perspective also allows us to prove a theoretical generation guarantee for the proposed flow model in terms of the χ2-divergence between the generated and true data distributions, leveraging the contraction property of the diffusion process. In practice, the stepwise structure of LFM is naturally amenable to model distillation, and various distillation techniques can be applied to accelerate generation. We empirically demonstrate that LFM achieves competitive generative performance compared to FM on unconditional generation of tabular and image datasets, and on conditional generation of robotic manipulation policies.
Chen Xu, Xiuyuan Cheng, Yao Xie
H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology. · Department of Mathematics, Duke University