Organizations: University of Hong Kong, Hong Kong · University of British Columbia, Canada · Vector Institute for AI, Toronto, Canada · Kling Team, Kuaishou Technology · Canada CIFAR AI Chair
While flow matching is elegant, its reliance on single-sample conditional velocities leads to high-variance training targets that destabilize optimization and slow convergence. By explicitly characterizing this variance, we identify 1) a high-variance regime near the prior, where optimization is challenging, and 2) a low-variance regime near the data distribution, where conditional and marginal velocities nearly coincide. Leveraging this insight, we propose Stable Velocity, a unified framework that improves both training and sampling. For training, we introduce Stable Velocity Matching (StableVM), an unbiased variance-reduction objective, along with Variance-Aware Representation Alignment (VA-REPA), which adaptively strengthen auxiliary supervision in the low-variance regime. For inference, we show that dynamics in the low-variance regime admit closed-form simplifications, enabling Stable Velocity Sampling (StableVS), a finetuning-free acceleration. Extensive experiments on ImageNet 256×256 and large pretrained text-to-image and text-to-video models, including SD3.5, Flux, Qwen-Image, and Wan2.2, demonstrate consistent improvements in training efficiency and more than 2× faster sampling within the low-variance regime without degrading sample quality. Our code is available at https://github.com/linYDTHU/StableVelocity.
Figures & tables
Figure 1 : Variance curves of VCFM(t) with 15%–85% quantile bands. Evaluated on GMMs of varying dimensionality, CIFAR-10 images, and 256×256 ImageNet latents obtained by the Stable Diffusion VAE. The y -axis reports VCFM(t) normalized by the square root of the data dimension. See Appendix F.2 for details.
Figure 2 : Illustration of CFM variance VCFM(t) . (a) The low-variance regime ( t≤ξ ), where the posterior pt(x0∣xt) is sharply concentrated and the conditional velocity vt(xt∣x0) nearly coincides with the true velocity vt(xt) , yielding VCFM(t)≈0 . (b) The high-variance regime ( t>ξ ), the posterior spreads over multiple reference samples, causing the conditional velocity to fluctuate and resulting in a large VCFM(t) .
Figure 3 : Motivation for variance-aware representation alignment. (a) In the low-variance regime , the alignment loss remains consistently low on a pretrained model from REPA ( Yu et al., 2024 ) , indicating a learnable and informative supervision signal. In contrast, in the high-variance regime , the loss stays high, reflecting the ill-posed nature of semantic recovery from noise. (b) Restricting representation alignment to the low-variance regime yields the best FID, while applying it only in the high-variance regime provides minimal meaningful improvement over the baseline. These results indicate that representation alignment should be activated adaptively rather than uniformly along the diffusion trajectory.
Model
Epochs
FID ↓
sFID ↓
IS ↑
Prec . ↑
Rec. ↑
Latent Diffusion Transformers
MaskDiT
1600
2.28
5.67
276.6
0.80
0.61
DiT-XL/2
1400
2.27
4.60
278.2
0.83
0.57
SiT-XL/2
1400
2.06
4.50
270.3
0.82
0.59
Faster-DiT
400
2.03
4.63
264.0
0.81
0.60
Representation Alignment Methods
Table 1 : Comparison of latent diffusion transformers with CFG. We compare our method against baselines including MaskDiT ( Zheng et al., 2023 ) , DiT-XL/2 ( Peebles and Xie, 2023b ) , SiT-XL/2 ( Ma et al., 2024 ) , Faster-DiT ( Yao et al., 2024 ) , REPA ( Yu et al., 2024 ) , iREPA ( Singh et al., 2025 ) , REG ( Wu et al., 2025b ) , and REPA-E ( Leng et al., 2025 ) . The first Ours block uses the standard REPA sampling protocol, while the second adopts class-balanced sampling (marked with ∗ ) following REPA-E. Methods marked with † require fine-tuning autoencoders.
Method
Iter
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
SiT-B/2 (130M)
REPA
100k
52.06
8.18
26.8
0.45
0.59
Ours
100k
49.69
8.18
28.5
0.46
0.60
SiT-L/2 (458M)
REPA
100k
22.75
5.52
59.9
0.61
0.63
Ours
100k
21.03
5.51
63.9
0.62
0.63
Table 2 : Variation in Model Scale and Checkpoints. Comparison of our full method (StableVM + VA-REPA) against vanilla-REPA. Results are reported without CFG.
Methods
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
REPA
18.59
5.39
70.6
0.64
0.62
+Ours
17.12
5.39
74.8
0.65
0.63
REG
8.90
5.50
125.3
0.72
0.59
+Ours
8.11
5.34
128.8
0.74
0.60
iREPA
16.62
5.31
76.7
0.65
0.63
+Ours
16.02
5.30
78.6
0.66
0.63
Table 3 : Compatibility of StableVM + VA-REPA with REPA variants. Integration results for vanilla REPA, REG, and iREPA. All metrics are reported at 100k iterations.
Iter
Split point ξ
FID ↓
sFID ↓
IS ↑
Prec. ↑
Rec. ↑
100k
0.6
17.38
5.36
73.7
0.65
0.63
0.7
17.63
5.33
73.2
0.65
0.62
0.8
17.85
5.34
72.3
0.65
0.62
200k
0.6
10.57
5.02
104.2
0.69
0.64
0.7
10.56
5.03
105.4
0.69
0.64
0.8
10.73
5.02
103.4
0.68
0.65
Table 4 : Ablation on split point ξ . Impact of different split points across training stages. The default setting ( ξ=0.7 ) is highlighted.
Figure 4 : Ablation on VA-REPA weighting and StableVM bank capacity. Left: effect of different weighting schemes w(t) , showing that soft weightings outperform hard thresholding. Right: effect of memory bank capacity K , where K=256 already achieves near-optimal performance. All results are evaluated at 100k iterations. REPA baseline is shown as a dashed line.
Figure 5 : Visual comparison across prompts on SD3.5 ( Esser et al., 2024 ) . Results are generated using the Euler solver with 30 and 20 steps, and with StableVS replacing Euler in the low-variance regime , all under the same random seeds. Compared to the standard 20-step solver, StableVS yields outputs that more closely resemble the 30-step results. Zoom in for details. Additional qualitative comparisons are provided in Appendix H .
Solver configuration
Reference metrics
T2V-CompBench metrics
Base solver
Solver in [0,ξ]
Total steps
PSNR ↑
SSIM ↑
LPIPS ↓
Consist ↑
Dynamic ↑
Spatial ↑
Motion ↑
Action ↑
Interact ↑
Numeracy ↑
UniPC
UniPC(19)
30
–
–
–
0.842
0.120
0.607
0.299
0.749
0.708
0.476
UniPC(13)
20
15.61
0.593
0.377
0.821
0.123
0.618
0.265
0.720
0.703
0.462
StableVS(9)
20
31.10
0.942
0.036
0.843
0.123
0.610
0.289
0.753
0.699
0.476
Table 5 : Evaluation on T2V-CompBench at 640×480 p for Wan2.2 ( Wan et al., 2025 ) . StableVS replaces the base solver in the low-variance regime [0,ξ] (steps in parentheses), keeping the high-variance regime [ξ,1] unchanged. Split point fixed at ξ=0.85 . Highlighted rows show StableVS matches or exceeds 30-step baselines with fewer steps.
Solver configuration
Overall ↑
Reference metrics
Base solver
Solver in [0,ξ]
Total steps
PSNR ↑
SSIM ↑
LPIPS ↓
SD3.5-Large
Euler
Euler(19)
30
0.723
–
–
–
Euler(13)
20
0.710
16.93
0.753
0.333
StableVS(9)
20
0.723
36.92
0.980
0.021
DPM++
DPM++(19)
30
0.724
–
–
–
Table 6 : Evaluation on GenEval at 1024×1024 resolution. StableVS replaces the base solver in the low-variance regime [0,ξ] , while keeping the high-variance regime [ξ,1] unchanged. Numbers in parentheses (e.g., Euler(19)) denote the number of sampling steps in the low-variance regime . The split point is fixed to ξ=0.85 for all models. Highlighted rows demonstrate that StableVS achieves comparable results to the 30-step baseline with fewer total sampling steps. Full results are reported in Tab. 10 .
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Iter
Model
FID ↓
IS ↑
sFID ↓
Prec . ↑
Rec. ↑
30k
CFM
10.76
7.97
4.55
0.58
0.57
STF (Eq. 23 )
38.37
6.02
9.94
0.52
0.45
STF (original impl.)
9.91
8.01
4.51
0.58
0.58
StableVM
9.31
8.02
4.62
0.59
0.57
50k
CFM
7.50
8.30
4.25
0.60
0.59
STF (Eq. 23 )
29.18
6.59
7.65
0.52
0.50
Appendix
Table 7 : Unconditional CIFAR-10 generation. Performance comparison between CFM, STF, and StableVM. STF instantiated directly from Eq. 23 performs poorly, whereas StableVM achieves faster convergence and better sample quality via unbiased variance reduction. All results use n=2048 and 50k generated samples.
Figure 6 : Qualitative comparison of generated samples on CIFAR-10. Samples generated by the checkpoint at 50k iterations (a) CFM, (b) STF instantiated directly from Eq. 23 , (c) the original STF implementation, and (d) our StableVM. StableVM produces sharper and more coherent samples, consistent with its improved convergence and variance reduction.
Algorithm 2 Stable Velocity Matching with Classifier-free Guidance
# Steps
SiT
StableVM ( K=256 )
IS ↑
FID ↓
sFID ↓
Precision ↑
Recall ↑
IS ↑
FID ↓
sFID ↓
Precision ↑
Recall ↑
40k (SiT) / 50k (StableVM)
212.5
4.33
6.29
0.71
0.69
219.2
3.92
6.04
0.72
0.69
80k
223.8
3.53
5.61
0.73
0.68
227.3
3.35
5.61
0.73
0.67
120k
234.4
3.07
5.34
0.74
0.67
234.8
3.02
5.40
0.74
0.67
160k
239.7
2.81
5.24
0.74
0.67
238.9
2.87
5.74
0.74
0.66
200k
244.7
2.62
5.20
0.75
0.67
247.0
2.59
5.21
0.75
0.67
Appendix
Table 8 : Two-stage training experiment, comparing CFM and StableVM (bank size K=256 ). Models are trained for 500k steps with checkpoints evaluated every 40k steps. The results are reported with classifier-free guidance. StableVM consistently achieves better FID and IS across training, demonstrating improved optimization stability when training is restricted to the high-variance regime.
Figure 7 : Comparison of CFM and StableVM with n=2048 on the synthetic GMM distribution. We plot the second-order moment as a function of training iterations at four time steps: t=0.20,0.30,0.40, and 0.50 .
Total steps
ξ
[0,ξ] steps
fβ
Overall ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Baseline
20
–
–
–
0.710
16.93
0.753
0.333
Default StableVS setting
20
0.85
9
0.0
0.723
36.92
0.980
0.021
Variance factor fβ
20
0.85
9
0.1
0.719
32.69
0.956
0.036
Appendix
Table 9 : Ablation study on hyperparameters of StableVS. We analyze the effects of the split point ξ , the number of steps in the low-variance regime [0,ξ] , and the variance factor fβ of Stable Velocity Sampling for SD3.5-Large ( Esser et al., 2024 ) at 1024×1024 resolution. All experiments use Euler as the base solver.
Solver configuration
GenEval Metrics
Reference metrics
Base solver
Solver in [0,ξ]
Total steps
Overall ↑
Single ↑
Two ↑
Counting ↑
Colors ↑
Position ↑
Color Attr ↑
PSNR ↑
SSIM ↑
LPIPS ↓
SD3.5-Large ( Esser et al., 2024 )
Euler
Euler(19)
30
0.723
0.994
0.907
0.697
0.843
0.293
0.608
–
–
–
Euler(13)
20
0.710
0.994
0.884
0.650
0.838
0.290
0.603
16.93
0.753
0.333
StableVS(9)
20
0.723
0.994
0.891
0.700
0.838
0.298
0.615
36.92
0.980
0.021
DPM++
DPM++(19)
30
0.724
0.997
0.917
0.700
0.825
0.278
0.630
–
–
–
Appendix
Table 10 : Detailed evaluation on GenEval at 1024×1024 resolution. We report the overall GenEval score together with per-category breakdowns. StableVS replaces the base solver in the low-variance regime [0,ξ] (with the number of steps indicated in parentheses), while keeping the high-variance regime [ξ,1] unchanged. The split point is fixed to ξ=0.85 for all models.
Figure 8 : Visual comparison of SD3.5-Large ( Esser et al., 2024 ) on different prompts. Results are generated using the Euler solver with 30 and 20 steps, and with StableVS replacing Euler in the low-variance regime , all under the same random seed. Compared to the standard 20-step solver, StableVS yields outputs that more closely resemble the 30-step results. Zoom in for details.
Figure 9 : Visual comparison of Flux-dev ( Labs, 2024 ) on different prompts. Results are generated using the Euler solver with 30 and 20 steps, and with StableVS replacing Euler in the low-variance regime , all under the same random seed. Compared to the standard 20-step solver, StableVS yields outputs that more closely resemble the 30-step results. Zoom in for details.
Figure 10 : Visual comparison of Qwen-Image-2512 ( Wu et al., 2025a ) on different prompts. Results are generated using the Euler solver with 30 and 17 steps, and with StableVS replacing Euler in the low-variance regime , all under the same random seed. Compared to the standard 17-step solver, StableVS yields outputs that more closely resemble the 30-step results. Zoom in for details.
Figure 11 : Visual comparison of Wan2.2 ( Wan et al., 2025 ) on different prompts. Results are generated using UniPC solver with 30 and 20 steps, and with StableVS replacing UniPC in the low-variance regime , all under the same random seed. Compared to the standard 20-step solver, StableVS yields outputs that more closely resemble the 30-step results. Zoom in for details.
Figure 12 : Visual comparison of Wan2.2 ( Wan et al., 2025 ) on different prompts. Results are generated using UniPC solver with 30 and 20 steps, and with StableVS replacing UniPC in the low-variance regime , all under the same random seed. Compared to the standard 20-step solver, StableVS yields outputs that more closely resemble the 30-step results. Zoom in for details.
While Flow Matching theoretically guarantees constant-velocity trajectories, we identify a critical breakdown in high-dimensional practice: the Velocity Deficit. We show that the MSE objective systematically underestimates velocity magnitude, causing generated samples to fail to reach the data manifold-a phenomenon we term Integration Lag. To rectify this, we propose Initial Energy Injection, instantiated via two complementary methods: the training-based Magnitude-Aware Flow Matching (MAFM) and the training-free Scale Schedule Corrector (SSC). Both are grounded in our discovery of a crucial asymmetry: velocity contraction causes harmful kinetic stagnation at the trajectory's start, yet acts as a beneficial denoising mechanism at its end. Empirically, SSC yields significant efficiency gains with zero retraining and just one line of code. On ImageNet-1k (256x256), it improves FID by 44.6% (from 13.68 to 7.58) and achieves a 5x speedup, enabling a 50-step generator (FID 7.58) to beat a 250-step baseline (FID 8.65). Furthermore, our methods generalize to Text-to-Image tasks and high-resolution generation, improving FID on MS-COCO by ~22%.
Flow-based models learn a target distribution by modeling a marginal velocity field, defined as the average of sample-wise velocities connecting each sample from a simple prior to the target data. When sample-wise velocities conflict at the same intermediate state, however, this averaged velocity can misguide samples toward low-density regions, degrading generation quality. To address this issue, we propose the Flow Divergence Sampler (FDS), a training-free framework that refines intermediate states before each solver step. Our key finding reveals that the severity of this misguidance is quantified by the divergence of the marginal velocity field that is readily computable during inference with a well-optimized model. FDS exploits this signal to steer states toward less ambiguous regions. As a plug-and-play framework compatible with standard solvers and off-the-shelf flow backbones, FDS consistently improves fidelity across various generation tasks including text-to-image synthesis, and inverse problems.
Recent work has popularized a practical recipe in diffusion and flow matching: predict the clean signal x, convert it to a velocity, and train through a velocity-space loss. The conversion contains a singular endpoint amplification and therefore appears prone to unstable optimization, yet recent systems obtain strong empirical results with this recipe. We investigate this tension through the integrability of the pre-optimizer stochastic-gradient second moment. Under stated initialization conditions, the moment diverges under Uniform sampling; boundary-suppressing sampling can restore integrability under an additional upper-growth condition. We then show that prediction--loss alignment eliminates this conversion-induced source of non-integrability. Under a uniform moment bound, alignment yields a finite second moment for every timestep density, including Uniform sampling. Controlled experiments across continuous and binary settings reproduce the predicted sampler-dependent instability and show that aligned objectives remain trainable across the tested samplers. These results reconcile pointwise amplification with sampler-dependent empirical success and support alignment as a principled route to more robust flow-matching training.
Jiadong Hong, Lei Liu, Xinyu Bian +2
Zhejiang University, Hangzhou, China · Huawei Technologies Co., Ltd, Hong Kong