Velocity Scaling in Flow Matching
Organizations: University of Geneva Geneva, Switzerland · University of Geneva and Meta FAIR Geneva, Switzerland
Abstract
Scaling a learned flow-matching velocity field by a gain was recently shown to greatly improve generation quality. Prior work argued that velocity fields trained with mean-squared error (MSE) systematically underestimate velocity magnitude and that scaling corrects this error. We show that MSE training does not create a velocity-magnitude deficit. We find instead that velocity scaling reduces population time lag: sampled states at model time resemble training states from an earlier time. Velocity scaling and moving model time back are two ways to address this population time lag. Across architectures and model sizes, measuring population time lag and using it to select a gain greatly improves generation quality, reducing FID from 28.0 to 12.2 (estimated by linear interpolation between FID measurements at neighboring gains) on ImageNet-256 at NFE 25 without guidance.
Figures & tables
| Model | NFE | Unscaled FID | Lag-gain FID (interpolated) | -adjusted FID | Lowest measured FID |
|---|---|---|---|---|---|
| SiT-S/2 | 25 | 71.1 | 50.9 | 44.4 | 44.4 |
| 250 | 50.6 | 49.9 | 44.1 | 43.8 | |
| SiT-B/2 | 25 | 47.9 | 30.2 | 25.9 | 25.9 |
| 250 | 29.1 | 28.5 | 24.8 | 24.8 | |
| SiT-L/2 | 25 | 30.2 | 16.1 | 13.4 | 13.4 |
| 250 | 14.6 | 14.4 | 12.2 | 12.2 |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Question | Main finding | Where |
|---|---|---|
| Is scaling repairing underestimated velocities? | The population MSE field gives exact transport. The tested pretrained SiT field has an MSE-optimal scalar multiplier near one on genuine training-path states. | § 4 ; Appendices B , F.1 |
| Why can coarse sampling produce lag? | Coarse integration can miss intermediate refinement. The two-step Euler example shows how missing sample-specific image structure registers as lag. | § 7.1 ; Appendix D.3 |
| Can we measure this lag? | Comparing mean probe readings on generated and genuine states reveals systematic population time lag, which persists after global scale normalization. | § 5.1 ; Appendix D |
| How can scaling help states catch up? | A first-order argument explains how scaling advances lagging states toward their scheduled time; moving model time backward addresses the same mismatch. Controls show that much of the benefit requires the model to predict new velocities at the advanced states. | §§ 5.2 – 5.3 , 6 ; Appendices D.4 – D.5 , G |
| Can lag select a useful gain? | Minimizing measured lag selects gains without an image-quality metric and improves coarse-sampling FID across the tested ImageNet models. The CIFAR-10 model has little lag and favors gain one. | § 5.4 ; Appendix E |
| What determines the useful gain beyond lag correction? | The useful gain depends on solver accuracy, guidance, and the image metric. In unguided ImageNet experiments, one multiplier per model roughly relates lowest-FID gains to lag-based gains across NFEs. For unguided SiT-XL/2, sharpening predicts gains with estimated FID within 0.2 of the best measured value. | § 7 ; Appendices E.3 , H – I |
| Data use | Representation | Count | Source or reference |
|---|---|---|---|
| ImageNet generation | RGB image decoded from a latent | 50,000 generated images | Fixed ADM evaluator references |
| ImageNet probe | latent | 12,800 training; 3,200 calibration; 16,000 held out | Disjoint ImageNet training images |
| CIFAR-10 generation | Centered RGB pixels | 50,000 generated images | Official 1-Rectified Flow reference statistics |
| CIFAR-10 probe | Centered RGB pixels | 16,000 training; 3,200 calibration; 4,000 held out | CIFAR-10 training images |
| Sampler | Update | Budget | Experiments |
|---|---|---|---|
| Hybrid SDE | SiT SDE; stochastic Euler updates through , then one Mean update; the public setting uses ( Ma et al., 2024 ) | 25–250 calls | Changed-state response and direction controls; model-time correction; guidance |
| Uniform stochastic Euler | Equation 25 on a uniform grid ( Ma et al., 2024 ) | 25–250 calls | Population time lag and gain grids; changed-state response across model sizes |
| Deterministic Euler ODE | Equation 24 | 50 or 250 calls | ODE response; U-Net and CIFAR-10 transfer |
| Euler–Maruyama with Mean | LFM SDE with and one Mean update ( Dao et al., 2023 ) | 25–250 calls | ADM U-Net transfer |
| Velocity-derived SDE | Equation 27 | 25–250 calls | CIFAR-10 SDE gain curves and lag calibration |
| Stochastic Heun | Predictor–corrector update defined below | 50 calls | Euler–Heun comparison |
| Probe | Training states | Calibration states | Held-out states | Epochs | Held-out error |
|---|---|---|---|---|---|
| SiT, ImageNet | 12,800 | 3,200 | 16,000 | 30 | 0.0258 |
| ADM U-Net, ImageNet | 12,800 | 3,200 | 16,000 | 30 | 0.0250 |
| Rectified Flow, CIFAR-10 | 16,000 | 3,200 | 4,000 | 20 | 0.0141 |
| Reduction in | Probe change from reducing | Probe change from perpendicular perturbation | Difference [95% interval] |
|---|---|---|---|
| 0.025 | |||
| 0.050 |
| Method | FID |
|---|---|
| Unscaled sampling | 22.7 |
| Model-time correction | 16.7 |
| Velocity scaling, gain 1.075 | 13.2 |
| Model | NFE | Lag-based gain [95%] | Lowest-FID gain | Unscaled lag RMS | Scalar lag RMS | Lowest-FID lag RMS |
|---|---|---|---|---|---|---|
| SiT-S/2 | 25 | 1.150 | 25.8 | 6.0 | 23.9 | |
| 50 | 1.100 | 12.1 | 3.7 | 21.9 | ||
| 100 | 1.075 | 5.4 | 2.4 | 20.4 | ||
| 250 | 1.050 | 2.7 | 2.3 | 16.1 | ||
| SiT-B/2 | 25 | 1.125 | 24.9 | 5.4 | 16.5 | |
| 50 | 1.075 | 11.5 | 3.3 | 14.0 |
| NFE | Lowest-FID gain | FID | Next higher gain | FID | FID [95%] |
|---|---|---|---|---|---|
| 25 | 1.125 | 14.7 | 1.150 | 15.0 | |
| 50 | 1.075 | 13.2 | 1.100 | 13.3 | |
| 100 | 1.075 | 12.9 | 1.100 | 14.3 | |
| 250 | 1.050 | 12.8 | 1.075 | 13.5 |
| Model | NFE 25 | NFE 50 | NFE 100 | NFE 250 |
|---|---|---|---|---|
| SiT-S/2 | 76 | 55 | 32 | 11 |
| SiT-B/2 | 81 | 62 | 36 | 13 |
| SiT-L/2 | 84 | 66 | 36 | 9 |
| SiT-XL/2 | 82 | 63 | 42 | 19 |
| ADM U-Net | 84 | 67 | 43 | 24 |
| Mean | 81 | 63 | 38 | 15 |
| Model | NFE | Predicted gain | Lowest-FID gain | Signed error |
|---|---|---|---|---|
| SiT-S/2 | 50 | 1.107 | 1.100 | |
| 100 | 1.086 | 1.075 | ||
| 250 | 1.076 | 1.050 | ||
| SiT-B/2 | 50 | 1.082 | 1.075 | |
| 100 | 1.062 | 1.050 | ||
| 250 | 1.053 | 1.050 |
| Model | Full-gain error | Excess error | |
|---|---|---|---|
| SiT-S/2 | 1.071 | 0.014 | 0.040 |
| SiT-B/2 | 1.048 | 0.007 | 0.030 |
| SiT-L/2 | 1.049 | 0.007 | 0.032 |
| SiT-XL/2 | 1.068 | 0.014 | 0.037 |
| ADM U-Net | 1.044 | 0.008 | 0.034 |
| Paths per candidate | Mean absolute gain change |
|---|---|
| 128 | 3.30 |
| 256 | 2.23 |
| 512 | 1.37 |
| 1024 | 0.55 |
| Training seed | Gain | Unscaled lag RMS | Scalar lag RMS | Reduction [95%] |
|---|---|---|---|---|
| 2026072307 | 1.04 | 13.6 | 5.0 | |
| 2026072308 | 1.04 | 12.9 | 4.2 | |
| 2026072309 | 1.04 | 11.8 | 4.4 |
| NFE | Unscaled | Fitted schedule | Direct scalar |
|---|---|---|---|
| 25 | 0.0254 | 0.0046 | 0.0058 |
| 50 | 0.0121 | 0.0020 | 0.0039 |
| 100 | 0.0061 | 0.0026 | 0.0026 |
| 250 | 0.0037 | 0.0023 | 0.0029 |
| Dynamics | NFE | Unscaled [95%] | SSC [95%] |
|---|---|---|---|
| ODE | 25 | ||
| ODE | 250 | ||
| SDE | 25 | ||
| SDE | 250 |
| Dynamics | Gain | Unscaled | Scaled | Change [95%] |
|---|---|---|---|---|
| ODE | SSC | |||
| SDE | SSC | |||
| SDE | Fitted |
| Quantity | Euler | Heun | Contrast | 95% interval |
|---|---|---|---|---|
| Strong endpoint error | 0.208 | 0.148 | ||
| Population time lag RMS | 0.019 | 0.005 |
| Endpoint property | Paired set | NFE | Intervention | Unscaled FID | Intervention FID | FID reduction [95%] |
|---|---|---|---|---|---|---|
| Global radius | A | 50 | Velocity scaling during sampling | 13.8 | 7.6 | |
| A | 50 | Endpoint-only control | 13.8 | 13.8 | ||
| Variance in 256 directions | B | 100 | Velocity scaling during sampling | 11.5 | 6.8 | |
| C | 100 | Endpoint-only control | 11.6 | 11.6 | ||
| Global latent mean and covariance | C | 100 | Velocity scaling during sampling | 11.6 | 6.8 | |
| C | 100 | Endpoint-only control | 11.6 | 12.3 |
| Model | Sampler, NFE, gain | FID: unscaled frozen full | FID-reduction ratio [95%] |
|---|---|---|---|
| SiT-S/2 | Uniform stochastic Euler, 25, 1.07 | 24.9% | |
| SiT-B/2 | Uniform stochastic Euler, 25, 1.07 | 23.8% | |
| SiT-L/2 | Uniform stochastic Euler, 25, 1.07 | 23.6% | |
| SiT-XL/2 | Public hybrid SDE, 50, 1.05 | 22.4% | |
| REPA-SiT-XL/2 | Public hybrid SDE, 50, 1.05 | 22.7% | |
| SiT-XL/2 | ODE, 50, 1.05 | 53.9% |
| Comparator | Realized FID | Comparator FID | FID contrast [95%] |
|---|---|---|---|
| Unscaled | 8.5 | 13.8 | |
| Norm-matched | 8.5 | 14.7 | |
| Scrambled | 8.5 | 25.3 |
| Gain | FID | sFID | Precision | Recall |
|---|---|---|---|---|
| 1.000 (unscaled) | 28.0 | 26.2 | 0.565 | 0.630 |
| 1.072 | 12.7 | 8.3 | 0.653 | 0.649 |
| 1.100 | 10.0 | 5.6 | 0.675 | 0.639 |
| 1.150 (lowest-FID gain) | 8.7 | 6.7 | 0.685 | 0.622 |
| Solver | NFE | SiT [95%] | Gaussian [95%] |
|---|---|---|---|
| Euler–Maruyama | 25 | ||
| Euler–Maruyama | 50 | ||
| Euler–Maruyama | 100 | ||
| Euler–Maruyama | 250 | ||
| Stochastic Heun | 2000 |
| Solver | Gain | FID | FID [95%] |
|---|---|---|---|
| Euler NFE 250 | 0.950 | 21.7 | |
| 0.975 | 5.9 | ||
| 1.000 | 3.0 | ||
| 1.025 | 8.3 | ||
| 1.050 | 16.6 | ||
| 1.075 | 25.3 |
| CFG | [95%] | Pointwise coefficient [95%] | Lowest-FID gain |
|---|---|---|---|
| 1.75 | 0.0258 [0.0252, 0.0264] | 0.0251 [0.0245, 0.0256] | 1.000 |
| 2.00 | 0.0279 [0.0272, 0.0286] | 0.0271 [0.0265, 0.0276] | 0.975 |
| 2.25 | 0.0298 [0.0292, 0.0305] | 0.0289 [0.0283, 0.0294] | 0.950 |
| Guidance | Lag-based gain [95%] | Unscaled lag RMS | Scalar lag RMS | Lowest-FID gain |
|---|---|---|---|---|
| 1.50 | 9.8 | 2.2 | — | |
| 1.75 | 8.6 | 1.5 | 1.000 | |
| 2.00 | 7.6 | 2.1 | 0.975 | |
| 2.25 | 7.2 | 3.5 | 0.950 | |
| 2.50 | 7.2 | 4.9 | — | |
| 4.00 | 14.8 | 14.7 | — |
| CFG | Precision gain 1 | Precision gain 0.95 | precision [95%] |
|---|---|---|---|
| 1.75 | 0.845 | 0.797 | |
| 2.00 | 0.871 | 0.837 | |
| 2.25 | 0.888 | 0.862 |