Global Transport Couplings for Classifier-Free Guided Flows
Authors: Katarina Petrović, Zander W. Blasingame, Danyal Rehman, İsmail İlkan Ceylan, Michael Bronstein, Stephen Y. Zhang, Lazar Atanackovic, Alexander Tong
Organizations: AITHYRA · University of Oxford · Mila - Québec AI Institute · Université de Montréal · Massachusetts Institute of Technology · Technische Universität Wien · Flatiron Institute · University of Alberta · Alberta Machine Intelligence Institute · Canada CIFAR AI Chair
Optimal-transport couplings have been shown to reduce training variance in unconditional flow models, but their role in conditional generation remains unclear. A natural approach constructs separate couplings for each condition, but this is impractical for large or continuous conditioning spaces found in modern image foundation models. We introduce Global Transport (GT), a global class-agnostic optimal-transport coupling, computed without class labels. GT can associate different conditions with different regions of the source noise, and consequently worsens performance without guidance. However, when combined with classifier-free guidance (CFG), GT consistently improves generation across domains, model scales, and sampling budgets. This reversal suggests that couplings for conditional flows should be evaluated both empirically and theoretically under the guided flow used at inference, rather than on unguided generation. We evaluate GT over both discrete class and continuous text conditioned image generation across model scales, and investigate how coupling choice alters guided trajectories. These results identify coupling design in the guided flow setting as a simple training time axis to improve performance without modifying existing architectures, samplers, or guidance mechanisms.
Figures & tables
Figure 1 : 2-dimensional depiction of the “yo-yo” effect. This effect pushes samples away from the center at high guidance scales for the independent and class-conditional OT couplings, degrading generative performance. Global Transport fixes this and is stable under increasing guidance w .
Figure 2 : Density over time for Independent, Class conditional, and GT couplings. Independent “yoyo”s, contracting towards the class origin then expanding outwards, whereas GT is stable.
Figure 3 : Guided dynamics on the 40-GMM across guidance scales w . From left: state norm along guided trajectories; integrated guided velocity norm; integrated prediction gap ∥vc−v∅∥ ; unconditional and conditional material acceleration ∥at[v]∥ .
Table 1 : Left: FID ↓ comparison against baselines. † denotes results obtained with guidance interval ( Kynkäänniemi et al., 2024a ) . Right: Curated class-conditional samples from SiT-XL/2 + GT on ImageNet-256.
Figure 4 : Left: Guidance scale w sweep for SiT-B/2 (Euler-64 steps) FID ↓ under a tuned guidance interval [0,0.7] . Right: Class-conditional samples for class 339 ( sorrel ) from SiT-L/2 (Euler-64 steps, w=2.5 ) under independent, class-conditional, and GT transport.
Model
Coupling
FID ↓
FD DINOv2 ↓
16
32
64
16
32
64
Independent
5.88
5.33
5.16
137.3
132.9
132.3
Class-cond. OT
5.79
5.20
5.08
136.4
132.2
131.4
GT + SD-OT
5.11
4.90
4.85
199.6
194.6
184.7
SiT-B/2 (130M)
GT
5.15
4.35
4.15
123.6
118.0
116.4
Independent
5.50
5.74
2.95
91.1
84.6
83.3
Table 2 : Comparison of different couplings across various NFEs and model sizes reported in FID and FD DINOv2 . We swept for the optimal guidance strength w∗ for each setting and report the best result.
Teacher (NFE)
Student (NFE)
Model
Coupling
16
32
64
1
2
4
Independent
5.88
5.33
5.16
5.82
5.19
5.19
Class-cond. OT
5.79
5.20
5.08
5.83
5.18
5.25
GT (from GT teacher)
5.15
4.35
4.15
5.66
4.17
4.14
DMF-B/2
GT (from ind. teacher)
5.88
5.32
5.22
5.88
4.21
4.18
Independent
5.50
5.74
2.95
3.31
2.82
2.68
Table 3: Left : FID ↓ on ImageNet-256 for DMF distilled with independent, class-conditional and GT couplings, at B/2 and L/2 scale. Right : Curated samples from GT DMF-B/2 at 4-NFE.
Independent
Class-cond. OT
GT
GT +SD-OT
w
FID
FD DINOv2
NFE
FID
FD DINOv2
NFE
FID
FD DINOv2
NFE
FID
FD DINOv2
NFE
1.0
25.01
586.47
50.83
25.23
589.00
47.73
29.65
618.75
49.17
42.39
772.55
38.00
2.0
5.17
206.41
60.99
5.08
208.96
60.56
4.06
220.08
61.54
7.56
341.00
50.00
3.0
11.13
138.30
76.39
10.98
137.98
75.84
8.77
130.56
73.45
4.79
226.95
62.80
4.0
15.77
132.94
90.04
15.56
132.06
89.86
13.37
116.34
85.63
5.47
194.00
77.09
Table 4: FID ↓ , FD DINOv2 ↓ , and NFE ↓ for SiT-B/2 over coupling plans and guidance scales w .
Figure 5 : Left: FID ↓ and FD DINOv2 ↓ for text-conditioned ImageNet-256 across guidance scales w . Right: Text-conditioned ImageNet samples from independent and GT coupling models at w=4 .
PBMC3K
Dentate gyrus
HLCA
MMD ( ↓ )
WD ( ↓ )
MMD ( ↓ )
WD ( ↓ )
MMD ( ↓ )
WD ( ↓ )
c-CFGen
0.45 ± 0.00
11.17 ± 0.02
0.06 ± 0.00
7.26 ± 0.03
0.06 ± 0.00
5.06 ± 0.00
c-CFGen-linear
0.41 ± 0.00
10.27 ± 0.08
0.06 ± 0.00
6.69 ± 0.01
0.07 ± 0.00
4.89 ± 0.01
scDiffusion
0.67 ± 0.06
12.11 ± 0.16
0.06 ± 0.00
5.89 ± 0.01
0.12 ± 0.00
5.42 ± 0.01
scVI
0.58 ± 0.01
13.39 ± 0.16
0.11 ± 0.00
7.34 ± 0.03
0.13 ± 0.00
6.39 ± 0.01
GT (ours)
0.39 ± 0.00
9.80 ± 0.02
0.05 ± 0.00
6.61 ± 0.01
0.07 ± 0.00
4.98 ± 0.00
Table 5 : Comparison of GT with single-cell generative models on distribution-matching metrics (RBF-kernel MMD and 2-Wasserstein distance), averaged over three seeds.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : 40-GMM ground truth data. Left: Mode centers by class, Right: Ground truth samples.
Figure 7 : Generated trajectories for Independent, Class-conditional and GT on 40-GMM.
Figure 8 : Velocity field heat maps for conditional vc and unconditional v∅ for Independent, Class-conditional and GT transport.
Figure 9 : Additional 40-GMM samples across a range of guidance scales.
B/2
L/2
XL/2
Backbone
Resolution
256×256
256×256
256×256
Params (M)
130
458
675
FLOPS (G)
23.1
80.7
118.6
Hidden dim.
768
1024
1152
Heads
12
16
16
Appendix
Table 6 : ImageNet-256 implementation details across model scales.
Optimizer
AdamW
Batch size
256
Learning rate
1e-4
Adam (β1,β2)
(0.9,0.95)
Adam ϵ
1e-8
Weight decay
0.0
EMA decay rate
0.9999
Appendix
Table 7 : Optimization hyperparameters, shared across all model scales and both training stages.
Coupling
ms / step
↑ cost
Independent
155.5
–
Class-cond. OT
157.9
↑ 1.6%
GT
160.4
↑ 3.1%
Appendix
Table 8: Compute cost per optimizer step of SiT-B/2 training on ImageNet-256 for the three coupling plans. The second column is the increase over the independent coupling.
Figure 10 : Sweeping βm for class-conditional OT.
SiT-B/2
Coupling
Metric
w=1.0
1.5
2.0
2.5
3.0
3.5
4.0
5.0
6.0
8.0
10.0
Independent
FID
26.44
6.68
5.16
7.93
11.04
13.65
15.76
18.63
20.42
22.40
23.28
FD DINOv2
605.11
337.16
212.66
159.87
139.10
132.46
132.27
139.64
149.79
169.51
189.02
Class-cond. OT
FID
26.72
6.70
5.08
7.81
10.90
13.47
15.59
18.44
20.28
22.25
23.12
FD DINOv2
607.86
340.46
215.04
160.77
139.17
131.64
131.41
138.27
148.37
168.91
189.17
GT
FID
31.16
8.13
4.15
5.88
8.72
11.32
13.41
16.51
18.62
20.81
21.62
Appendix
Table 9 : FID ↓ and FD DINOv2 ↓ vs. guidance scale w for independent, class-conditional OT and GT couplings across SiT scales (64 Euler steps), under the full guidance interval [0,1] .
Figure 11 : SiT-XL/2 curated samples for class 339 sorrel across guidance scales w and coupling choices. Under GT the sample object characteristics remain stable, while for independent and class-conditional object changes its composition, pose, background at higher w .
Figure 12 : Measuring yo-yo effect and prediction gap across trajectory using SiT-XL/2 (64 steps).
Full
GI [0, 0.7]
Model
Coupling
w=1.0
1.5
2.0
2.5
3.0
4.0
1.5
2.0
2.5
3.0
4.0
SiT-B/2
Independent
26.44
6.68
5.16
7.93
11.04
15.76
11.93
5.95
3.95
3.72
5.07
Class-cond. OT
26.72
6.70
5.08
7.81
10.90
15.59
12.13
6.00
3.92
3.68
5.00
GT
31.16
8.13
4.15
5.88
8.72
13.41
15.63
8.03
4.71
3.52
3.70
SiT-L/2
Independent
13.85
2.95
5.85
10.10
13.53
17.89
4.53
2.35
2.62
3.65
6.00
GT
17.72
2.96
3.88
7.38
10.66
15.17
6.64
2.96
2.17
2.48
4.02
Appendix
Table 10 : FID ↓ across model scales and couplings, under the full guidance interval and tuned interval [0,0.7] .
Parameter
Autoencoder
Latent Flow Matching
Batch size
256
256 (64 for PBMC3K)
Epochs
300
1500
Optimizer
AdamW
AdamW
Learning rate
1×10−3
1×10−4
Weight decay
1×10−5
1×10−6
Gradient clipping
1.0
1.0
Appendix
Table 11: Training configuration for uni-modal generation with CFGen.
Classifier-free guidance (CFG) improves conditional generation in Flow Matching, but strong guidance can distort the generated distribution and reduce diversity. We provide a geometric account of this behavior by viewing Flow Matching as a time-varying gradient flow and characterizing how CFG reshapes its underlying potential. This view explains how stronger alignment can be accompanied by mean displacement and trajectory concentration, and motivates controlling guidance through the model-implied terminal posterior mean. We therefore propose Posterior-Mean-Capped CFG (PMC-CFG), a training-free, per-sample method that adaptively retains the strongest feasible guidance without additional network evaluations. Experiments on synthetic and large-scale image-generation benchmarks show that PMC-CFG limits guidance-induced distortion and concentration while improving the alignment--diversity trade-off, with particularly strong benefits when nominal guidance is large.
Jishen Peng, Zheng Ma
School of Mathematical Sciences, Shanghai Jiao Tong University, Shanghai, China · Institute of Natural Sciences, MOE-LSC, Shanghai Jiao Tong University, Shanghai, China · Qing Yuan Research Institute, Shanghai Jiao Tong University, Shanghai, China +1
Flow matching models learn to transport samples from a simple prior distribution to a complex data distribution. When prior-data pairs are coupled via optimal transport (OT), the learned trajectories are straight and non-crossing, enabling fast, even single-step, generation. However, computing the OT coupling in high dimensions is intractable, and existing methods attempt to solve the OT problem, at the cost of persistent bias or significant overhead. Rather than solving for the OT coupling, we reformulate the problem. Once the prior is treated as a design choice rather than a fixed input, the OT coupling between prior and data is no longer unique. Many priors admit an OT-optimal identity coupling to the data, leaving us free to choose one that is also tractable to sample. We identify low-frequency projection of natural images as such a choice. The identity coupling between data and its low-frequency representation is empirically OT-optimal, the prior is structured enough to be sampled by a lightweight model at inference, and the remaining flow-matching task reduces to synthesizing high-frequency detail. Interpolating the prior with Gaussian noise further improves generation quality while preserving the OT coupling. The approach requires no modifications to the flow model itself, and integrates naturally with latent-space models, classifier-free guidance, and one-step generation frameworks. Across all benchmarks, our method reduces trajectory curvature by more than 2× compared to existing flow matching methods, yielding better generation quality in the few-step regime.
Diffusion and flow-matching models are typically trained by corrupting data through independently sampled Gaussian noise. While simple and scalable, this forward process induces arbitrary data-noise couplings, forcing the network to learn high-curvature transports between unrelated endpoints. Existing optimal-transport methods reduce this burden by reassigning fixed noise samples to data, but the source noise distribution itself remains passive. To address this, we introduce Contrastive Noise Alignment (CNA), a training-time method that creates dynamic, contrastive couplings by optimizing the noise representations directly. By modeling the noise batch as an interacting particle system, CNA employs a cross-modal InfoNCE objective to align noise particles with their paired data targets. To prevent spatial collapse, this alignment is regularized using an angular entropy term and a radial norm penalty. We show theoretically that this equilibrium asymptotically preserves Gaussian structures, maintaining tractability during inference. Empirically, CNA improves the alignment between noise and data, reduces flow curvature, and provides better generation quality with fewer required sampling steps. For few-step, pixel-space generation (2-4 NFEs), CNA reduces FID by over 50% compared to standard rectified flow, and by at least 24% against Optimal Transport baselines.