Global Transport Couplings for Classifier-Free Guided Flows
Authors: Katarina Petrović, Zander W. Blasingame, Danyal Rehman, İsmail İlkan Ceylan, Michael Bronstein, Stephen Y. Zhang, Lazar Atanackovic, Alexander Tong
Organizations: AITHYRA · University of Oxford · Mila - Québec AI Institute · Université de Montréal · Massachusetts Institute of Technology · Technische Universität Wien · Flatiron Institute · University of Alberta · Alberta Machine Intelligence Institute · Canada CIFAR AI Chair
Optimal-transport couplings have been shown to reduce training variance in unconditional flow models, but their role in conditional generation remains unclear. A natural approach constructs separate couplings for each condition, but this is impractical for large or continuous conditioning spaces found in modern image foundation models. We introduce Global Transport (GT), a global class-agnostic optimal-transport coupling, computed without class labels. GT can associate different conditions with different regions of the source noise, and consequently worsens performance without guidance. However, when combined with classifier-free guidance (CFG), GT consistently improves generation across domains, model scales, and sampling budgets. This reversal suggests that couplings for conditional flows should be evaluated both empirically and theoretically under the guided flow used at inference, rather than on unguided generation. We evaluate GT over both discrete class and continuous text conditioned image generation across model scales, and investigate how coupling choice alters guided trajectories. These results identify coupling design in the guided flow setting as a simple training time axis to improve performance without modifying existing architectures, samplers, or guidance mechanisms.
Figures & tables
Figure 1 : 2-dimensional depiction of the “yo-yo” effect. This effect pushes samples away from the center at high guidance scales for the independent and class-conditional OT couplings, degrading generative performance. Global Transport fixes this and is stable under increasing guidance w .
Figure 2 : Density over time for Independent, Class conditional, and GT couplings. Independent “yoyo”s, contracting towards the class origin then expanding outwards, whereas GT is stable.
Figure 3 : Guided dynamics on the 40-GMM across guidance scales w . From left: state norm along guided trajectories; integrated guided velocity norm; integrated prediction gap ∥vc−v∅∥ ; unconditional and conditional material acceleration ∥at[v]∥ .
Table 1 : Left: FID ↓ comparison against baselines. † denotes results obtained with guidance interval ( Kynkäänniemi et al., 2024a ) . Right: Curated class-conditional samples from SiT-XL/2 + GT on ImageNet-256.
Figure 4 : Left: Guidance scale w sweep for SiT-B/2 (Euler-64 steps) FID ↓ under a tuned guidance interval [0,0.7] . Right: Class-conditional samples for class 339 ( sorrel ) from SiT-L/2 (Euler-64 steps, w=2.5 ) under independent, class-conditional, and GT transport.
Model
Coupling
FID ↓
FD DINOv2 ↓
16
32
64
16
32
64
Independent
5.88
5.33
5.16
137.3
132.9
132.3
Class-cond. OT
5.79
5.20
5.08
136.4
132.2
131.4
GT + SD-OT
5.11
4.90
4.85
199.6
194.6
184.7
SiT-B/2 (130M)
GT
5.15
4.35
4.15
123.6
118.0
116.4
Independent
5.50
5.74
2.95
91.1
84.6
83.3
Table 2 : Comparison of different couplings across various NFEs and model sizes reported in FID and FD DINOv2 . We swept for the optimal guidance strength w∗ for each setting and report the best result.
Teacher (NFE)
Student (NFE)
Model
Coupling
16
32
64
1
2
4
Independent
5.88
5.33
5.16
5.82
5.19
5.19
Class-cond. OT
5.79
5.20
5.08
5.83
5.18
5.25
GT (from GT teacher)
5.15
4.35
4.15
5.66
4.17
4.14
DMF-B/2
GT (from ind. teacher)
5.88
5.32
5.22
5.88
4.21
4.18
Independent
5.50
5.74
2.95
3.31
2.82
2.68
Table 3: Left : FID ↓ on ImageNet-256 for DMF distilled with independent, class-conditional and GT couplings, at B/2 and L/2 scale. Right : Curated samples from GT DMF-B/2 at 4-NFE.
Independent
Class-cond. OT
GT
GT +SD-OT
w
FID
FD DINOv2
NFE
FID
FD DINOv2
NFE
FID
FD DINOv2
NFE
FID
FD DINOv2
NFE
1.0
25.01
586.47
50.83
25.23
589.00
47.73
29.65
618.75
49.17
42.39
772.55
38.00
2.0
5.17
206.41
60.99
5.08
208.96
60.56
4.06
220.08
61.54
7.56
341.00
50.00
3.0
11.13
138.30
76.39
10.98
137.98
75.84
8.77
130.56
73.45
4.79
226.95
62.80
4.0
15.77
132.94
90.04
15.56
132.06
89.86
13.37
116.34
85.63
5.47
194.00
77.09
Table 4: FID ↓ , FD DINOv2 ↓ , and NFE ↓ for SiT-B/2 over coupling plans and guidance scales w .
Figure 5 : Left: FID ↓ and FD DINOv2 ↓ for text-conditioned ImageNet-256 across guidance scales w . Right: Text-conditioned ImageNet samples from independent and GT coupling models at w=4 .
PBMC3K
Dentate gyrus
HLCA
MMD ( ↓ )
WD ( ↓ )
MMD ( ↓ )
WD ( ↓ )
MMD ( ↓ )
WD ( ↓ )
c-CFGen
0.45 ± 0.00
11.17 ± 0.02
0.06 ± 0.00
7.26 ± 0.03
0.06 ± 0.00
5.06 ± 0.00
c-CFGen-linear
0.41 ± 0.00
10.27 ± 0.08
0.06 ± 0.00
6.69 ± 0.01
0.07 ± 0.00
4.89 ± 0.01
scDiffusion
0.67 ± 0.06
12.11 ± 0.16
0.06 ± 0.00
5.89 ± 0.01
0.12 ± 0.00
5.42 ± 0.01
scVI
0.58 ± 0.01
13.39 ± 0.16
0.11 ± 0.00
7.34 ± 0.03
0.13 ± 0.00
6.39 ± 0.01
GT (ours)
0.39 ± 0.00
9.80 ± 0.02
0.05 ± 0.00
6.61 ± 0.01
0.07 ± 0.00
4.98 ± 0.00
Table 5 : Comparison of GT with single-cell generative models on distribution-matching metrics (RBF-kernel MMD and 2-Wasserstein distance), averaged over three seeds.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : 40-GMM ground truth data. Left: Mode centers by class, Right: Ground truth samples.
Figure 7 : Generated trajectories for Independent, Class-conditional and GT on 40-GMM.
Figure 8 : Velocity field heat maps for conditional vc and unconditional v∅ for Independent, Class-conditional and GT transport.
Figure 9 : Additional 40-GMM samples across a range of guidance scales.
B/2
L/2
XL/2
Backbone
Resolution
256×256
256×256
256×256
Params (M)
130
458
675
FLOPS (G)
23.1
80.7
118.6
Hidden dim.
768
1024
1152
Heads
12
16
16
Appendix
Table 6 : ImageNet-256 implementation details across model scales.
Optimizer
AdamW
Batch size
256
Learning rate
1e-4
Adam (β1,β2)
(0.9,0.95)
Adam ϵ
1e-8
Weight decay
0.0
EMA decay rate
0.9999
Appendix
Table 7 : Optimization hyperparameters, shared across all model scales and both training stages.
Coupling
ms / step
↑ cost
Independent
155.5
–
Class-cond. OT
157.9
↑ 1.6%
GT
160.4
↑ 3.1%
Appendix
Table 8: Compute cost per optimizer step of SiT-B/2 training on ImageNet-256 for the three coupling plans. The second column is the increase over the independent coupling.
Figure 10 : Sweeping βm for class-conditional OT.
SiT-B/2
Coupling
Metric
w=1.0
1.5
2.0
2.5
3.0
3.5
4.0
5.0
6.0
8.0
10.0
Independent
FID
26.44
6.68
5.16
7.93
11.04
13.65
15.76
18.63
20.42
22.40
23.28
FD DINOv2
605.11
337.16
212.66
159.87
139.10
132.46
132.27
139.64
149.79
169.51
189.02
Class-cond. OT
FID
26.72
6.70
5.08
7.81
10.90
13.47
15.59
18.44
20.28
22.25
23.12
FD DINOv2
607.86
340.46
215.04
160.77
139.17
131.64
131.41
138.27
148.37
168.91
189.17
GT
FID
31.16
8.13
4.15
5.88
8.72
11.32
13.41
16.51
18.62
20.81
21.62
Appendix
Table 9 : FID ↓ and FD DINOv2 ↓ vs. guidance scale w for independent, class-conditional OT and GT couplings across SiT scales (64 Euler steps), under the full guidance interval [0,1] .
Figure 11 : SiT-XL/2 curated samples for class 339 sorrel across guidance scales w and coupling choices. Under GT the sample object characteristics remain stable, while for independent and class-conditional object changes its composition, pose, background at higher w .
Figure 12 : Measuring yo-yo effect and prediction gap across trajectory using SiT-XL/2 (64 steps).
Full
GI [0, 0.7]
Model
Coupling
w=1.0
1.5
2.0
2.5
3.0
4.0
1.5
2.0
2.5
3.0
4.0
SiT-B/2
Independent
26.44
6.68
5.16
7.93
11.04
15.76
11.93
5.95
3.95
3.72
5.07
Class-cond. OT
26.72
6.70
5.08
7.81
10.90
15.59
12.13
6.00
3.92
3.68
5.00
GT
31.16
8.13
4.15
5.88
8.72
13.41
15.63
8.03
4.71
3.52
3.70
SiT-L/2
Independent
13.85
2.95
5.85
10.10
13.53
17.89
4.53
2.35
2.62
3.65
6.00
GT
17.72
2.96
3.88
7.38
10.66
15.17
6.64
2.96
2.17
2.48
4.02
Appendix
Table 10 : FID ↓ across model scales and couplings, under the full guidance interval and tuned interval [0,0.7] .
Parameter
Autoencoder
Latent Flow Matching
Batch size
256
256 (64 for PBMC3K)
Epochs
300
1500
Optimizer
AdamW
AdamW
Learning rate
1×10−3
1×10−4
Weight decay
1×10−5
1×10−6
Gradient clipping
1.0
1.0
Appendix
Table 11: Training configuration for uni-modal generation with CFGen.
School of Mathematical Sciences, Shanghai Jiao Tong University, Shanghai, China · Institute of Natural Sciences, MOE-LSC, Shanghai Jiao Tong University, Shanghai, China · Qing Yuan Research Institute, Shanghai Jiao Tong University, Shanghai, China +1