We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation. Full observation recovers deterministic encoding, no observation recovers the latent marginal, and intermediate subsets define conditional latent inference, all within a single conditional flow model. A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently. To learn these conditional distributions from incomplete training examples, we extend conditional flow matching through self-distillation: predictions conditioned on richer available observations supervise the same model conditioned on smaller subsets at the same intermediate latent state. When the richer-evidence trajectory follows the exact conditional flow, this provides the same expected learning signal as full-target denoising. Across PolyMNIST-D-Q, FFHQ64, and image-text-audio, MUNITE achieves competitive or better generation quality and source-target alignment, with higher joint-generation coherence. In particular, it attains the highest coherence in all one-to-many and unconditional image-text-audio comparisons, showing the effectiveness of unified latent inference across diverse multimodal settings.
Figures & tables
Figure 1: Latent inference and decoding. Left: latent distributions under different observation subsets. Right: factorized decoding from a latent value.
Figure 2: MUNITE training objectives. (a) Reconstruction learns a shared representation, using target detaching to discourage copying modality-private information. (b) Self-distillation uses detached predictions conditioned on more modalities to supervise subset-conditioned predictions at the same noisy latent state, with both branches using the same model. (c) An auxiliary contrastive loss encourages semantic alignment between complementary modality subsets using in-batch negatives.
Figure 3: Latent inference architecture. A shared Transformer performs latent inference under arbitrary observation subsets by aggregating modality information in latent registers. Keeping modality streams separate allows reconstruction gradients to be stopped on the target’s key/value paths to the registers (slashes), while retaining its information in the forward pass.
Route
Input → output
Metric
MUNI
CFM
DFM
MUNITE
PolyMNIST-D-Q
Uncond.
All five modalities
FD ↓
3.0407
71.6626
19.9587
1.5230
Digit Coh. (all) ↑
0.5799
0.0237
0.2973
0.9826
Quad. Coh. (all) ↑
0.9041
0.9902
0.8222
1.0000
1→1
d→mi
Digit Acc. ↑
0.9843
0.2402
0.7224
0.9941
q→mi
Quad. Acc. ↑
1.0000
0.9936
0.9372
1.0000
Table 1: Results on PolyMNIST-D-Q (top) and FFHQ64 (bottom). Acc. and Coh. denote accuracy and coherence, d and q denote digit and quadrant, and scores for mi are averaged over the three views. FD uses verifier features, and Normal Err. is 1−cos . Best results are in bold.
1→N coherence
Unconditional coherence
Method
T→(I,A)
I→(T,A)
A→(T,I)
CLIP
CLAP
AIS
AIS
CLAP
CLIP
(T,I)
(T,A)
(I,A)
CoDi
63.866
7.418
23.414
—
—
—
OmniFlow
77.027
14.781
22.697
21.17
14.23
50.95
MUNI
81.558
32.237
25.850
26.949
28.467
81.989
CFM
74.702
21.695
24.034
23.622
19.896
74.208
Table 2: Image–text–audio coherence between jointly generated modalities. T , I , and A denote text, image, and audio. OmniFlow’s unconditional scores are from Yeo et al. (2026, Table 4) , and dashes denote unsupported unconditional generation. Higher is better, and best results are in bold.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Width
Depth
Time cond.
PolyMNIST-D-Q
MUNITE
256
7
FiLM
CFM
256
16
FiLM
DFM
256
24
N/A
FFHQ64
MUNITE
512
11
FiLM
Appendix
Table 3: Transformer backbones.
Observed subset
Rows
%
m0,m1,m2
55,000
25.0
m0,m1
18,333
8.3
m0,m2
18,334
8.3
m1,m2
18,333
8.3
m0
36,667
16.7
m1
36,666
16.7
Appendix
Table 4: Training observations grouped by available modality subset. For PolyMNIST-D-Q and FFHQ64, categorical labels are observed in all groups. T , I , and A denote text, image, and audio, respectively.
Table 8
Task
FlowBind
MUNI
CFM
DFM
MUNITE
PolyMNIST-D-Q
—
32.09
20.09
20.57
19.33
FFHQ64
—
130.99
105.11
107.27
106.64
Image–text–audio
567.97
328.90
281.21
281.16
278.72
Appendix
Table 7: Trainable parameters in millions, including learned encoders, priors, and decoders. Frozen foundation models and the frozen VQ-VAEs of DFM, which have 13.27M parameters on PolyMNIST-D-Q and FFHQ64, are excluded.
Method
Epochs
Updates
Batch
LR
PolyMNIST-D-Q
MUNI, MUNITE
100
85,600
256
1.5×10−4
CFM, DFM
200
171,200
256
1.5×10−4
FFHQ64
MUNI, MUNITE
500
252,000
128
1.5×10−4
CFM, DFM
1,000
504,000
128
1.5×10−4
Appendix
Table 8: Training budgets. Batch sizes are global, and all methods use a constant learning rate. FlowBind’s epoch count assumes 702 updates per epoch.
Method
Conditional
Unconditional
Decoding
PolyMNIST-D-Q
MUNITE
Euler 50
Euler 50
Image Euler 50, CFG 1.5
MUNI
Posterior sampling
Euler 50
Image Euler 50, CFG 1.5
CFM
Euler 150
Euler 150
—
DFM
Random-order AR
Random-order AR
VQ decoder
FFHQ64
Appendix
Table 9: Inference and decoding methods. Euler step counts equal the number of function evaluations, CFG denotes classifier-free guidance, and AR denotes autoregressive sampling. Rendering of image–text–audio features is not included.
TV ↓
Coherence ↑
Route
MUNI
MUNITE
Ref.
MUNI
MUNITE
Indep.
Chance
d→(m0,m1,m2)
0.029
0.029
0.022
0.2528
1.0000
0.2500
0.25
q→(m0,m1,m2)
0.109
0.034
0.024
0.1074
0.9904
0.1014
0.10
Appendix
Table 10: Missing attribute in PolyMNIST-D-Q one-to-many generation. TV is the total variation distance between the missing-label distribution of a generated view and the uniform data conditional, averaged over input labels and views; Ref. is its expected value for the same number of samples drawn from that conditional. Indep. is MUNITE with m1 and m2 decoded from separately drawn latents, keeping m0 and all decoder noise fixed. Chance is the coherence of independent uniform views.
Image: FID ↓
Audio: FAD ↓
Text: CIDEr ↑
Method
T→I
A→I
T→A
I→A
I→T
A→T
CoDi
24.97
53.60
8.67
13.92
11.97
7.70
OmniFlow
20.52
97.56
4.08
4.87
26.60
31.42
MUNI
15.615
22.786
3.507
2.246
44.400
51.794
CFM
13.455
28.971
3.564
2.105
13.085
19.729
DFM
14.400
28.374
3.670
1.991
11.263
20.909
Appendix
Table 11: Image–text–audio one-to-one fidelity on the evaluation sets of Table 6 , grouped by target modality. CoDi and OmniFlow scores are taken from Yeo et al. (2026, Table 8) . Bold marks the best score in each column.
Text–image: CLIP
Text–audio: CLAP
Image–audio: AIS
Method
T→I
I→T
T→A
A→T
I→A
A→I
CoDi
29.70
26.05
16.72
32.72
59.45
84.47
OmniFlow
30.76
27.48
30.07
44.74
76.52
64.22
MUNI
29.718
28.373
30.596
40.102
90.390
94.773
CFM
27.710
24.918
25.127
32.094
82.888
80.842
DFM
27.321
24.353
24.883
32.262
79.903
79.933
Appendix
Table 12: Image–text–audio one-to-one alignment. CoDi and OmniFlow scores are taken from Yeo et al. (2026, Table 9) . Higher is better, and bold marks the best score in each column.
(I,A)→T
(T,A)→I
(T,I)→A
Method
CLIP
CLAP
CLIP
AIS
CLAP
AIS
CoDi
24.05
33.72
24.98
85.52
11.06
65.31
OmniFlow
24.73
36.26
26.41
81.51
13.50
63.55
MUNI
27.922
36.695
27.056
85.168
30.524
82.621
CFM
25.221
31.889
25.712
80.305
24.907
76.098
DFM
24.855
33.259
25.398
80.347
25.141
75.846
Appendix
Table 13: Image–text–audio many-to-one alignment on 975 triplets, with each generated target scored against both source modalities. CoDi and OmniFlow scores are taken from Yeo et al. (2026, Table 3) . Higher is better, and bold marks the best score in each column.
Route
Input → output
Metric
Control
No target detach
No contrastive
Full-only distillation
1→1
T→I
FID ↓
15.939
16.700
17.308
15.230
CLIP ↑
29.007
28.543
27.750
29.593
I→T
CIDEr ↑
37.409
35.414
35.741
39.667
CLIP ↑
27.526
27.261
26.919
27.882
T→A
FAD ↓
3.059
3.145
3.321
3.234
CLAP ↑
30.695
30.458
31.461
30.540
Appendix
Table 14: Component ablations on image–text–audio. Control is the full MUNITE model of Tables 2 and 11 – 13 . Each variant removes one component and otherwise keeps the control’s architecture, optimization, and evaluation protocol. The 1→N and unconditional rows report coherence among jointly generated modalities. Best results in each row are in bold.
Figure 4: PolyMNIST-D-Q single-target generation. Icons show the observed labels: a digit, the conditioned quadrant shaded in a 2×2 grid, or the digit placed in that quadrant. Each column generates the view listed under its input.
Figure 5: Unconditional PolyMNIST-D-Q co-generation. Each group is one joint sample of the three views and both labels, and its icon shows the generated digit in the generated quadrant.
Figure 6: PolyMNIST-D-Q one-to-many generation. Each row generates all three views from a digit (a) or a quadrant (b), so the views must agree on the unobserved attribute.
Figure 7: FFHQ64 RGB generation from age (a), gender (b), or both attributes (c).
Figure 8: Unconditional FFHQ64 co-generation of all five modalities. The generated gender and age labels of each joint sample appear below its maps.
Figure 9: FFHQ64 many-to-many generation of RGB, segmentation, and surface normals from age and gender. The three maps of each group come from one joint sample.
Figure 10: Image–text–audio one-to-one and many-to-one generation of images. Audio inputs are shown as waveforms.
Figure 11: Image–text–audio one-to-one and many-to-one generation of text. Audio inputs are shown as waveforms.
Figure 12: Unconditional image–text–audio co-generation by MUNITE. Each column decodes one latent sample into text, image, and audio.
Figure 13: One-to-many image–text–audio generation by MUNITE. In each column, the two outputs are decoded independently from one latent sample inferred from the input.