We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The latter requires modality-dependent objectives and sampling procedures. Fully continuous modeling avoids these trade-offs and enables a shared generative process, but remains underexplored for multimodal pretraining. Multimodal Flow introduces a unified continuous architecture that integrates multimodal continuous representations with a shared chunk-causal flow backbone. It organizes text blocks and images as ordered continuous hyperchunks, preserving textual token order and visual spatial structure. The backbone learns a single vector field over these hyperchunks through Flow Matching. Joint attention enables cross-modal interaction, while modality-specific feed-forward networks process each modality. The model predicts multiple target chunks in parallel during training and generates hyperchunks sequentially at inference. We instantiate MF-1 and pretrain it on multimodal data. Across 0.6B, 1.2B, and 1.6B scales, continued pretraining consistently improves multimodal modeling. With only 150B pretraining tokens, MF-1 achieves an average score of 82.8 across GenEval and DPG-Bench and 75.3 across VQAv2, MMBench, and POPE, remaining competitive with unified models trained on substantially more data. Under matched data, optimization, and parameter budgets, Multimodal Flow further outperforms representative hybrid and discrete models. These results establish continuous chunk-based embedding flow modeling as a new fully continuous paradigm for unified multimodal modeling. The related code and model are publicly released at https://github.com/hustvl/Multimodal-Flow.
Figures & tables
Figure 1: Comparison of multimodal modeling paradigms. Multimodal Flow combines continuous states with a shared generative objective and sampling procedure for language and vision. Blue and green denote discrete and continuous states respectively.
Figure 2: Multimodal Flow architecture and ordered hyperchunk representation. Frozen multimodal encoders map text blocks and images to continuous hyperchunks. After normalization and projection, a shared chunk-causal backbone applies joint attention and modality-specific FFNs. Multimodal decoders map generated hyperchunks back to text or images.
Figure 3: Parallel training, mixed multimodal pretraining, and sequential inference. The chunk-causal formulation predicts multiple target chunks in parallel during training and generates chunks sequentially at inference. Different chunk sequences express unimodal modeling, cross-modal generation, and downstream finetuning within the same Flow Matching objective.
Type
Model
Params
Single
Two
Count.
Colors
Pos.
Color Attr.
Overall ↑
Gen. only
SDXL ( Podell et al., 2024 )
2.6B
0.98
0.74
0.39
0.85
0.15
0.23
0.55
Hunyuan-DiT ( Li et al., 2024b )
1.5B
0.97
0.77
0.71
0.88
0.13
0.30
0.63
SD3 Medium ( Esser et al., 2024 )
2.0B
0.98
0.74
0.63
0.67
0.34
0.36
0.62
Unified
Fully discrete
LWM ( Liu et al., 2025a )
7B
0.93
0.41
0.46
0.79
0.09
0.15
0.47
Show-o ( Xie et al., 2025 )
1.3B
0.95
0.52
0.49
0.82
0.11
0.28
0.53
Table 1: Category-level comparison on GenEval. Unified models are grouped by multimodal modeling paradigm. Best and second-best results are shown in bold and underlined respectively.
Model
Params
Global
Entity
Attribute
Relation
Other
Overall ↑
SDXL ( Podell et al., 2024 )
2.6B
83.27
82.43
80.91
86.76
80.41
74.65
Hunyuan-DiT ( Li et al., 2024b )
1.5B
84.59
80.59
88.01
74.36
86.41
78.87
SD3 Medium ( Esser et al., 2024 )
2.0B
87.90
91.01
88.83
80.70
88.68
84.08
Show-o ( Xie et al., 2025 )
1.3B
79.33
75.44
78.02
84.45
60.80
67.27
Janus ( Wu et al., 2025 )
1.3B
82.33
87.38
87.70
85.46
86.41
79.68
Emu3-Gen ( Wang et al., 2024b )
8B
85.21
86.68
86.84
90.22
83.15
80.60
Table 2: Detailed text-to-image generation performance on DPG-Bench.
Paradigm
Model
Params
PT Tok.
POPE
MMB
SEEDB
VQAv2
GQA
OK-VQA
With Pretrained LLM Initialization
Und. only
MobileVLM-V2 ( Chu et al., 2024 )
2.7B
∼ 1.3T
84.7
63.2
–
–
61.1
–
LLaVA-Phi ( Zhu et al., 2024 )
2.7B
∼ 1.4T
85.0
59.8
–
71.4
–
–
LLaVA-v1.5 ( Liu et al., 2024a )
7.0B
∼ 2.0T
85.9
64.3
58.6
78.5
62.0
–
Qwen-VL-Chat ( Bai et al., 2023 )
7.0B
∼ 3.1T
83.7
60.6
58.2
78.2
57.5
56.6
Discrete
Janus ( Wu et al., 2025 )
1.3B
∼ 0.9T
87.0
69.4
63.7
77.3
59.1
–
Table 3: Multimodal understanding performance and pretraining scale. PT Tok. denotes cumulative pretraining tokens along the backbone inheritance chain.
Architecture
Architecture Type
GenEval ↑
GQA ↑
VQAv2 ↑
MMBench ↑
SEEDB ↑
Multimodal Flow
Fully Continuous
0.7134
55.60
69.03
46.74
51.85
Transfusion-Style
Discrete–Continuous Hybrid
0.6693
52.81
68.49
33.68
31.31
Chameleon-Style
Fully discrete
0.3744
45.83
56.37
38.40
42.40
Table 4: Controlled comparison of multimodal architectures under matched data, optimization schedules, and trainable parameter budgets.
Initialization
GenEval ↑
DPG-Bench ↑
SEEDB ↑
MMB ↑
OK-VQA ↑
Random
0.527
57.81
31.6
36.0
22.9
Mixed pretraining
0.821
83.44
62.4
67.2
39.1
Table 5: Downstream performance with random and mixed-pretrained initialization.
Figure 4: Training progress under mixed multimodal pretraining at different model capacities. We compare the Flow Matching objective, GPT-2-large PPL, GenEval, and CLIPScore for the 0.6B, 1.2B, and 1.6B models.
Attn. projection
FFN
Text PPL ↓
GenEval ↑
DPG ↑
CIDEr ↑
CLIPScore ↑
Specific
Specific
28.09
0.255
71.01
48.14
0.820
Specific
Shared
29.67
0.233
70.82
47.31
0.820
Shared
Specific
27.43
0.226
71.04
47.91
0.820
Shared
Shared
28.35
0.237
69.70
42.02
0.788
Table 6: Comparison of cross-modal parameterizations after 50B pretraining tokens.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Setting
Representation
Setting
Parameters
1.6B
Text encoder
T5-small
Transformer layers
32
Text embedding dim.
512
Hidden dimension
1600
Tokens per text chunk
8
Attention heads
25
Text input bottleneck
128
Head dimension
64
Vision encoder
SigLIP2-so400m
FFN hidden dimension
4224
Image input resolution
224
Appendix
Table 7: Backbone and representation settings of the main MF-1 model.
Experiment
Reference
Model size
PT Tokens
FT Tokens
Main benchmark comparisons
Tables 1 – 3
1.6B
150B
5B
Controlled architectures
Table 4
1.6B
50B
5B
Pretraining transfer
Table 5
1.6B
150B
5B
Attention/FFN parameterization
Table 6
0.6B
50B
0B
Visual representations and CFG
Figure 5
0.6B
50B
0B
Appendix
Table 8: Training token budgets for the main and controlled experiments. PT and FT denote pretraining and finetuning respectively.
Architecture
Text
Vision
Attention / FFNs
Multimodal Flow
T5 latent (512; 8-token blocks), flow
SigLIP2 latent ( 256×1152 ), flow
Shared / modality-specific (4224)
Chameleon-Style
T5 token IDs, AR
Janus VQ-16 (576 IDs; 16,384 codes), AR
Shared / modality-specific (4000)
Transfusion-style
T5 token IDs, AR
SigLIP2 latent ( 256×1152 ), flow
Shared / modality-specific (4062)
Appendix
Table 9: Representations and objectives in the controlled architecture comparison. AR denotes autoregressive cross-entropy; flow denotes velocity MSE. Parenthesized values in the last column are FFN hidden dimensions.
Task
Sampler
Steps
CFG scale
SDE γ
Visual question answering
SDE
16
3.0
1.0
Text-to-image generation
ODE
64
5.0
–
Appendix
Table 10: Inference settings for visual question answering and image generation.