Organizations: The Hong Kong University of Science and Technology · ATH, Alibaba Group · The Hong Kong University of Science and Technology (Guangzhou) · Northwestern Polytechnical University · Zhejiang University
Multimodal Large Language Models (MLLMs) are increasingly being applied in a wider range of real-world scenarios. However, due to the substantial labor cost, creating high-quality multimodal instruction datasets for MLLMs remains a significant challenge. Although some methods propose to generate instruction data, they often face limitations in modality support and struggle with generating multi-round instructions. To address these problems, we introduce UniData, a universal instruction generation pipeline, to transform simple user requirements into multi-round, multimodal instructions. Specifically, UniData first expands user requirements into multiple diverse events. Using these events, UniData then integrates an any-to-any large model for multimodal instruction generation. Finally, UniData enhances data quality by correcting irrelevant and redundant inference flow, leveraging correlations between instruction rounds. To train this pipeline, we also build UniDataset, a dataset comprising 20,000 entries across nine modalities for improved multimodal generation. Our experiments demonstrate that UniData achieves SOTA performance in data quality and can also enhance the understanding and generation capabilities of other multimodal models.
Figures & tables
Figure 1: Usage of UniData. (I) Users only input one simple requirement, combined with random emotions and occasions to create diverse events. (II) Each diverse event will generate one multimodal instruction. (III) The generated multimodal instructions include various modalities.
Figure 2: Comparison of different instruction generation frameworks. (A) Self-Instruct Wang et al. (2023) can only support language modality. (B) VIGC Wang et al. (2024a) integrates visual modality into its generation. (C) Multimodal Self-Instruct Zhang et al. (2024) can output abstract images. (D) Ours (UniData) can freely understand and generate different modalities in instructions.
Figure 3: Procedure of data construction. Stage (I) prepares diverse events by random topic pairs. Stage (II) leverages these events to produce multi-round instructions. Stage (III) forecasts eight types of multimodal indicator tokens and integrates these into instructions. Stage (IV) uses a toolkit to generate multimodal content.
Figure 4: Pipeline of Diverse Event Generator ( DEG ). Users can generate diverse events by inputting their needs.
Figure 5: Pipeline of Multimodal Instruction Generator ( MIG ). With just one event provided by the user, MIG can iteratively generate multi-round multimodal instructions based on the context.
Figure 6: Inference chain for correcting instruction flow.
Method
Backbones
GPT-Guided Metrics (↑) OpenAI (2024)
Avg. #Multimodalities (↑)
Avg. #Rounds
Reasonableness
Clarity
Detail
Relevance
#Image
#Music
#Others
Self-Instruct Wang et al. (2023)
GPT-4 OpenAI (2024)
0.624
0.444
0.594
0.520
-
-
-
≈ 3
VIGC Wang et al. (2024a)
Vicuna 7B Chiang et al. (2023)
0.460
0.359
0.534
0.569
-
-
-
≈ 1
Ours (UniData)
LLaMA-3 7B Grattafiori et al. (2024)
0.661
0.527
0.667
0.682
1.92
0.74
7.37
≈ 17.5
Table 1: Quantitative performance in data quality. Red indicates the best performance. (#Others includes all other modalities (#Emoji, #Code, #Math, #Map, #Link, #QR Code)).
Figure 7: Qualitative comparison in data quality.
Table 9
Figure 8: ( left ) Difference distribution of the events in three random generations. ( right ) Example of diverse events based on the same keywords.
Figure 9: Example of flow error correction.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
e-commerce
academic discussions
lifestyle
sports
mathematics
business
technology
travels
health
entertainment
art
history
cooking
parenting
fashion
finance
politics
literature
gardening
astronomy
music
education
computer science
programming
design
research
internship
food
beauty
entrepreneurship
startups
algorithm
bugs
scholar
IT
vision
supervision
classroom
assignment
game
psychology
social media
Appendix
Table 8: The 90-domain topic pool used to seed UniDataset generation. At sampling time two domains are intersected (Eq. 9 ) to form a composite theme.
Tool
Modality
Source
Role in the pipeline
GPT-4o OpenAI (2024)
<language>
OpenAI API
Orchestrator and language backbone; expands events, drafts dialogue, and validates outputs.
Google Search
<link>
https://www.google.com
Provides up-to-date factual snippets and reference URLs.
Google Maps
<map>
https://www.google.com/maps
Supplies geographical references and routing information.
DALL ⋅ E 3
<image>
https://openai.com/index/dall-e-3/
Synthesizes photo-realistic and stylized images from text prompts.
MusicGen Copet et al. (2023)
<music>
Open-source model
Generates instrumental clips from textual descriptions.
Apple Emoji Archive
<emoji>
https://www.apple.com/
Provides culturally familiar emoji symbols for affective cues.
Appendix
Table 9: External tools used by UniData during multimodal realization. Each tool is responsible for one modality, allowing the pipeline to be extended modularly.
Dataset
Input Modality
Output Modality
Avg. #Rounds
#Instruction
Num
Class
Num
Class
LLaVA Liu et al. (2023)
2
L/I
1
L
≈ 1.0
00 5K
VideoChat Li et al. (2023b)
2
L/V
1
L
≈ 1.8
0 11K
FIRE Li et al. (2024b)
2
L/I
1
L
≈ 1.0
100K
LAMM Yin et al. (2023)
3
L/I/3D
1
L
≈ 3.3
196K
mPLUG-DocOwl Ye et al. (2023)
4
L/I/T/L
1
L
-
-
Appendix
Table 10: Comparison with current multimodal instruction datasets. Num is the number of distinct modality types; Class lists the modalities themselves. Legend: L: <language> ; I: <image> ; V: <video> ; 3D: <point cloud> ; LK: <link> ; A: <audio> ; M: <music> ; E: <emoji> ; C: <code> ; MP: <map> ; MT: <math> ; Q: <QR code> .
Method
Reason. (↑)
Clarity (↑)
Detail (↑)
Relev. (↑)
Self-Instruct Wang et al. (2023)
18.5
0 9.7
0 9.8
0 8.5
VIGC Wang et al. (2024a)
10.3
0 5.3
14.4
10.9
Ours (UniData)
71.2
85.0
75.8
80.6
Appendix
Table 11: User study: percentage of samples (%) on which each method is selected as the best by 15 human evaluators across 50 samples (per-dimension forced choice). Red indicates the best performance.
Method
Backbones
GPT-Guided Metrics (↑)
Avg. #Multimodalities (↑)
Avg. #Rounds
Reasonableness
Clarity
Detail
Relevance
#Image
#Music
#Others
Self-Instruct Wang et al. (2023)
GPT-4 OpenAI (2024)
0.735
0.475
0.608
0.580
-
-
-
≈ 3
VIGC Wang et al. (2024a)
Vicuna 7B Chiang et al. (2023)
0.714
0.507
0.732
0.737
-
-
-
≈ 1
Ours (UniData)
LLaMA-3 7B Grattafiori et al. (2024)
0.752
0.701
0.808
0.845
1.04
0.27
3.82
≈ 10
Appendix
Table 12: Quality of generated instructions on the out-of-distribution (OOD) set. Bold marks the best score in each column. #Others aggregates the emoji, code, math, map, link, and QR-code modalities.
Position in dialogue
Multimodal frequency (%)
0 0 – 20 %
20.5
20 – 40 %
19.4
40 – 60 %
19.5
60 – 80 %
21.3
80 – 100 %
19.3
Appendix
Table 13: Fraction of multimodal content (%) as a function of the normalized position within a dialogue. Position is binned into five equal-width buckets.
Input type
#Image
#Music
#Others
General
0 1.92
0 0.74
7.37
Music-related
0 1.10
12.30
2.10
Image-related
14.90
0 0.10
5.20
Appendix
Table 14: Average count of each modality per dialogue when the user specifies a desired modality at input time. Red indicates the targeted modality in each row.
Figure 10: Diverse event generation by DEG . For the same keyword pair, three independently sampled events differ substantially in content, scenario, and intended modality, illustrating the diversity that drives downstream instruction variety.
Figure 11: End-to-end qualitative example (I): a multi-round, multimodal instruction produced by UniData. Image, music, and text are interleaved into a coherent narrative.
Figure 12: End-to-end qualitative example (II): a different topic pair, showing variation in dialogue length and modality mix while preserving topical coherence.
Figure 13: End-to-end qualitative example (III): a long-horizon dialogue that interleaves text, image, code, and emoji content, illustrating UniData’s behavior on extended interactions.
Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends unified modeling to the video domain. UniVideo adopts a dual-stream design, combining a Multimodal Large Language Model (MLLM) for instruction understanding with a Multimodal DiT (MMDiT) for video generation. This design preserves the MLLM's original text generation capabilities, enables accurate interpretation of complex multimodal instructions, and maintains visual consistency in the generated content. Built on this architecture, UniVideo unifies diverse video generation and editing tasks under a single multimodal instruction paradigm and is jointly trained across them. Extensive experiments demonstrate that UniVideo matches or surpasses state-of-the-art task-specific baselines in text/image-to-video generation, in-context video generation and in-context video editing. Notably, the unified design of UniVideo enables two forms of generalization. First, UniVideo supports task composition, such as combining editing with style transfer, by integrating multiple capabilities within a single instruction. Second, even without explicit training on free-form video editing, UniVideo transfers its editing capability from large-scale image editing data to this setting, handling unseen instructions such as changing the environment or altering materials within a video. Beyond these core capabilities, UniVideo also supports visual-prompt-based video generation, where the MLLM interprets visual prompts and guides the MMDiT during synthesis. To foster future research, we released our model and code.
Cong Wei, Quande Liu, Zixuan Ye +5
University of Waterloo · Kling Team, Kuaishou Technology
We introduce MUNITE, a latent-variable framework for flexible any-to-any multimodal generation that treats encoding and latent generation as the same inference problem under different amounts of observed evidence. Given any subset of modalities, MUNITE models the conditional distribution over the latent representation associated with the complete observation. Full observation recovers deterministic encoding, no observation recovers the latent marginal, and intermediate subsets define conditional latent inference, all within a single conditional flow model. A shared latent sample captures variation that must remain consistent across generated targets, while modality-specific generative decoders model the remaining uncertainty independently. To learn these conditional distributions from incomplete training examples, we extend conditional flow matching through self-distillation: predictions conditioned on richer available observations supervise the same model conditioned on smaller subsets at the same intermediate latent state. When the richer-evidence trajectory follows the exact conditional flow, this provides the same expected learning signal as full-target denoising. Across PolyMNIST-D-Q, FFHQ64, and image-text-audio, MUNITE achieves competitive or better generation quality and source-target alignment, with higher joint-generation coherence. In particular, it attains the highest coherence in all one-to-many and unconditional image-text-audio comparisons, showing the effectiveness of unified latent inference across diverse multimodal settings.
We introduce MUNI, an end-to-end multimodal latent diffusion framework for any-to-any generation that unifies subset-conditioned cross-modal generation and unconditional joint sampling through a shared stochastic latent. Existing multimodal generative models are largely LLM-based, which limits leveraging modality-specific generators and requires text-paired data for training. Recent diffusion- and flow-based any-to-any extensions take a different direction but still rely on text-aligned embeddings, fully-paired training, or matched-dimensionality deterministic mappings. MUNI rests on two complementary contributions, one architectural and one in the training objective. First, we extend latent diffusion to multimodal any-to-any generation end-to-end: instead of the standard two-stage recipe that precomputes a frozen latent space and then fits a prior over it, MUNI jointly trains modality-specific encoders, expressive decoders, and a single shared flow-based prior under one objective. Second, we identify that the standard aggregation rules of multimodal variational inference are insufficient once coupled with a learned prior and expressive decoders. A suitable shared latent must simultaneously satisfy coherence across generated modalities, predictive sufficiency of subset latents, and minimality of the latent content. We propose a routed training objective whose structural choices align the latent with these criteria and admit a minimal-sufficiency characterization in the realizable setting. Experiments on PolyMNIST-Quadrant-Labels and a large-scale image-text-audio benchmark show MUNI matching or exceeding the strongest baselines on conditional generation while opening its largest margins on unconditional coherence. Project page: https://muni-proj.github.io/.