Organizations: The Hong Kong University of Science and Technology · ATH, Alibaba Group · The Hong Kong University of Science and Technology (Guangzhou) · Northwestern Polytechnical University · Zhejiang University
Multimodal Large Language Models (MLLMs) are increasingly being applied in a wider range of real-world scenarios. However, due to the substantial labor cost, creating high-quality multimodal instruction datasets for MLLMs remains a significant challenge. Although some methods propose to generate instruction data, they often face limitations in modality support and struggle with generating multi-round instructions. To address these problems, we introduce UniData, a universal instruction generation pipeline, to transform simple user requirements into multi-round, multimodal instructions. Specifically, UniData first expands user requirements into multiple diverse events. Using these events, UniData then integrates an any-to-any large model for multimodal instruction generation. Finally, UniData enhances data quality by correcting irrelevant and redundant inference flow, leveraging correlations between instruction rounds. To train this pipeline, we also build UniDataset, a dataset comprising 20,000 entries across nine modalities for improved multimodal generation. Our experiments demonstrate that UniData achieves SOTA performance in data quality and can also enhance the understanding and generation capabilities of other multimodal models.
Figures & tables
Figure 1: Usage of UniData. (I) Users only input one simple requirement, combined with random emotions and occasions to create diverse events. (II) Each diverse event will generate one multimodal instruction. (III) The generated multimodal instructions include various modalities.
Figure 2: Comparison of different instruction generation frameworks. (A) Self-Instruct Wang et al. (2023) can only support language modality. (B) VIGC Wang et al. (2024a) integrates visual modality into its generation. (C) Multimodal Self-Instruct Zhang et al. (2024) can output abstract images. (D) Ours (UniData) can freely understand and generate different modalities in instructions.
Figure 3: Procedure of data construction. Stage (I) prepares diverse events by random topic pairs. Stage (II) leverages these events to produce multi-round instructions. Stage (III) forecasts eight types of multimodal indicator tokens and integrates these into instructions. Stage (IV) uses a toolkit to generate multimodal content.
Figure 4: Pipeline of Diverse Event Generator ( DEG ). Users can generate diverse events by inputting their needs.
Figure 5: Pipeline of Multimodal Instruction Generator ( MIG ). With just one event provided by the user, MIG can iteratively generate multi-round multimodal instructions based on the context.
Figure 6: Inference chain for correcting instruction flow.
Method
Backbones
GPT-Guided Metrics (↑) OpenAI (2024)
Avg. #Multimodalities (↑)
Avg. #Rounds
Reasonableness
Clarity
Detail
Relevance
#Image
#Music
#Others
Self-Instruct Wang et al. (2023)
GPT-4 OpenAI (2024)
0.624
0.444
0.594
0.520
-
-
-
≈ 3
VIGC Wang et al. (2024a)
Vicuna 7B Chiang et al. (2023)
0.460
0.359
0.534
0.569
-
-
-
≈ 1
Ours (UniData)
LLaMA-3 7B Grattafiori et al. (2024)
0.661
0.527
0.667
0.682
1.92
0.74
7.37
≈ 17.5
Table 1: Quantitative performance in data quality. Red indicates the best performance. (#Others includes all other modalities (#Emoji, #Code, #Math, #Map, #Link, #QR Code)).
Figure 7: Qualitative comparison in data quality.
Table 9
Figure 8: ( left ) Difference distribution of the events in three random generations. ( right ) Example of diverse events based on the same keywords.
Figure 9: Example of flow error correction.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
e-commerce
academic discussions
lifestyle
sports
mathematics
business
technology
travels
health
entertainment
art
history
cooking
parenting
fashion
finance
politics
literature
gardening
astronomy
music
education
computer science
programming
design
research
internship
food
beauty
entrepreneurship
startups
algorithm
bugs
scholar
IT
vision
supervision
classroom
assignment
game
psychology
social media
Appendix
Table 8: The 90-domain topic pool used to seed UniDataset generation. At sampling time two domains are intersected (Eq. 9 ) to form a composite theme.
Tool
Modality
Source
Role in the pipeline
GPT-4o OpenAI (2024)
<language>
OpenAI API
Orchestrator and language backbone; expands events, drafts dialogue, and validates outputs.
Google Search
<link>
https://www.google.com
Provides up-to-date factual snippets and reference URLs.
Google Maps
<map>
https://www.google.com/maps
Supplies geographical references and routing information.
DALL ⋅ E 3
<image>
https://openai.com/index/dall-e-3/
Synthesizes photo-realistic and stylized images from text prompts.
MusicGen Copet et al. (2023)
<music>
Open-source model
Generates instrumental clips from textual descriptions.
Apple Emoji Archive
<emoji>
https://www.apple.com/
Provides culturally familiar emoji symbols for affective cues.
Appendix
Table 9: External tools used by UniData during multimodal realization. Each tool is responsible for one modality, allowing the pipeline to be extended modularly.
Dataset
Input Modality
Output Modality
Avg. #Rounds
#Instruction
Num
Class
Num
Class
LLaVA Liu et al. (2023)
2
L/I
1
L
≈ 1.0
00 5K
VideoChat Li et al. (2023b)
2
L/V
1
L
≈ 1.8
0 11K
FIRE Li et al. (2024b)
2
L/I
1
L
≈ 1.0
100K
LAMM Yin et al. (2023)
3
L/I/3D
1
L
≈ 3.3
196K
mPLUG-DocOwl Ye et al. (2023)
4
L/I/T/L
1
L
-
-
Appendix
Table 10: Comparison with current multimodal instruction datasets. Num is the number of distinct modality types; Class lists the modalities themselves. Legend: L: <language> ; I: <image> ; V: <video> ; 3D: <point cloud> ; LK: <link> ; A: <audio> ; M: <music> ; E: <emoji> ; C: <code> ; MP: <map> ; MT: <math> ; Q: <QR code> .
Method
Reason. (↑)
Clarity (↑)
Detail (↑)
Relev. (↑)
Self-Instruct Wang et al. (2023)
18.5
0 9.7
0 9.8
0 8.5
VIGC Wang et al. (2024a)
10.3
0 5.3
14.4
10.9
Ours (UniData)
71.2
85.0
75.8
80.6
Appendix
Table 11: User study: percentage of samples (%) on which each method is selected as the best by 15 human evaluators across 50 samples (per-dimension forced choice). Red indicates the best performance.
Method
Backbones
GPT-Guided Metrics (↑)
Avg. #Multimodalities (↑)
Avg. #Rounds
Reasonableness
Clarity
Detail
Relevance
#Image
#Music
#Others
Self-Instruct Wang et al. (2023)
GPT-4 OpenAI (2024)
0.735
0.475
0.608
0.580
-
-
-
≈ 3
VIGC Wang et al. (2024a)
Vicuna 7B Chiang et al. (2023)
0.714
0.507
0.732
0.737
-
-
-
≈ 1
Ours (UniData)
LLaMA-3 7B Grattafiori et al. (2024)
0.752
0.701
0.808
0.845
1.04
0.27
3.82
≈ 10
Appendix
Table 12: Quality of generated instructions on the out-of-distribution (OOD) set. Bold marks the best score in each column. #Others aggregates the emoji, code, math, map, link, and QR-code modalities.
Position in dialogue
Multimodal frequency (%)
0 0 – 20 %
20.5
20 – 40 %
19.4
40 – 60 %
19.5
60 – 80 %
21.3
80 – 100 %
19.3
Appendix
Table 13: Fraction of multimodal content (%) as a function of the normalized position within a dialogue. Position is binned into five equal-width buckets.
Input type
#Image
#Music
#Others
General
0 1.92
0 0.74
7.37
Music-related
0 1.10
12.30
2.10
Image-related
14.90
0 0.10
5.20
Appendix
Table 14: Average count of each modality per dialogue when the user specifies a desired modality at input time. Red indicates the targeted modality in each row.
Figure 10: Diverse event generation by DEG . For the same keyword pair, three independently sampled events differ substantially in content, scenario, and intended modality, illustrating the diversity that drives downstream instruction variety.
Figure 11: End-to-end qualitative example (I): a multi-round, multimodal instruction produced by UniData. Image, music, and text are interleaved into a coherent narrative.
Figure 12: End-to-end qualitative example (II): a different topic pair, showing variation in dialogue length and modality mix while preserving topical coherence.
Figure 13: End-to-end qualitative example (III): a long-horizon dialogue that interleaves text, image, code, and emoji content, illustrating UniData’s behavior on extended interactions.