cs.AIOct 8, 2026

UniData: Universal Multimodal Instruction Generation Pipeline

Authors: Jiaqi Tang, Yi-Feng Wu, Yuting Zhang, Hao Lu, Bowen Fu, Qing-Guo Chen, Xiaogang Xu, Yuwei Hu, +7 more

Organizations: The Hong Kong University of Science and Technology · ATH, Alibaba Group · The Hong Kong University of Science and Technology (Guangzhou) · Northwestern Polytechnical University · Zhejiang University

Abstract

Multimodal Large Language Models (MLLMs) are increasingly being applied in a wider range of real-world scenarios. However, due to the substantial labor cost, creating high-quality multimodal instruction datasets for MLLMs remains a significant challenge. Although some methods propose to generate instruction data, they often face limitations in modality support and struggle with generating multi-round instructions. To address these problems, we introduce UniData, a universal instruction generation pipeline, to transform simple user requirements into multi-round, multimodal instructions. Specifically, UniData first expands user requirements into multiple diverse events. Using these events, UniData then integrates an any-to-any large model for multimodal instruction generation. Finally, UniData enhances data quality by correcting irrelevant and redundant inference flow, leveraging correlations between instruction rounds. To train this pipeline, we also build UniDataset, a dataset comprising 20,000 entries across nine modalities for improved multimodal generation. Our experiments demonstrate that UniData achieves SOTA performance in data quality and can also enhance the understanding and generation capabilities of other multimodal models.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. UniVideo: Unified Understanding, Generation, and Editing for Videos

    Oct 9, 2025Cong Wei, Quande Liu, Zixuan Ye +5Image-to-Video GenerationUnified Multimodal Models