cs.AIMay 23, 2025

MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation

Authors: Jihan Yao, Yushi Hu, Wenyuan Wang, Bin Han, Shangbin Feng, Guang Yang, Yujie Yi, Bingbing Wen, +5 more

Organizations: University of Washington · Allen Institute for AI

Abstract

Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation, especially for complex tasks that involve multiple modalities. We present MMMG, the first benchmark to bring the verifiable-task paradigm to multimodal generation, spanning 4 modality combinations (image, audio, interleaved text and image, interleaved text and audio). As few multimodal outputs can be checked by programs alone, MMMG targets tasks that are either verifiable or near-verifiable: by providing references, and constraining model judges with explicit rubrics. We keep tasks challenging for generation models while enabling reliable automatic evaluation through a combination of models and programs. MMMG encompasses 55 tasks (including 31 newly developed ones), each with a carefully designed evaluation pipeline, and 1288 instructions to systematically assess reasoning, controllability, and other key capabilities of multimodal generation models. Extensive validation demonstrates that MMMG is highly aligned with human judgment, achieving an average agreement of 94.4%. Benchmarking results on 29 models reveal that even though the state-of-the-art model, GPT Image, achieves 70.7% accuracy for image generation, it falls short on interleaved generation. Furthermore, results suggest considerable improvement space in audio generation, highlighting an important future direction.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MMGR: Multi-Modal Generative Reasoning Benchmark and Evaluation

    Date pendingZefan Cai, Haoyi Qiu, Tianyi Ma +13Multimodal Reasoning BenchmarksVideo Generation

  2. UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models

    Feb 9, 2026Cheng Yang, Chufan Shi, Bo Shui +7Multimodal GenerationMultimodal Model

  3. Multimodal Language Models as Text-to-Image Model Evaluators

    May 1, 2025Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat +5Modern Text-To-Image ModelsText-To-Image