MMMG: a Comprehensive and Reliable Benchmark for Multitask Multimodal Generation
Organizations: University of Washington · Allen Institute for AI
Abstract
Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align with human evaluation, especially for complex tasks that involve multiple modalities. We present MMMG, the first benchmark to bring the verifiable-task paradigm to multimodal generation, spanning 4 modality combinations (image, audio, interleaved text and image, interleaved text and audio). As few multimodal outputs can be checked by programs alone, MMMG targets tasks that are either verifiable or near-verifiable: by providing references, and constraining model judges with explicit rubrics. We keep tasks challenging for generation models while enabling reliable automatic evaluation through a combination of models and programs. MMMG encompasses 55 tasks (including 31 newly developed ones), each with a carefully designed evaluation pipeline, and 1288 instructions to systematically assess reasoning, controllability, and other key capabilities of multimodal generation models. Extensive validation demonstrates that MMMG is highly aligned with human judgment, achieving an average agreement of 94.4%. Benchmarking results on 29 models reveal that even though the state-of-the-art model, GPT Image, achieves 70.7% accuracy for image generation, it falls short on interleaved generation. Furthermore, results suggest considerable improvement space in audio generation, highlighting an important future direction.
Figures & tables
| Task | Subtask | Description | In. | Out. | # | Evaluation |
| Object Generation | Inclusion | Include one or two unrelated objects in the scene. | 40 | VLM | ||
| Exclusion | Exclude one related object from the scene. | 40 | VLM | |||
| Count | Generate exactly N objects. | 40 | VLM | |||
| Attribution | Generate an object with uncommon attributes. | 40 | VLM | |||
| Knowledge | Reason the answer object to a multi-hop question. | 40 | VLM | |||
| Commonsense | Reason the answer object/scene by commonsense. | 40 | VLM |
| Dataset | # Samples | # Tasks | Generation Modality | Evaluation | Tested Capability | ||||||||
| + | + | human | mllm | score | code | gen | edit | reason | |||||
| GenEval ( Ghosh et al., 2023 ) | 553 | 6 | ✔ | ✘ | ✘ | ✘ | ✘ | ✘ | ✔ | ✔ | ✔ | ✘ | ✘ |
| DrawBench ( Saharia et al., 2022 ) | 200 | 11 | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ | ✘ |
| GenAI-Bench ( Li et al., 2024 ) | 1,600 | 8 | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ | ✘ |
| AudioTime ( Xie et al., 2024 ) | 500 | 4 | ✘ | ✔ | ✘ | ✘ | ✘ | ✘ | ✔ | ✔ | ✘ | ✘ | |
| MusicEval ( Liu et al., 2025 ) | 384 | 1 | ✘ | ✔ | ✘ | ✘ | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ | ✘ |
| Model | Arena | GenEval | Draw | GenAI | MMMG |
| Imagen 3 | 1064 | 0.707 | 0.861 | 0.793 | 0.474 |
| Recraft v3 | 1018 | 0.732 | 0.826 | 0.817 | 0.441 |
| Luma Photon | 997 | 0.738 | 0.766 | 0.804 | 0.587 |
| Flux 1.1 Pro | 992 | 0.588 | 0.725 | 0.736 | 0.431 |
| Ideogram 2 | 1011 | 0.615 | 0.757 | 0.782 | 0.508 |
| Dalle 3 | 978 | 0.627 | 0.809 | 0.811 | 0.352 |
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Example | Input | Output |
| Table Generation | Create a 2x2 table image. In the first column, place the text ’apple’ in the top cell and ’pear’ in the bottom cell. In the second column, place an image of an apple in the top cell and an image of a pear in the bottom cell. | ||
| Figure Generation | Create a histogram to visualize the given data. <data> | ||
| Format Color | Create a watermelon farm using only varying shades of red. | ||
| Format Symmetric | Generate an image of a futuristic cityscape. The image must be axisymmetric along the vertical center line. | ||
| Art Style | Create a painting of a dandelion sea in Impressionist style. | ||
| Photography | Create q zoomed out photo of a small bag of coffee beans from below. |
| Statistics | Number |
| Total number of modality combinations | 4 |
| Total number of tasks | 55 |
| - I : A : I-T : A-T | 15 : 12 : 22 : 6 |
| Total number of questions | 1288 |
| - I : A : I-T : A-T | 410 : 238 : 480 : 160 |
| Total number of images | 542 |
| You are a multimodal assistant capable of generating both text and images. When visual content would enhance your response or is specifically requested, you can generate or edit images through advanced diffusion models. To generate or edit an image: 1. Identify when visual content would be beneficial or requested. 2. Insert an image generation/editing placeholder using the following format: <image_start><image_prompt="Detailed image generation or editing prompt here."><image_ref=[reference identifiers]><image_end> 3. The post-processing system replaces this placeholder with an image created or edited based on your instructions. 4. Naturally incorporate references to the generated or edited image in your ongoing conversation. When crafting image prompts, follow these guidelines: For image prompts: • Provide detailed, specific descriptions (15-30 words) for optimal results. • Include artistic styles (photorealistic, cartoon, watercolor, etc.) or style transfers. • Specify key objects and their attributes (colors, textures, etc.), or modifications. • Detail composition elements (spatial relationships, perspective, lighting, etc.), or compositional changes. • Ensure instructions are clear and concise. For image references: Three reference types are available: 1. Image generation (no reference): <image_ref=[]> 2. Editing user-provided images: Format: <image_ref=[i]> where i is the index of the provided image (indices starting at 0). Example: <image_ref=[0]> references the first provided image. Multiple images example: <image_ref=[0,2]> references the first and third provided images. 3. Editing previously generated images: Format: <image_ref=[#N]> , where N is the sequential number of previously generated images (starting from 0). Example: <image_ref=[#3]> references the fourth generated image. Multiple images example: <image_ref=[#0,#2]> references the first and third generated images. Important: Use only one reference type within each placeholder. Different reference types may be used across multiple placeholders. Provide concise and direct responses following user instructions precisely. Always maintain the exact placeholder format for proper parsing, ensuring that both images and text appear in the required order. Do not omit any necessary text following image placeholders. |
| You are a multimodal assistant capable of generating both text and audio. When audio content would enhance your response or is specifically requested, you can generate audio through text-to-audio models. To generate audio: 1. Identify when audio content would be beneficial or requested. 2. Insert an audio generation placeholder using the format: <audio_start><audio_type="sound" OR "speech" OR "music"><audio_text="Text to be spoken here."><audio_style="Descriptive text here." OR audio reference ID><audio_end> 3. The post-processing system replaces this placeholder with generated audio based on your specifications. 4. Naturally incorporate references to the generated audio in your ongoing conversation. When crafting audio prompts, follow these guidelines: Audio Type: • Must be exactly one of: "sound" , "speech" , or "music" . • "speech" : For human speech. • "sound" : For environmental sounds or effects. • "music" : For musical compositions or instrumental pieces. Audio Text: • For "speech" : Provide the exact transcript. • For "sound" or "music" : Leave as empty string (""). • Keep speech concise (typically under 50 words). Audio Style: 1. Descriptive Text: • For "speech" : Specify voice characteristics (gender, emotion, pace, pitch, accent). • For "sound" : Specify sound source, environment, qualities. • For "music" : Specify genre, mood, tempo, instruments. 2. Reference Audio: • For consistency, particularly with speech: – Previously generated audio: <audio_style=#N> ( N is sequential number starting at 0). – User-provided audio: <audio_style=N> ( N is sequential number of provided audio starting at 0). • Important: Only reference audio that itself does not reference previous audio to avoid circular references. Provide concise, direct responses precisely following user instructions. In multi-speaker scenarios, maintain consistent and distinctive voice characteristics for each speaker. Always maintain the exact placeholder format for correct parsing |
| Task | GPT-4o | Gemini 2.5 | Qwen2.5-VL | IAA | ||||
| agree | corr | agree | corr | agree | corr | agree | corr | |
| Object Inclusion | 0.925 | 0.776 | 0.900 | 0.715 | 0.888 | 0.657 | 1.000 | 1.000 |
| Object Exclusion | 0.963 | 0.924 | 0.913 | 0.823 | 0.925 | 0.855 | 1.000 | 1.000 |
| Object Count | 0.875 | 0.709 | 0.963 | 0.912 | 0.900 | 0.763 | 0.975 | 0.943 |
| Object Knowledge | 0.963 | 0.925 | 0.963 | 0.925 | 0.938 | 0.875 | 1.000 | 1.000 |
| Object Commonsense | 0.913 | 0.787 | 0.888 | 0.696 | 0.938 | 0.822 | 1.000 | 1.000 |
| Task | Imagen 3 | Recraft v3 | Luma Photon | Flux 1.1 Pro | Ideo -gram 2 | Dalle 3 | SD 3.5 | Gemini 2 Image | GPT Image | BILP-o3 | Janus Pro | Gemini 2.5 Image |
| Object Inclusion | 0.838 | 0.688 | 0.831 | 0.444 | 0.863 | 0.788 | 0.544 | 0.844 | 0.869 | 0.475 | 0.706 | 0.919 |
| Object Exclusion | 0.338 | 0.300 | 0.425 | 0.325 | 0.469 | 0.244 | 0.013 | 0.281 | 0.819 | 0.138 | 0.063 | 0.681 |
| Object Count | 0.269 | 0.319 | 0.369 | 0.375 | 0.319 | 0.119 | 0.256 | 0.356 | 0.569 | 0.138 | 0.263 | 0.450 |
| Object Knowledge | 0.494 | 0.481 | 0.656 | 0.325 | 0.419 | 0.463 | 0.150 | 0.706 | 0.531 | 0.238 | 0.081 | 0.588 |
| Object Commonsense | 0.256 | 0.306 | 0.288 | 0.206 | 0.288 | 0.288 | 0.306 | 0.275 | 0.163 | 0.194 | 0.175 | 0.363 |
| Object Attribution | 0.325 | 0.206 | 0.319 | 0.256 | 0.275 | 0.294 | 0.244 | 0.375 | 0.619 | 0.331 | 0.363 | 0.475 |
| Task | Seed Llama | Anole | GPT-4o + GPT Image | Gemini 2.5 + GPT Image | Gemini 2 Image | GPT Image | Gemini 2.5 Image |
| Semantic Consistency | 0.000 | 0.000 | 0.613 | 0.763 | 0.013 | 0.675 | 0.325 |
| Multi-angle Consistency | 0.000 | 0.000 | 0.230 | 0.461 | 0.352 | 0.448 | 0.367 |
| Multi-view Consistency | 0.000 | 0.000 | 0.064 | 0.221 | 0.143 | 0.188 | 0.191 |
| Composition Consistency | 0.000 | 0.000 | 0.800 | 0.738 | 0.000 | 0.075 | 0.225 |
| Decomposition Consistency | 0.000 | 0.000 | 0.600 | 0.875 | 0.013 | 0.575 | 0.638 |
| Self Count | 0.000 | 0.038 | 0.100 | 0.850 | 0.213 | 0.763 | 0.000 |
| Task | Stable Audio | Audio LDM 2 | AudioGen | Make-An -Audio 2 | Tango 2 | MusicGen | Tango Music | Yue |
| Sound Begin-End | 0.525 | 0.450 | 0.475 | 0.631 | 0.525 | - | - | - |
| Sound Inclusion | 0.700 | 0.413 | 0.450 | 0.575 | 0.513 | - | - | - |
| Sound Reasoning | 0.014 | 0.014 | 0.042 | 0.611 | 0.194 | - | - | - |
| Sound Silence | 0.063 | 0.019 | 0.019 | 0.131 | 0.006 | - | - | - |
| Instrument Inclusion | 0.817 | 0.833 | - | - | - | 0.833 | 0.950 | 0.600 |
| Instrument Exclusion | 0.225 | 0.163 | - | - | - | 0.200 | 0.050 | 0.525 |
| Task | Gemini 2.5 + VoxInstruct | Gemini 2.5 + VoiceLDM | Spirit LM |
| Voice Attribution | 0.684 | 0.568 | 0.000 |
| Voice Replication | 0.625 | 0.109 | 0.002 |
| Speech Multi-lingual | 0.654 | - | - |
| Transcript Generation | 0.638 | 0.438 | 0.200 |
| Transcript Editing | 0.200 | 0.350 | 0.000 |
| Conversation Generation | 0.788 | 0.375 | 0.000 |
| Task | Imagen 3 | Recraft v3 | Luma Photon | Flux 1.1 Pro | Ideo -gram 2 | Dalle 3 | SD 3.5 | Gemini 2 Image | GPT Image | BILP-o3 | Janus Pro | Average |
| Object Inclusion | 4.24 | 3.16 | 5.05 | 5.43 | 4.24 | 4.24 | 5.79 | 3.08 | 3.67 | 4.90 | 3.08 | 4.26 |
| Object Exclusion | 4.24 | 5.29 | 7.21 | 5.29 | 6.44 | 2.35 | 2.45 | 6.44 | 1.22 | 4.24 | 1.41 | 4.24 |
| Object Count | 3.08 | 2.35 | 4.18 | 5.29 | 5.05 | 6.75 | 1.23 | 5.43 | 5.43 | 5.10 | 3.16 | 4.28 |
| Object Knowledge | 1.23 | 3.08 | 2.35 | 2.00 | 3.67 | 5.10 | 7.21 | 1.23 | 5.05 | 5.83 | 2.35 | 3.35 |
| Object Commonsense | 2.35 | 6.44 | 4.69 | 2.35 | 1.41 | 4.69 | 3.67 | 2.00 | 4.24 | 4.64 | 6.93 | 3.95 |
| Object Attribution | 3.46 | 3.08 | 5.05 | 5.79 | 7.75 | 1.22 | 3.67 | 4.47 | 3.08 | 2.35 | 2.45 | 3.85 |
| Task | Seed Llama | Anole | GPT-4o + GPT Image | Gemini 2.5 + GPT Image | Gemini 2 Image | GPT Image | Average |
| Semantic Consistency | 0.00 | 0.00 | 7.35 | 2.45 | 2.45 | 6.33 | 3.10 |
| Multi-angle Consistency | 0.00 | 0.00 | 2.54 | 0.55 | 4.30 | 2.69 | 1.68 |
| Multi-view Consistency | 0.00 | 0.00 | 1.03 | 0.20 | 2.01 | 0.78 | 0.67 |
| Composition Consistency | 0.00 | 0.00 | 6.93 | 9.28 | 0.00 | 6.33 | 3.76 |
| Decomposition Consistency | 0.00 | 0.00 | 6.93 | 2.83 | 2.45 | 6.33 | 3.09 |
| Self Count | 0.00 | 2.45 | 6.93 | 6.93 | 4.69 | 8.37 | 4.89 |
| Task | Stable Audio | Audio LDM 2 | AudioGen | Make-An -Audio 2 | Tango 2 | MusicGen | Tango Music | Yue | Average |
| Sound Begin-End | 5.66 | 2.83 | 13.12 | 8.95 | 5.25 | - | - | - | 7.16 |
| Sound Inclusion | 10.59 | 8.37 | 5.66 | 2.45 | 5.13 | - | - | - | 6.44 |
| Sound Reasoning | 2.72 | 2.72 | 2.72 | 5.44 | 1.94 | - | - | - | 3.11 |
| Sound Silence | 2.45 | 3.67 | 1.23 | 1.23 | 0.63 | - | - | - | 1.84 |
| Instrument Inclusion | 6.26 | 3.77 | - | 0.00 | - | 3.77 | 3.27 | 0.00 | 2.84 |
| Instrument Exclusion | 2.83 | 7.35 | - | 0.00 | - | 5.66 | 5.66 | 2.83 | 4.05 |
| Task | Gemini 2.5 + VoxInstruct | Gemini 2.5 + VoiceLDM | Spirit LM | Average |
| Voice Attribution | 4.14 | 6.08 | 0.01 | 5.11 |
| Voice Replication | 5.91 | 2.80 | 0.13 | 4.35 |
| Speech Multi-lingual | 4.06 | 0.17 | 0.00 | 2.12 |
| Transcript Generation | 7.35 | 12.89 | 0.00 | 10.12 |
| Transcript Editing | 5.66 | 12.00 | 0.00 | 8.83 |
| Conversation Generation | 7.35 | 8.49 | 0.00 | 7.92 |
| Task | Obj. Inc. | Obj. Exc. | Obj. Cou. | Obj. Kno. | Obj. Com. | Obj. Att. | Com. Rel. | Uni. Rel. | Rel. Spa. | Abs. Spa. | Reg. Fill | Bor. Fill | Sin. TR | Dou. TR | Mul. TR |
| Object Include | - | 0.677 | 0.464 | 0.712 | 0.441 | 0.557 | 0.734 | 0.654 | 0.720 | 0.501 | 0.473 | 0.096 | 0.570 | 0.611 | 0.627 |
| Object Exclude | 0.677 | - | 0.859 | 0.671 | 0.232 | 0.742 | 0.884 | 0.930 | 0.889 | 0.540 | 0.691 | 0.389 | 0.601 | 0.799 | 0.656 |
| Object Count | 0.464 | 0.859 | - | 0.609 | 0.198 | 0.623 | 0.673 | 0.919 | 0.857 | 0.680 | 0.696 | 0.378 | 0.767 | 0.894 | 0.838 |
| Object Knowl. | 0.712 | 0.671 | 0.609 | - | 0.466 | 0.363 | 0.764 | 0.671 | 0.700 | 0.529 | 0.530 | 0.309 | 0.634 | 0.772 | 0.663 |
| Object Common. | 0.441 | 0.232 | 0.198 | 0.466 | - | -0.120 | 0.304 | 0.298 | 0.196 | -0.068 | 0.087 | -0.458 | 0.477 | 0.453 | 0.220 |
| Object Attribute | 0.557 | 0.742 | 0.623 | 0.363 | -0.120 | - | 0.664 | 0.664 | 0.740 | 0.615 | 0.823 | 0.612 | 0.188 | 0.449 | 0.550 |
| You are a multimodal assistant capable of generating interleaved text and images based on user instructions. • Follow the required modality structure and number in user’s instruction exactly, especially when multiple images are implied or requested. • Generate separate images for each described part, do not combine multiple concepts into one image unless told to. • Interleave images and text in the order described. Your goal is to match the user’s intent with exact number and sequence of image and text. |
| Task | Gemini Image w/ prompt | Gemini Image w/o prompt |
| Semantic Consistency | 0.263 | 0.013 |
| Multi-Angel Consistency | 0.135 | 0.352 |
| Multi-View Consistency | 0.094 | 0.143 |
| Compose Consistency | 0.013 | 0.000 |
| Decompose Consistency | 0.000 | 0.013 |
| Interleaved Object Adding | 0.399 | 0.545 |
| Error Category | % in All | Observation | Plausible Causes |
| Audio generation failure | 20.0% | Generated audio consists of pure noise with no meaningful structure. | CLAPScore is not OOD-generalizable. When generation quality is low, it cannot make reliable judgments. |
| Audio generation quality unsatisfactory | 26.7% | Generated audio contains recognizable and reasonable sounds but is overall of low quality with distortion. For example, the models repeatedly failed to generate proper reggae music, often producing tracks dominated by drums. | CLAPScore is not OOD-generalizable. When generation quality is low, it cannot make reliable judgments. |
| Lack of fine-grained understanding | 26.7% | Some generated audio includes fine-grained details that are easily recognizable by humans, such as brief dog barks or brief sneezes, but their duration is too short. Others contain noisy background such as sirens in loud traffic, which is treated as a meaningless piece by the CLAP model. | Since CLAPScore computes the similarity of the whole generated audio with the reference audio, short duration or noisy background will underestimate such similarity. Future methods could attempt to locate the target sound first before computing similarity. |
| Underrepresented reference audio | 26.7% | Some generated audio is not representative in the reference audio. A peaceful drum sequence is not considered as a drum piece by the CLAP model since all the reference drum music is metallic-like. | The reference datasets we use are standard and may not fully capture a sound, an instrument, or a music genre. Enriching the dataset diversity is important for more reliable evaluation. |
| Error Category | % in All | Observation | Plausible Causes |
| Generated speech with noise or distortion | 21.4% | Some generated speeches contain noticeable background white noise. Sometimes, the noise is intermittent volume fluctuation (sometimes slightly lower, sometimes higher) or long pauses, which may interfere with the model’s judgment. | WavLM is not OOD-generalizable. When generation quality is low, it cannot make reliable judgments. |
| Generated speech with synthetic pattern | 57.1% | Some generated speeches can be considered as coming from the same speaker, but with clear synthetic patterns, making it hard even for human judges to decide whether the generated speeches are from the same speaker. | WavLM is not OOD-generalizable. It is trained on real-world human voice and may not effectively judge synthetic speeches. |
| Computed similarity score close to threshold | 21.4% | The generated speeches are very similar to the references, aside from slight white noise. This likely caused the model’s judgment to be only minimally affected, placing the scores near the decision threshold. | WavLM is not well-calibrated. It will only give high scores to the same speaker and low scores to different speakers. When the score is close to the threshold, it is likely to make mistakes. |
| Error Category | Examples | Possible Reasons | # (%) of All Err. | # (%) of Gem Err. | # (%) of GPT Err. |
| Fail to render required scene properly | Instruction: “Create an image of a waterfall flowing into the sky, containing one pumpkin and one suitcase.” Error: The waterfall is rendered normally, not flowing into the sky. | The model relies heavily on physical laws and realistic priors embedded in the training data, and struggles to synthesize scenes that violate common-sense physics. | 3 (1.2%) | 0 (0.0%) | 3 (3.0%) |
| Fail to render objects of the required number | Instruction: “Generate an image of a historical site. Please include a single spaceship and a giraffe in the image.” Error: Two spaceships are generated even though the instruction asks for one. | The model demonstrates limited numeracy and counting capabilities. This may be due to the attention mechanism, which often fails to disentangle identical object instances and leads to superfluous objects. | 49 (19.9%) | 27 (18.5%) | 22 (22.0%) |
| Object rendered with distortion or unrealistic appearance | Instruction: “Create an image of an endless mirrored hallway, with one cactus and one saxophone.” Error: The mirrored hallway is distorted. | This may be due to under-training. Models may not be trained on the target objects enough to reproduce full details. | 15 (6.1%) | 6 (4.1%) | 9 (9.0%) |
| Required object not present, partially present, or wrong object present | Instruction: “Create an image of a whale flying through a sunset sky, featuring one suitcase and one mailbox.” Error: The image does not include a suitcase. | Failed instruction following may indicate under post-training. Instead of generating the required image, the model generates the most plausible image. | 1 (0.4%) | 0 (0.0%) | 1 (1.0%) |
| Unable to remove the correlated object | Instruction: “Generate an image of people camping. Do not include tents in the image.” Error: The generated image still contains tents. | Negative constraints paradoxically activate associated concepts in the latent space. In addition, strong semantic co-occurrence between scenes and objects (spurious correlation) makes it difficult to decouple these associations. | 31 (12.6%) | 24 (16.4%) | 7 (7.0%) |
| Fail to render required attribute (single object) | Instruction: “Generate an image of a single chair with 5 legs.” Error: The generated chair has 4 legs. | Strong visual priors, such as the normative four-legged chair, override textual prompts and create a conflict between internal knowledge and counterfactual instructions. | 28 (11.4%) | 20 (13.7%) | 8 (8.0%) |
| Error Category | Examples | Possible Reasons | # (%) of All Err. | # (%) of Gem Err. | # (%) of GPT Err. |
| Fail to apply attribute requirements to all objects (multiple objects) | Instruction: “Generate an image of birds on a wire where all birds are facing the same direction.” Error: Not all birds are facing the same direction. | The model has difficulty maintaining attribute consistency across multiple entities. Stochastic generation processes often fail to enforce constraints uniformly on every instance. | 6 (2.4%) | 4 (2.7%) | 2 (2.0%) |
| Fail to apply attribute requirements only to the required objects (multiple objects) | Instruction: “Generate an image of a bookshelf where all books are standing vertically except one lying horizontally.” Error: More than one book is lying horizontally. | High logical complexity hampers the execution of exception logic, causing the model to over-generalize rules to excluded entities. | 6 (2.4%) | 5 (3.4%) | 1 (1.0%) |
| Wrong / insufficient knowledge to reason the target object, wrong object generated | Instruction: “Create an image featuring only the national flag of a country that has a major city located on both the European and Asian continents.” Error: The model generates the French flag instead of the Turkish flag. | Unlike LLMs, image generation models are usually not trained on extensive world knowledge, so they can only follow the literal instruction and may fail to infer the intended target object. | 20 (8.1%) | 8 (5.5%) | 12 (12.0%) |
| Wrong / insufficient knowledge to reason the target object, multiple objects generated | Instruction: “Generate an image of a small, handheld stringed instrument, commonly used in folk and country music and important in black American music.” Error: The model generates multiple instruments. | This shows signs of reward hacking in post-training. When the reward model does not punish guessing behavior, the model may generate multiple objects instead of following the instruction. | 1 (0.4%) | 0 (0.0%) | 1 (1.0%) |
| Fail to render the required logical relation between objects | Instruction: “Create an image featuring a single nail and a single snake, and the nail is longer than the snake.” Error: The nail is still shorter than the snake. | The unrealistic logical relation in the prompt conflicts with the model’s internal semantic priors, and the model favors the prior. | 25 (10.2%) | 17 (11.6%) | 8 (8.0%) |
| Fail to render the required spatial relation between objects | Instruction: “Generate an image of a peaceful garden, with a single watering can at the upper right quarter of the image.” Error: The watering can is only on the right side. | Models are trained on image-caption pairs that stress semantic consistency and overlook perception accuracy. | 12 (4.9%) | 6 (4.1%) | 6 (6.0%) |
| Error Category | Examples | Possible Reasons | # (%) of All Err. | # (%) of Gem Err. | # (%) of GPT Err. |
| Fail to render image format at all | Instruction: “Generate a space scene with planets and nebulae. The entire image should be surrounded by a simple and flat, solid and black border of approximately 10% of the image width on all sides.” Error: The scene is generated without any border. | Formatting constraints are ignored because the model interprets them as soft suggestions rather than hard rules. This is likely an OOD problem, where such image formats are rarely requested in training data. | 1 (0.4%) | 0 (0.0%) | 1 (1.0%) |
| Rendered text with misspelling | Instruction: “Generate the text ‘always move forward’.” Error: The generated text is misspelled as “always move forwad.” | Text rendering is a known problem for image generation. Models operate at the pixel level rather than the character level, and without explicit orthographic verification they tend to generate pseudo-text that is visually similar but orthographically incorrect. | 10 (4.1%) | 6 (4.1%) | 4 (4.0%) |
| Rendered text with duplication | Instruction: “Generate the text ‘capture the moment’.” Error: The generated text becomes “capture the the moment.” | Repetition artifacts arise from loops or redundancies in the text-generation attention mechanism. | 3 (1.2%) | 2 (1.4%) | 1 (1.0%) |
| Required text not rendered | Instruction: “Generate the text ‘Washington WA 98105 Evergreen State’.” Error: No text is generated. | Lack of text rendering training makes the model neglect text requirements. | 0 (0.0%) | 0 (0.0%) | 0 (0.0%) |
| Required text rendered in distortion | Instruction: “Generate the text ‘Enjoy The Little Things Always’.” Error: The word “Enjoy” has distorted glyphs. | Lack of text rendering training makes the model unable to generate high-quality text. | 1 (0.4%) | 1 (0.7%) | 0 (0.0%) |
| Required text in the wrong place or overflow | Instruction: “Generate an image of exactly two posters on an office wall…” Error: The model generates three posters, and the second text appears on the middle poster. | Lack of text rendering training makes the model unable to place text correctly. Adding a layout controller or training on such data may help. | 5 (2.0%) | 2 (1.4%) | 3 (3.0%) |
| Error Category | Examples | Possible Reasons | # (%) of All Err. | # (%) of Gem Err. | # (%) of Age Err. |
| Generate wrong number of images | Instruction: “Using the provided image as the reference angle, create four additional images…” Error: The model generates only three images and misses the final view. | Difficulty in precise counting and planning; the model loses track of the count during the sequential generation loop. | 40 (12.2%) | 33 (13.9%) | 7 (7.6%) |
| Only image description provided but no image | Instruction: “Create an image depicting a musician’s room that includes exactly one microphone, one guitar case, and one music stand…” Error: The model provides the text evaluation but does not generate the actual image. | Tool-use failure; the model prioritizes text reasoning instead of actually calling the image generation tool. | 2 (0.6%) | 2 (0.8%) | 0 (0.0%) |
| Low image quality or distorted, unrealistic object | Instruction: “Based on the provided image showing a frontal view, create four more images depicting the scene from specific angles…” Error: Later views become severely distorted and blurry. | This may be due to under-training. Models may not be trained on the target objects enough to reproduce full details. | 4 (1.2%) | 4 (1.7%) | 0 (0.0%) |
| Combine multiple images together instead of generating multiple images | Instruction: “Based on the provided image showing the frontal view, create four additional images…” Error: The model generates a single image grid instead of four separate images. | Models, especially modality-unified ARMs, are likely to entangle multiple images in one output due to continuous latent representations and a training bias toward single-turn unified outputs. | 3 (0.9%) | 3 (1.3%) | 0 (0.0%) |
| Inconsistent image sequence: objects, style, or scene change during generation | Instruction: “Create four images that sequentially show the addition of a passport, a map, a camera, and a pair of sunglasses…” Error: Previously added objects disappear and the suitcase style changes. | Agent models do not show strong planning capabilities on image generation tasks. ARMs also lack cross-image consistency mechanisms, so state is not maintained across steps. | 16 (4.9%) | 8 (3.4%) | 8 (8.7%) |
| Image sequence shows no variance: image repetition | Instruction: “Using the provided image as the reference angle, create four additional images…” Error: All generated images are identical to the original reference view. | Mode output collapse. This is likely related to the same phenomenon as repetitive outputs in LLMs, and may indicate lack of post-training. | 11 (3.3%) | 8 (3.4%) | 3 (3.3%) |
| Error Category | Examples | Possible Reasons | # (%) of All Err. | # (%) of Gem Err. | # (%) of Age Err. |
| Fail to render objects of the required number | Instruction: “Create an image of an office desk that includes a stapler, a mouse, and a pen…” Error: The generated image contains two pens, violating the “exactly once” constraint. | The model demonstrates limited numeracy and counting capabilities. This may be due to the attention mechanism, which often fails to disentangle identical object instances and leads to superfluous objects. | 12 (3.6%) | 10 (4.2%) | 2 (2.2%) |
| Failed spatial generation: images are not in the position / angle asked | Instruction: “Using the provided image as the reference angle, create four additional images…” Error: The model generates two images at 30 degrees to the right and two at 30 degrees to the left. | Poor understanding of 3D position. Models are trained on image-caption pairs that stress semantic consistency and overlook other perception abilities. | 4 (1.2%) | 2 (0.8%) | 2 (2.2%) |
| Failed temporal generation: images do not show the required adding / removing order | Instruction: “Create three images that sequentially show the addition of a coffee mug, a notebook, and a pen…” Error: The first image already shows the final result. | Agent models do not show strong planning capabilities on image generation tasks. ARMs also lack cross-image consistency mechanisms, so state is not maintained across steps. | 21 (6.4%) | 20 (8.4%) | 1 (1.1%) |
| Failed logic generation: images do not show the required logical order | Instruction: “Create three images, each featuring a single balloon of a different color… Arrange the images in alphabetical order…” Error: The generated order is still red, yellow, and purple. | Agent models do not show strong planning capabilities on image generation tasks. ARMs also lack cross-image consistency mechanisms, so state is not maintained across steps. | 11 (3.3%) | 7 (3.0%) | 4 (4.3%) |
| Wrong editing region | Instruction: “Create an image after the utensils have been removed from the photo.” Error: The model removes the food but leaves the utensils untouched. | Current training for image editing lacks verifiable and accurate reward, which makes the editing region inaccurate. | 5 (1.5%) | 3 (1.3%) | 2 (2.2%) |
| Oversized / undersized editing region | Instruction: “Create an image that displays the result after removing the man’s wig…” Error: The edit also removes part of the forehead and background. | Current training for image editing lacks verifiable and accurate reward, which makes the editing region inaccurate. | 2 (0.6%) | 1 (0.4%) | 1 (1.1%) |
| Error Category | Examples | Possible Reasons | # (%) of All Err. | # (%) of Gem Err. | # (%) of Age Err. |
| Invalid edit: edit not applied to the required text / object | Instruction: “Create an image that shows the result after replacing the cop with a smiling clown…” Error: The output image is identical to the input image. | Under-training for image editing tasks. Models cannot understand the editing instruction well. | 20 (6.1%) | 16 (6.8%) | 4 (4.3%) |
| Partial edit: edit only applies to a single object / text | Instruction: “Generate an image showing the editing result after making all the sprinkles on the cupcakes blue…” Error: Only the front cupcake is edited. | Under-training for image editing tasks. Models cannot understand the editing instruction well. | 4 (1.2%) | 3 (1.3%) | 1 (1.1%) |
| Global changes outside the editing region | Instruction: “Create an image displaying the result after replacing the highest kite with an eagle…” Error: The eagle is added correctly, but the whole sky changes color. | Current training for image editing lacks verifiable and accurate reward, which makes the editing scope difficult to control. | 11 (3.3%) | 4 (1.7%) | 7 (7.6%) |
| Generated image-text interleaved content not in the required modality order | Instruction: “For each phase, start with an image that illustrates the phase, followed by a written explanation…” Error: The model outputs all text first and all images at the end. | Lack of instruction-following training for interleaved image-text content. Models cannot maintain the required output structure and modality sequence. | 6 (1.8%) | 5 (2.1%) | 1 (1.1%) |
| Parsing error: wrong function-calling format for image generation | Instruction: “Develop a 3-step guide…” Error: The model outputs raw tool code or no image at all. | Lack of training on agentic data. Models cannot reliably output the formatted text required by tool agents. | 0 (0.3%) | 1 (0.4%) | 0 (0.0%) |
| Parsing error: failed to generate text in the required format | Instruction: “…return only a JSON object that maps each item to its corresponding color…” Error: The model returns a bulleted list instead of valid JSON. | Lack of training on agentic data. Models cannot reliably output the formatted text required by tool agents. | 19 (5.8%) | 19 (8.0%) | 0 (0.0%) |
| Error Category | Examples | Possible Reasons | # (%) of All Err. | # (%) of Gem Err. | # (%) of Age Err. |
| Image inconsistent with the self-generated object description in text | Instruction: “Create four images that sequentially show the result after removing the sunglasses, the camera, the map, and the passport…” Error: The text says the passport is removed, but the final image still shows the passport. | Lack of training on interleaved image-text tasks. Models cannot produce self-consistent image-text pairs. For agent models, this is often because the image generation tool does not strictly follow the plan, and the planning model cannot correct it with feedback. | 23 (7.0%) | 15 (6.3%) | 8 (8.7%) |
| Image inconsistent with the self-generated relation description in text | Instruction: “…answer the following two questions…” Error: The model answers “Left,” but the generated image places the object on the right. | Lack of training on interleaved image-text tasks. Models cannot produce self-consistent image-text pairs. For agent models, this is often because the image generation tool does not strictly follow the plan, and the planning model cannot correct it with feedback. | 45 (13.7%) | 31 (13.1%) | 14 (15.2%) |
| Image inconsistent with the self-generated OCR recognition results in text | Instruction: “…after generating the image, output only the announcement text in XML format…” Error: The XML output says “Science Fair,” but the rendered image contains gibberish. | Lack of training on interleaved image-text tasks. Models cannot produce self-consistent image-text pairs. For agent models, this is often because the image generation tool does not strictly follow the plan, and the planning model cannot correct it with feedback. | 16 (4.9%) | 14 (5.9%) | 2 (2.2%) |
| No step-by-step reasoning analysis | Instruction: “Analyze the following SVG code step-by-step…” Error: The model gives only a short summary instead of the required step-by-step reasoning. | Under-training for multimodal reasoning tasks involving image generation. RLVR may improve such multimodal reasoning capabilities if more verifiable data like MMMG is released. | 2 (0.6%) | 2 (0.8%) | 0 (0.0%) |
| Wrong reasoning analysis | Instruction: “What geometric shape does this SVG code describe?…” Error: The model reasons incorrectly and concludes that the code describes a circle. | The model hallucinates the function of the code or cannot mentally simulate the SVG geometry. | 46 (14.0%) | 22 (9.3%) | 24 (26.1%) |
| Failed generation | Instruction: “…” Error: The output is a completely black image or a corrupted file that cannot be opened. | Activation of safety filters or internal system generation failures. | 5 (1.5%) | 4 (1.7%) | 1 (1.1%) |