Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
Organizations: Independent Researcher · University of Maryland · University of Waterloo · MBZUAI
Abstract
Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are genuinely aligned and integrated within a single model remains unclear. To investigate this question, we introduce GapEval, a bidirectional benchmark designed to quantify the gap between understanding and generation capabilities, and quantitatively measure the cognitive coherence of the two "unified" directions. Each question can be answered in both modalities (image and text), enabling a symmetric evaluation of a model's bidirectional inference capability and cross-modal consistency. Experiments reveal a persistent gap between the two directions across a wide range of UMMs with different architectures, suggesting that current models achieve only surface-level unification rather than deep cognitive convergence of the two. To further explore the underlying mechanism, we conduct an empirical study from the perspective of knowledge manipulation to illustrate the underlying limitations. Our findings indicate that knowledge within UMMs often remains disjoint. The capability emergence and knowledge across modalities are unsynchronized, paving the way for further exploration.
Figures & tables
| Benchmark | Size | Category | Annotation | Und. Task | Gen. Task | Features | ||||||
| WK. | RS. | VP. | IF | WK | RS | SYN. | BI. | GQ. | ||||
| MMMU [ 67 ] | 11,550 | I2T | Human | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| MMBench [ 28 ] | 3,217 | I2T | Mixed | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| GenEval [ 14 ] | 553 | T2I | Mixed | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| DPG-Bench [ 18 ] | 1,065 | T2I | (M)LLM | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ |
| T2I-CoReBench [ 25 ] | 1,080 | T2I | (M)LLM | ✗ | ✗ | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ |
| Model | World Knowledge | Numerical Perception | Instruction Following | Reasoning | Gap ↓ | ||||||||||||
| Succ. | Und. | Gen. | Gap↓ | Succ. | Und. | Gen. | Gap↓ | Succ. | Und. | Gen. | Gap↓ | Succ. | Und. | Gen. | Gap↓ | ||
| Open-source UMM | |||||||||||||||||
| Bagel | 52.24 | 88.17 | 58.50 | 56.67 | 2.40 | 17.00 | 8.40 | 84.87 | 46.38 | 59.15 | 75.74 | 57.38 | 2.49 | 35.67 | 5.38 | 87.14 | 71.52 |
| OneCAT | 57.14 | 88.26 | 62.92 | 33.14 | 2.75 | 9.00 | 9.75 | 62.33 | 23.67 | 50.00 | 41.22 | 37.51 | 0.66 | 20.66 | 1.56 | 85.36 | 54.73 |
| UniWorld-V1 | 87.84 | 93.28 | 92.74 | 12.60 | 1.00 | 10.60 | 3.60 | 84.68 | 28.51 | 62.13 | 46.60 | 70.46 | 1.11 | 36.78 | 1.44 | 86.11 | 63.47 |
| UniPic2 | 39.29 | 79.60 | 41.16 | 65.23 | 1.75 | 15.00 | 7.50 | 82.37 | 51.86 | 66.75 | 58.77 | 45.69 | 7.86 | 34.18 | 3.93 | 82.61 | 68.97 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Judge Model | Bagel | OneCAT | UniWorld-V1 | UniPic2 |
|---|---|---|---|---|
| Gemini3-Flash | 72.96 | 60.67 | 66.64 | 73.39 |
| GPT5-mini | 71.52 | 54.73 | 63.47 | 68.97 |
| Judge Model | Show-o2 | OmniGen2 | Gemini2.5-F-I | GPT-Image-1 |
| Gemini3-Flash | 72.33 | 88.87 | 62.65 | 47.47 |
| GPT5-mini | 67.83 | 89.74 | 62.91 | 50.61 |
| Prompt for Understanding on World Knowledge Task. |
|---|
| [Image] |
| [Reference_Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. It includes the core information present in the [Answer] |
| Prompt for Generating on World Knowledge Task. |
|---|
| [Image] |
| [Reference_Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. Compare the generated image [Image] with the reference image [Reference_Image] and the caption [Question] , and decide whether the image should be judged as pass (score 1) or fail (score 0). |
| Prompt for Understanding on Reasoning Task. |
|---|
| [Image] |
| [Reference_Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. You are given a reasoning problem [Question] , an authoritative reference_answer (the expected result or phenomenon), and a model-generated answer [Answer] . Your goal is to determine whether the final outcome/result expressed in [Answer] matches the reference_answer. |
| Prompt for Generating on Reasoning Task. |
|---|
| [Image] |
| [Reference_Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. Compare [Image] with [Reference_Image] and [Answer] and check whether all required answer-relevant elements are present: the main result, key objects, and core information needed to visually answer the physics question posed by the problem. Major answer-relevant objects must not be missing or clearly misrepresented. |
| Prompt for Understanding on Numerical Perception Task. |
|---|
| [Image] |
| [Reference_Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. Use the JSON task in [Question] (its "objects" and "number" fields) as the authoritative specification of which object types and exact counts are required. The final confirmed result stated in [Answer] must include only those specified object types, with counts that exactly match the JSON, and must not introduce any extra or non-specified objects. |
| Prompt for Generating on Numerical Perception Task. |
|---|
| [Image] |
| [Reference_Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. Use the JSON specification in [Question] (its "objects" list and corresponding "number" fields) as the exact target: [Image] passes (score = 1) only if every specified object type appears with exactly the required quantity and class, regardless of other non-target real objects that may be present. |
| Prompt for Understanding on Instruction Following Task. |
|---|
| [Image] |
| [Reference_Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. From [Question] , understand the rule that modifies the image scenario and the reference text that concisely describes the expected result or core feature after this rule is applied. Identify the core aspect, feature, or outcome that must appear once the rule is in effect. |
| Prompt for Generating on Instruction Following Task. |
|---|
| [Image] |
| [Reference_Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. From [Question] , fully understand the rule’s intent and logic, and how it is supposed to modify the original image (objects, features, or arrangements in the original scenario). Identify the core effect or result that must appear after the rule is applied. |
| Prompt for evalution of und task in edit and inject knoledge. |
|---|
| [Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. Ensure the subject described in the [Answer] matches the subject in the ground_truth (whether it’s an animal, object, person, etc.). |
| 2. If the output_text and ground_truth both describe the basic features, position, state, or other relevant characteristics of the subject consistently, it is considered correct. |
| Prompt for evaluation of gen task in edit knowledge. |
|---|
| [Image] |
| Here is the question: [Question] |
| Here is the answer: [Answer] |
| Please judge the correctness of the answer. You should follow the following rules: |
| 1. Ensure the subject depicted in the [Image] is the same as the subject in the ground_truth (whether it’s an animal, object, person, etc.). |
| 2. If the [Image] clearly depicts the same main subject as the ground_truth, even if there are variations in its state, expression, angle, or other minor details, it is considered correct. |