Authors: Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, Chengxiang Huang, Dongzhan Zhou, Kai Chen, Qi Zhang, +3 more
Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Beijing Key Laboratory of Research on Large Models and Intelligent Governance · Beihang University · Beijing University of Posts and Telecommunications · Shanghai Artificial Intelligence Laboratory · AresoX
In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal models are reaching a bottleneck: existing approaches focus primarily on accommodating additional modalities while overlooking the distinct characteristics of each modality and the relations among them. Motivated by the multistage property of human multisensory perception, we propose a multimodal interplay pyramid that organizes multimodal modeling as a progression from modality-specific processing, through cross-modal alignment, to deeper multimodal integration. Guided by this pyramid, Gestalt adopts a unified discrete diffusion framework and an interplay-partitioned architecture, with learnable interplay tokens mediating cross-modal exchange and integration. The pyramid also structures its data organization and training strategy. Strong performance across image generation, multimodal understanding, and text-only evaluation shows that Gestalt significantly improves cross-modal integration while preserving modality-specific information, effectively harnessing the strengths of diffusion-based multimodal models and offering a promising path toward unified multimodal intelligence.
Figures & tables
Figure 1 : (a) The multimodal interplay pyramid progresses from modality-specific modeling to cross-modal alignment and synergy. (b) Combining the textual size-collar relation with the visible collar identifies the depicted dog as the smaller one. (c) Gestalt supports multimodal understanding, image generation, and language tasks within a unified model.
Figure 2 : Overview of the Gestalt architecture. Gestalt processes discrete vision (orange), interplay (green), and text (blue) tokens through Interplay Zone I ( N1 bottleneck interplay layers) and Interplay Zone II ( N2 full multimodal interplay layers) under a unified objective. Zone I uses restricted interplay attention, with cross-modal exchange mediated by interplay tokens, together with modality-specific FFNs; Zone II uses full multimodal self-attention and a shared FFN.
Figure 3 : Training recipe for interplay-oriented data organization during SFT.
Figure 4 : Qualitative comparison of fine-grained text-to-image generation. Red text highlights key semantic constraints in the prompts, such as spatial relations, attributes, negation, and properties.
Model
Overall
Style
World
Attr.
Action
Rel.
Comp.
Grammar
Layout
Logic
Text
Gen. Only
DALL-E-3
70.82
95.08
92.71
84.98
68.36
77.90
73.88
68.19
71.76
57.11
18.26
SD-3.5-Large
64.35
88.12
88.15
78.78
59.63
67.62
62.21
65.23
71.19
44.90
17.66
OmniGen2
71.39
94.35
84.83
83.03
66.57
73.06
70.49
76.40
80.63
56.55
27.99
Unified
Emu3
50.95
89.36
76.16
66.81
43.80
51.70
46.00
50.25
56.67
27.43
1.36
Table 1 : Evaluations on UniGenBench English Long for fine-grained text-to-image generation. The best and the second-best results are highlighted in bold and underline , respectively.
Model
Overall
Basic
Advanced
Designer
Short
Long
Short
Long
Short
Long
Short
Long
Gen. Only
PixArt-Sigma
62.00
58.12
70.66
75.25
57.65
49.50
62.11
52.41
FLUX.1 Pro
67.32
69.89
79.08
78.91
61.10
65.37
71.80
68.80
MidJourney V7
68.74
65.69
77.41
76.00
64.66
60.53
68.83
63.61
SD 3.5 Large
71.15
66.96
78.34
79.56
67.67
61.18
64.43
66.39
Table 2 : Evaluations on TIIF-Bench for text-to-image instruction following under short and long prompts. The best and the second-best results are highlighted in bold and underline , respectively.
Models
General
Vision-Centric
MME-P
GQA
MMStar-P
POPE
RWQA
MMVP
CVB 2d
CVB 3d
AR-Based
BAGEL
1687.0 †
66.4
70.9
88.2
67.6
69.3 †
77.7
84.2
Diffusion-based
MMaDA
1410.7 †
61.3 †
43.0
86.1 †
48.2
17.3
55.3
54.8
Lumina-DiMOO
1534.2 †
43.3
-
87.4 †
35.9
34.0
54.3
52.0
Table 3 : Evaluations on multimodal understanding benchmarks. The symbol † denotes the results from the official paper, while the rest of the results are evaluated using the official checkpoint and inference scripts. The best and the second-best results in diffusion-based models are highlighted.
Figure 5 : Visualization of image-text hidden representation across diffusion and AR architecture.
Type
Model
Norm Mean
Norm Std
Mean Distance
Cosine Sim.
CKA
Base
LLaDA
7.82
1.11
–
–
–
Projector Alignment
LLaDA-V
7.82
1.11
0.043
0.99998
0.99997
LaViDa
7.83
1.11
0.160
0.99976
0.99964
LLaDA-O
8.02
1.16
1.520
0.98136
0.97424
Unified Token Space
Lumina-DiMOO
0.85
0.49
7.426
0.58795
0.43894
Gestalt
5.29
0.76
2.590
0.99677
0.99494
Table 4 : Comparison of distribution shift (Norm and Distance) and representation preservation (Cosine Similarity and CKA) of LLaDA-based multimodal models in the language embedding space.
[][9pt] Model
MMLU
TruthfulQA
WinoGrande
HellaSwag
ARC-E
ARC-C
AR-Based
Show-o2
71.70
46.94
74.03
76.67
84.43
58.53
Janus-Pro
49.90
41.72
67.17
68.41
65.74
40.70
BAGEL
28.02
40.51
50.75
28.59
27.53
23.63
Diffusion-based
MMaDA
40.14
43.81
54.85
45.81
46.72
28.67
Table 5 : Evaluation of language capabilities based on six text-only benchmarks. The best and the second-best results among diffusion-based unified models are highlighted.
Model
Visual
Synergy
MIB-V
CoreCog-SM
MM-IMDb
SRBench
AR-Based
BAGEL
65.96
65.00
60.60
51.89
Diffusion-Based
MMaDA
44.23
43.20
30.39
36.72
Lumina-DiMOO
56.43
42.20
31.22
45.50
Table 6 : Evaluation on multimodal benchmarks requiring visual-specific and synergistic interplay.
Figure 6 : Visualization of interplay token distribution across samples with different multimodal interplay.
Primary Task
Secondary Capability
Samples
Percentage
Interplay
Multimodal Understanding (MMU)
MMU Total
9.743M
100.0%
R , Uv , Ut , S
Image Captioning & Scene Understanding
3.086M
31.7%
R , Uv
Knowledge-intensive Visual Understanding
2.105M
21.6%
Ut , S
General VQA & Instruction Following
1.630M
16.7%
Uv , S
Object, Attribute, State & Counting
1.558M
16.0%
Uv
Scene Text & OCR
0.437M
4.5%
Uv
Table 7 : Composition of the 13.7M instruction-tuning examples, including multimodal understanding and text-only data, together with their primary interplay types, including redundant information ( R ), visual-unique information ( Uv ), text-unique information ( Ut ), and synergistic information ( S )
Metric
Phase 1
Phase 2
Phase 3
Uniform
Curriculum
Diff.
Uniform
Curriculum
Diff.
Uniform
Curriculum
Diff.
MMStar-P
49.52
45.08
-4.44
49.42
49.72
+0.30
51.67
53.87
+2.20
MME-P
917.1
1030.9
+113.8
1135.8
1178.5
+42.7
1249.5
1211.6
-37.9
MMVP
26.67
28.00
+1.33
32.67
26.00
-6.67
29.33
31.33
+2.00
CVBench-2D
68.64
68.08
-0.56
73.02
72.11
-0.90
73.99
74.62
+0.63
Table 8 : Comparison between uniform data mixing and interplay-oriented curriculum learning using 10% of the SFT data.
Model
RS Data
FGRS
FGRC
BAGEL
–
42.51
36.65
GeoChat
–
53.47
21.90
SkySenseGPT
1.415M
79.76
55.50
Lumina-DiMOO ∗
50K
77.77
–
Gestalt
50K
79.34
38.20
Table 9 : Evaluation on remote-sensing benchmarks. RS Data denotes the number of remote-sensing samples used for fine-tuning.
Figure 7 : Qualitative results on UniGenBench. Gestalt demonstrates strong fine-grained generation capabilities on complex prompts involving multiple objects, attributes, and spatial relations.