Authors: Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, Chengxiang Huang, Dongzhan Zhou, Kai Chen, Qi Zhang, +3 more
Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Beijing Key Laboratory of Research on Large Models and Intelligent Governance · Beihang University · Beijing University of Posts and Telecommunications · Shanghai Artificial Intelligence Laboratory · AresoX
In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal models are reaching a bottleneck: existing approaches focus primarily on accommodating additional modalities while overlooking the distinct characteristics of each modality and the relations among them. Motivated by the multistage property of human multisensory perception, we propose a multimodal interplay pyramid that organizes multimodal modeling as a progression from modality-specific processing, through cross-modal alignment, to deeper multimodal integration. Guided by this pyramid, Gestalt adopts a unified discrete diffusion framework and an interplay-partitioned architecture, with learnable interplay tokens mediating cross-modal exchange and integration. The pyramid also structures its data organization and training strategy. Strong performance across image generation, multimodal understanding, and text-only evaluation shows that Gestalt significantly improves cross-modal integration while preserving modality-specific information, effectively harnessing the strengths of diffusion-based multimodal models and offering a promising path toward unified multimodal intelligence.
Figures & tables
Figure 1 : (a) The multimodal interplay pyramid progresses from modality-specific modeling to cross-modal alignment and synergy. (b) Combining the textual size-collar relation with the visible collar identifies the depicted dog as the smaller one. (c) Gestalt supports multimodal understanding, image generation, and language tasks within a unified model.
Figure 2 : Overview of the Gestalt architecture. Gestalt processes discrete vision (orange), interplay (green), and text (blue) tokens through Interplay Zone I ( N1 bottleneck interplay layers) and Interplay Zone II ( N2 full multimodal interplay layers) under a unified objective. Zone I uses restricted interplay attention, with cross-modal exchange mediated by interplay tokens, together with modality-specific FFNs; Zone II uses full multimodal self-attention and a shared FFN.
Figure 3 : Training recipe for interplay-oriented data organization during SFT.
Figure 4 : Qualitative comparison of fine-grained text-to-image generation. Red text highlights key semantic constraints in the prompts, such as spatial relations, attributes, negation, and properties.
Model
Overall
Style
World
Attr.
Action
Rel.
Comp.
Grammar
Layout
Logic
Text
Gen. Only
DALL-E-3
70.82
95.08
92.71
84.98
68.36
77.90
73.88
68.19
71.76
57.11
18.26
SD-3.5-Large
64.35
88.12
88.15
78.78
59.63
67.62
62.21
65.23
71.19
44.90
17.66
OmniGen2
71.39
94.35
84.83
83.03
66.57
73.06
70.49
76.40
80.63
56.55
27.99
Unified
Emu3
50.95
89.36
76.16
66.81
43.80
51.70
46.00
50.25
56.67
27.43
1.36
Table 1 : Evaluations on UniGenBench English Long for fine-grained text-to-image generation. The best and the second-best results are highlighted in bold and underline , respectively.
Model
Overall
Basic
Advanced
Designer
Short
Long
Short
Long
Short
Long
Short
Long
Gen. Only
PixArt-Sigma
62.00
58.12
70.66
75.25
57.65
49.50
62.11
52.41
FLUX.1 Pro
67.32
69.89
79.08
78.91
61.10
65.37
71.80
68.80
MidJourney V7
68.74
65.69
77.41
76.00
64.66
60.53
68.83
63.61
SD 3.5 Large
71.15
66.96
78.34
79.56
67.67
61.18
64.43
66.39
Table 2 : Evaluations on TIIF-Bench for text-to-image instruction following under short and long prompts. The best and the second-best results are highlighted in bold and underline , respectively.
Models
General
Vision-Centric
MME-P
GQA
MMStar-P
POPE
RWQA
MMVP
CVB 2d
CVB 3d
AR-Based
BAGEL
1687.0 †
66.4
70.9
88.2
67.6
69.3 †
77.7
84.2
Diffusion-based
MMaDA
1410.7 †
61.3 †
43.0
86.1 †
48.2
17.3
55.3
54.8
Lumina-DiMOO
1534.2 †
43.3
-
87.4 †
35.9
34.0
54.3
52.0
Table 3 : Evaluations on multimodal understanding benchmarks. The symbol † denotes the results from the official paper, while the rest of the results are evaluated using the official checkpoint and inference scripts. The best and the second-best results in diffusion-based models are highlighted.
Figure 5 : Visualization of image-text hidden representation across diffusion and AR architecture.
Type
Model
Norm Mean
Norm Std
Mean Distance
Cosine Sim.
CKA
Base
LLaDA
7.82
1.11
–
–
–
Projector Alignment
LLaDA-V
7.82
1.11
0.043
0.99998
0.99997
LaViDa
7.83
1.11
0.160
0.99976
0.99964
LLaDA-O
8.02
1.16
1.520
0.98136
0.97424
Unified Token Space
Lumina-DiMOO
0.85
0.49
7.426
0.58795
0.43894
Gestalt
5.29
0.76
2.590
0.99677
0.99494
Table 4 : Comparison of distribution shift (Norm and Distance) and representation preservation (Cosine Similarity and CKA) of LLaDA-based multimodal models in the language embedding space.
[][9pt] Model
MMLU
TruthfulQA
WinoGrande
HellaSwag
ARC-E
ARC-C
AR-Based
Show-o2
71.70
46.94
74.03
76.67
84.43
58.53
Janus-Pro
49.90
41.72
67.17
68.41
65.74
40.70
BAGEL
28.02
40.51
50.75
28.59
27.53
23.63
Diffusion-based
MMaDA
40.14
43.81
54.85
45.81
46.72
28.67
Table 5 : Evaluation of language capabilities based on six text-only benchmarks. The best and the second-best results among diffusion-based unified models are highlighted.
Model
Visual
Synergy
MIB-V
CoreCog-SM
MM-IMDb
SRBench
AR-Based
BAGEL
65.96
65.00
60.60
51.89
Diffusion-Based
MMaDA
44.23
43.20
30.39
36.72
Lumina-DiMOO
56.43
42.20
31.22
45.50
Table 6 : Evaluation on multimodal benchmarks requiring visual-specific and synergistic interplay.
Figure 6 : Visualization of interplay token distribution across samples with different multimodal interplay.
Primary Task
Secondary Capability
Samples
Percentage
Interplay
Multimodal Understanding (MMU)
MMU Total
9.743M
100.0%
R , Uv , Ut , S
Image Captioning & Scene Understanding
3.086M
31.7%
R , Uv
Knowledge-intensive Visual Understanding
2.105M
21.6%
Ut , S
General VQA & Instruction Following
1.630M
16.7%
Uv , S
Object, Attribute, State & Counting
1.558M
16.0%
Uv
Scene Text & OCR
0.437M
4.5%
Uv
Table 7 : Composition of the 13.7M instruction-tuning examples, including multimodal understanding and text-only data, together with their primary interplay types, including redundant information ( R ), visual-unique information ( Uv ), text-unique information ( Ut ), and synergistic information ( S )
Metric
Phase 1
Phase 2
Phase 3
Uniform
Curriculum
Diff.
Uniform
Curriculum
Diff.
Uniform
Curriculum
Diff.
MMStar-P
49.52
45.08
-4.44
49.42
49.72
+0.30
51.67
53.87
+2.20
MME-P
917.1
1030.9
+113.8
1135.8
1178.5
+42.7
1249.5
1211.6
-37.9
MMVP
26.67
28.00
+1.33
32.67
26.00
-6.67
29.33
31.33
+2.00
CVBench-2D
68.64
68.08
-0.56
73.02
72.11
-0.90
73.99
74.62
+0.63
Table 8 : Comparison between uniform data mixing and interplay-oriented curriculum learning using 10% of the SFT data.
Model
RS Data
FGRS
FGRC
BAGEL
–
42.51
36.65
GeoChat
–
53.47
21.90
SkySenseGPT
1.415M
79.76
55.50
Lumina-DiMOO ∗
50K
77.77
–
Gestalt
50K
79.34
38.20
Table 9 : Evaluation on remote-sensing benchmarks. RS Data denotes the number of remote-sensing samples used for fine-tuning.
Figure 7 : Qualitative results on UniGenBench. Gestalt demonstrates strong fine-grained generation capabilities on complex prompts involving multiple objects, attributes, and spatial relations.
We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity scaling or text-image-dominant designs, Lance explores a practical paradigm for unified multimodal modeling via collaborative multi-task training. It is grounded in two core principles: unified context modeling and decoupled capability pathways. Specifically, Lance is trained from scratch and employs a dual-stream mixture-of-experts architecture on shared interleaved multimodal sequences, enabling joint context learning while decoupling the pathways for understanding and generation. We further introduce modality-aware rotary positional encoding to mitigate interference among heterogeneous visual tokens and boost cross-task alignment. During training, Lance adopts a staged multi-task training paradigm with capability-oriented objectives and adaptive data scheduling to strengthen both semantic comprehension and visual generation performance. Experimental results demonstrate that Lance substantially outperforms existing open-source unified models in image and video generation, while retaining strong multimodal understanding capabilities. The homepage is available at https://lance-project.github.io.
The long-standing goal of multimodal AI is to build unified models in which visual understanding and visual generation mutually enhance one another. Despite recent works such as BAGEL, BLIP3o achieves remarkable progress; In practice, however, this unification remains one-directional: understanding routinely guides generation, yet how and why generation can support understanding is rarely investigated. We revisit this asymmetry and propose Generation-to-Understanding (G2U) synergy, where visual generation becomes an explicit intermediate reasoning step. Our framework enables a model to perform controlled generative acts, such as detail enhancement, context expansion or structural visualisation, to produce self-generated visual thoughts, which are then fed back into the model to refine perception without retraining or external tools. Through a comprehensive evaluation on twelve benchmarks, this reversed information flow consistently improves multimodal understanding. We show that generative fidelity bounds perceptual gain and that distinct families of edit prompts govern transfer efficiency. We further analyse whether models can decide what to imagine. While they can produce plausible edits, these self-generated visual thoughts lack stable task alignment, revealing that current large multimodal models fall short of true self-reflection. This work exposes a missing mechanism in unified cognition and suggests that imagination is not the end of understanding but its beginning.
Yujun Tong, Dongliang Chang, Zijin Yin +3
Beijing University of Posts and Telecommunications · Beijing Key Laboratory of Multimodal Data Intelligent Perception and Governance
Multimodal modeling represents a vital step from modality-agnostic reasoning toward world modeling. While early approaches predominantly rely on late-fusion that assembles encoders and frozen language backbones with output heads, recent efforts have shifted the paradigm toward native multimodal modeling (NMM) with the intrinsic integration of modalities for superior multimodal performance. Despite its potential, the design space of native architectures remains insufficiently defined. In this paper, we present the community with a formalized roadmap for this transition. Specifically, we formally define the architectural nativity, distinguishing mid-fusion and early-fusion from non-native paradigms. We further organize the existing native models through the lens of input-output duality into three categories: (i) Multi-to-Text for cross-modal comprehension with text-only output; (ii) Multi-to-Target for scenario-oriented generation, e.g., image, audio and video generation, and (iii) Multi-to-Multi for unified modeling with symmetric input-output. We deliver a comprehensive and industrial-grade investigation into the transition toward the definitive NMM framework, where understanding and generation seamlessly coexist within a unified transformer paradigm. We systematically unpack the end-to-end pipeline from industrial perspectives from architectural coordination, massive data curation, to full-stack training recipes, inference & deployment, and the comprehensive evaluation for truly native modeling.