cs.CVSep 30, 2026

Gestalt: Large Multimodal Interplay Model

Authors: Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, Chengxiang Huang, Dongzhan Zhou, Kai Chen, Qi Zhang, +3 more

Organizations: Gaoling School of Artificial Intelligence, Renmin University of China · Beijing Key Laboratory of Research on Large Models and Intelligent Governance · Beihang University · Beijing University of Posts and Telecommunications · Shanghai Artificial Intelligence Laboratory · AresoX

Abstract

In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal models are reaching a bottleneck: existing approaches focus primarily on accommodating additional modalities while overlooking the distinct characteristics of each modality and the relations among them. Motivated by the multistage property of human multisensory perception, we propose a multimodal interplay pyramid that organizes multimodal modeling as a progression from modality-specific processing, through cross-modal alignment, to deeper multimodal integration. Guided by this pyramid, Gestalt adopts a unified discrete diffusion framework and an interplay-partitioned architecture, with learnable interplay tokens mediating cross-modal exchange and integration. The pyramid also structures its data organization and training strategy. Strong performance across image generation, multimodal understanding, and text-only evaluation shows that Gestalt significantly improves cross-modal integration while preserving modality-specific information, effectively harnessing the strengths of diffusion-based multimodal models and offering a promising path toward unified multimodal intelligence.

Figures & tables

Explore similar work

CardsList
  1. Lance: Unified Multimodal Modeling by Multi-Task Synergy

    May 18, 2026Fengyi Fu, Mengqi Huang, Shaojin Wu +10Multimodal UnderstandingMulti--Task Learning

  2. Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models

    May 15, 2026Yujun Tong, Dongliang Chang, Zijin Yin +3Multimodal GenerationLarge Multimodal Models

  3. Toward Native Multimodal Modeling: A Roadmap

    May 25, 2026Siyu An, Junru Lu, Junnan Dong +18Multimodal ModelCross-Modal