Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image--text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solutions, most rely on additional training with augmented data, which is computationally expensive and remains limited in preserving visual subjects, temporal consistency, and physical plausibility. In this work, we propose self-correction optimization (SCO), an effective training-free method for consistent interleaved generation. SCO treats the classifier-free guidance update as a reference and performs minimal self-correction under two complementary constraints, including new-event and state-preserving constraints. Specifically, the new-event constraint promotes temporal consistency across image--text sequences, while the state-preserving constraint maintains the coherence of visual subjects throughout subsequent generation steps. Experiments on challenging interleaved multimodal generation benchmarks demonstrate significant improvements in temporal coherence and visual-subject preservation. Furthermore, SCO can be extended to video generation and improves the modeling of physically grounded processes, including robot manipulation and long-horizon handcrafting.
Figures & tables
Figure 1: The mechanism of self-correction optimization, which consists of new-event and state-preserving constraints for the new-event realization and visual subject consistency.
Figure 2: Qualitative results on various interleaved generation tasks when evaluating the efficacy of SCO, including iterative image editing, novel view synthesis, and embodied-AI scenarios.
Type
Method
FID ↓
CLIP-I ↑
CLIP-T ↑
GPT-J1 ↑
Cs↑
Cb↑
GPT-J2 ↑
Modular Model
MiniGPT-5 Zheng et al. (2023)
91.3
0.151
0.269
0.327
0.13
0.20
0.243
MiniGPT-5 + SCO
88.5
0.170
0.285
0.365
0.16
0.22
0.279
Show-o2 Xie et al. (2026)
83.6
0.146
0.270
0.337
0.12
0.17
0.248
Show-o2 + SCO
79.7
0.169
0.281
0.349
0.13
0.19
0.264
MM-Interleaved Tian et al. (2024)
76.2
0.201
0.329
0.397
0.18
0.22
0.351
MM-Interleaved + SCO
73.6
0.227
0.337
0.419
0.20
0.23
0.384
Table 1: Quantitative results on the ISG-Bench dataset when plugging SCO into various open-source models, including modular and unified models.
Type
Method
FID ↓
CLIP-I ↑
CLIP-T ↑
GPT-J1 ↑
Cs↑
Cb↑
GPT-J2 ↑
Modular Model
MiniGPT-5 Zheng et al. (2023)
90.4
0.149
0.257
0.445
0.18
0.23
0.344
MiniGPT-5 + SCO
89.9
0.138
0.247
0.423
0.17
0.21
0.319
Show-o2 Xie et al. (2026)
88.5
0.173
0.280
0.471
0.19
0.26
0.385
Show-o2 + SCO
84.0
0.179
0.291
0.496
0.19
0.28
0.401
MM-Interleaved Tian et al. (2024)
94.1
0.140
0.269
0.432
0.15
0.19
0.327
MM-Interleaved + SCO
81.3
0.164
0.293
0.450
0.17
0.22
0.369
Table 2: Quantitative results on the Opening dataset when plugging SCO into various open-source models, including modular and unified models.
Figure 3: Sensitivity analysis for hyper-parameters of SCO ( α0 , sstate , Mstate , Mnew , and τend ).
MM-Interleaved
SenseNova-U1
Setting
FID ↓
CLIP-I ↑
GPT-J1 ↑
Cs↑
GPT-J2 ↑
FID ↓
CLIP-I ↑
GPT-J1 ↑
Cs↑
GPT-J2 ↑
Baseline
94.1
0.140
0.432
0.15
0.327
52.1
0.385
0.677
0.36
0.591
+astate
85.6
0.149
0.445
0.18
0.320
51.1
0.393
0.689
0.39
0.620
+anew
89.4
0.153
0.440
0.13
0.362
50.6
0.397
0.686
0.34
0.563
+Both
81.3
0.164
0.450
0.17
0.369
49.7
0.409
0.695
0.38
0.624
Table 4: Ablation analysis of two key constraints. MM-interleaved and SenseNova-U1 are respectively adopted as the baseline model on the Opening and ISG-Bench datasets.
Figure 7
Figure 5: Qualitative results when applied to robotic arm manipulation and handcrafting scenarios.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Qualitative results on the effectiveness of SCO when applied to embodiment and handcrafting fields.
Figure 7: The relationship between the new-event constraint/state-preserving constraint and how to rewrite text prompts.
Figure 8: Qualitative results on the interleaved multimodal generation for the storytelling task.
Figure 9: Qualitative results on the task of iterative image editing.
Figure 10: Qualitative results on the interleaved multimodal generation for the sequential object removal task.
The advancement of generative AI models capable of producing text and image marks a critical step forward in the realm of multimodal intelligence, particularly for tasks involving the interleaving of both modalities. To advance this intelligence to the next stage, it is crucial for models to autonomously generate free-form interleaved text-image sequences. In this paper, we introduce ILLUME-X, an advanced unified multimodal paradigm that enables high-quality, free-form interleaved text-image generation by improving multimodal data efficiency and stabilizing the multimodal training process. ILLUME-X comprises three key components: (i) an expanded training data pipeline optimized for interleaved text-image generation, (ii) a progressive training strategy with self-adaptive objectives for free-length multimodal token sequences, and (iii) an objective and comprehensive evaluation method ILScore for interleaved text-image sequences. Notably, our ILLUME-X outperforms previous unified models across multiple interleaved text-image generation tasks like style transfer, image decomposition and storytelling.
Chonghuinan Wang, Zhikai Chen, Chunwei Wang +9
Harbin Institute of Technology, Harbin, China · Huawei Noah’s Ark Lab, Shenzhen, China · Zhengzhou Advanced Research Institute of Harbin Institute of Technology, Zhengzhou, China +1
While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggle to maintain performance under complex interleaved instructions. This limitation stems from the structural separation of images and text in current paradigms, which forces models to bridge difficult long-range dependencies to match descriptions with visual targets. To address these challenges, we propose \texttt{I}mages i\texttt{N} \texttt{SE}n\texttt{T}ences (\textit{a.k.a}, INSET), a unified generation model that seamlessly embeds images as native vocabulary within textual instructions. By positioning visual features directly at their corresponding semantic slots, INSET leverages the contextual locality of transformers for precise object binding, effectively treating images as dense, expressive language tokens. Furthermore, we introduce a scalable data engine that synthesizes 15M high-quality interleaved samples from standard image and video datasets, utilizing VLMs and LLMs to construct rich, long-horizon sequences. Evaluation results on InterleaveBench demonstrate that INSET significantly outperforms state-of-the-art methods in multi-image consistency and text alignment, with performance gaps widening as input complexity increases. Beyond standard generation, our approach inherently extends to multimodal image editing, integrating visual content as part of the instruction to facilitate highly expressive and creative visual manipulations.
Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their architectures, they cannot achieve interleaved generation (text-image sequence), which has crucial applications in visual narratives, guidance, and embodied manipulation. Even the latest open-source Unified Multimodal Models (UMMs) exhibit limited performance in this regard. In this paper, we introduce InterleaveThinker, the first multi-agent pipeline designed to endow any existing image generator with interleaved generation capabilities. Specifically, we employ a planner agent to organize the image-text input sequence, instructing the image generator on the required execution at each step. Subsequently, we introduce a critic agent to evaluate the generator's outputs, identify samples that deviate from the planned instructions, and refine the instructions for regeneration. To implement this pipeline, we construct the Interleave-Planner-SFT-80k and Interleave-Critic-SFT-112k to perform a format cold-start. Then we develop Interleave-Critic-RL-13k to reinforce the step-wise instruction correction capability within a generation trajectory using GRPO. Since a single interleaved generation trajectory may involve over 25 generator calls, optimizing the entire trajectory is computationally impractical. Therefore, we propose accuracy reward and step-wise reward, allowing single-step RL to effectively guide the entire generation trajectory. The results show that InterleaveThinker improves performance across various image generators. On interleaved generation benchmarks, it achieves performance comparable to Nano Banana and GPT-5. Surprisingly, it also significantly enhances the base model on reasoning-based benchmarks; for example, on 4-step FLUX.2-klein, we observe substantial gains on WISE and RISE.