cs.CVOct 7, 2026

Self-correction Optimization for Interleaved Multimodal Generation

Authors: Xin You, Zhiwei Ning, Zukai Chen, Minghui Zhang, Xuanke Shi, Hanxiao Zhang, Jingsong Liu, Jie Yang, +2 more

Organizations: Shanghai Jiao Tong University · SenseTime Research · Technical University of Munich

Abstract

Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image--text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solutions, most rely on additional training with augmented data, which is computationally expensive and remains limited in preserving visual subjects, temporal consistency, and physical plausibility. In this work, we propose self-correction optimization (SCO), an effective training-free method for consistent interleaved generation. SCO treats the classifier-free guidance update as a reference and performs minimal self-correction under two complementary constraints, including new-event and state-preserving constraints. Specifically, the new-event constraint promotes temporal consistency across image--text sequences, while the state-preserving constraint maintains the coherence of visual subjects throughout subsequent generation steps. Experiments on challenging interleaved multimodal generation benchmarks demonstrate significant improvements in temporal coherence and visual-subject preservation. Furthermore, SCO can be extended to video generation and improves the modeling of physically grounded processes, including robot manipulation and long-horizon handcrafting.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generation

    Jun 29, 2026Chonghuinan Wang, Zhikai Chen, Chunwei Wang +9Unified Multimodal ModelsMultimodal Model Evaluation

  2. InterleaveThinker: Reinforcing Agentic Interleaved Generation

    Jun 11, 2026Dian Zheng, Harry Lee, Manyuan Zhang +4Agentic Image GenerationMultimodal Generation