cs.CVSep 28, 2026

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Authors: Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu

Organizations: Nanyang Technological University · Shanghai Jiao Tong University · The University of Tokyo

Abstract

Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Apr 15, 2026cs.CV

UniRect-CoT: Enhancing Generation in Unified Multimodal Models via Reflective Rectification with Inherent Understanding

Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability often outperforms their generation capability. This mismatch suggests that the model's rich internal knowledge, while effective for understanding tasks, is not fully utilized during generation. To address this, we draw inspiration from the human ``Thinking-While-Drawing'' paradigm, where humans continuously reflect on intermediate results and revise them according to their understanding of the intended target. In this paper, we propose UniRect-CoT, a training-free unified rectification chain-of-thought framework. We regard the multi-step denoising process in UMMs as an intrinsic visual reasoning process, whose intermediate states provide opportunities for continuous reflection and rectification. By leveraging the UMM's inherent understanding to reflect on intermediate results in light of the target instruction and rectify them accordingly, our framework forms a reflective chain of thought within a single denoising trajectory, enabling the model's internal knowledge to guide generation throughout the process. Specifically, the UMM generates a state-conditioned reflection, whose autoregressive negative log-likelihood under the same model defines a model-native differentiable semantic loss that translates textual reflection into gradients for latent rectification. Extensive experiments demonstrate that UniRect-CoT generalizes across existing flow-based UMMs, consistently enhancing overall text-to-image generation performance.
Apr 28, 2026cs.CV

Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models

Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified capability allows UMMs to refine outputs after their initial generation, potentially extending the performance upper bound. Current UMM-based refinement methods primarily follow a refinement-via-editing (RvE) paradigm, where UMMs produce editing instructions to modify misaligned regions while preserving aligned content. However, editing instructions often describe prompt-image misalignment only coarsely, leading to incomplete refinement. Moreover, pixel-level preservation, though necessary for editing, unnecessarily restricts the effective modification space for refinement. To address these limitations, we propose Refinement via Regeneration (RvR), a novel framework that reformulates refinement as conditional image regeneration rather than editing. Instead of relying on editing instructions and enforcing strict content preservation, RvR regenerates images conditioned on the target prompt and the semantic tokens of the initial image, enabling more complete semantic alignment with a larger modification space. Extensive experiments demonstrate the effectiveness of RvR, improving Geneval from 0.78 to 0.91, DPGBench from 84.02 to 87.21, and UniGenBench++ from 61.53 to 77.41.
Sep 8, 2025cs.CV

Reconstruction Alignment Improves Unified Multimodal Models

Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture. However, conventional training relies on image-text pairs (or sequences) whose captions are typically sparse and miss fine-grained visual details, even when they use hundreds of words to describe a simple image. We introduce Reconstruction Alignment (RECA), a resource-efficient post-training method that leverages visual understanding encoder embeddings as dense "text prompts", providing rich supervision without captions. Concretely, RECA conditions a UMM on its own visual understanding embeddings and optimizes it to reconstruct the input image with a self-supervised reconstruction loss, thereby realigning understanding and generation. Despite its simplicity, RECA is broadly applicable: across autoregressive, masked-autoregressive, and diffusion-based UMMs, it consistently improves generation and editing fidelity. With only 27 GPU hours, post-training with RECA substantially improves image generation performance on GenEval (0.73 →\rightarrow 0.90) and DPGBench (80.93 →\rightarrow 88.15), while also boosting editing benchmarks (ImgEdit 3.38 →\rightarrow 3.75, GEdit 6.94 →\rightarrow 7.27). Notably, RECA surpasses much larger open-source models and applies broadly across diverse UMM architectures, establishing it as an efficient and general post-training alignment strategy for UMMs.