Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability often outperforms their generation capability. This mismatch suggests that the model's rich internal knowledge, while effective for understanding tasks, is not fully utilized during generation. To address this, we draw inspiration from the human ``Thinking-While-Drawing'' paradigm, where humans continuously reflect on intermediate results and revise them according to their understanding of the intended target. In this paper, we propose UniRect-CoT, a training-free unified rectification chain-of-thought framework. We regard the multi-step denoising process in UMMs as an intrinsic visual reasoning process, whose intermediate states provide opportunities for continuous reflection and rectification. By leveraging the UMM's inherent understanding to reflect on intermediate results in light of the target instruction and rectify them accordingly, our framework forms a reflective chain of thought within a single denoising trajectory, enabling the model's internal knowledge to guide generation throughout the process. Specifically, the UMM generates a state-conditioned reflection, whose autoregressive negative log-likelihood under the same model defines a model-native differentiable semantic loss that translates textual reflection into gradients for latent rectification. Extensive experiments demonstrate that UniRect-CoT generalizes across existing flow-based UMMs, consistently enhancing overall text-to-image generation performance.
Figures & tables
Figure 1: Capability mismatch in UMMs and overview of UniRect-CoT. (a) A UMM may fail to follow the generation instruction while its understanding branch can identify the error and describe the desired correction. (b) UniRect-CoT realizes “Thinking-While-Drawing” by interleaving state-conditioned reflection and rectification within the same denoising trajectory.
Figure 2: Overall pipeline of UniRect-CoT. (a) Reflection and rectification are interleaved with denoising within a selected rectification window. (b) At each rectification timestep, GITO performs iterative trajectory exploration and greedy selection. ISR combines state-conditioned reflection with likelihood-guided rectification. The reflective target is generated in the first ISR update and reused by subsequent updates at the same timestep. The resulting candidates, including the original state, are evaluated using the user instruction to determine the state retained for subsequent denoising.
Figure 3: Judgment accuracy across denoising timesteps. We measure how often each UMM’s judgment of whether a look-ahead image satisfies the user instruction agrees with the GenEval correctness label for the same image.
DPG ↑
GenEval ↑
Model
Score
Single Obj.
Two Obj.
Counting
Colors
Position
Color Attri.
Overall
TokenFlow
73.4
0.97
0.66
0.40
0.84
0.17
0.26
0.55
OpenUni
79.0
0.99
0.71
0.55
0.82
0.25
0.42
0.62
Emu3
80.6
0.98
0.71
0.34
0.81
0.17
0.21
0.54
Show-o2
85.0
0.99
0.86
0.55
0.86
0.46
0.63
0.73
Janus
79.6
0.97
0.68
0.30
0.84
0.46
0.42
0.61
Table 1: Quantitative comparison on GenEval and DPG-Bench. We compare our method with state-of-the-art unified multimodal models (top) and verify the effectiveness of our method on standard baselines (bottom). Results marked with † are reproduced in this work by generating 4 consecutive samples per prompt, starting with a fixed initial seed. ↑ denotes that higher is better. We highlight improvements over the baseline and the best overall performance.
Configuration
GenEval ↑
Method
Pre
Multi
Traj.
State
Single Obj.
Two Obj.
Counting
Colors
Position
Color Attri.
Overall
Training-free test-time scaling on BAGEL
BAGEL
✗
✗
✗
✗
99.1
95.7
74.3
86.7
49.0
61.0
77.6
Qwen Rewrite †
✓
✗
✗
✗
100.0
95.7
73.4
89.9
47.3
64.0
78.4
UiG †
✓
✓
✗
✓
99.1
95.7
75.3
87.8
48.8
61.5
78.1
TiR †
✓
✓
✗
✓
98.4
95.7
80.0
83.5
54.0
67.2
79.8
Table 2: Comparison of test-time strategies on BAGEL. “Pre”: pre-generation reasoning or prompt enhancement; “Multi”: multiple full generation/refinement rounds; “Traj.”: intervention during denoising; “State”: feedback conditioned on the current visual result. † marks results reproduced under our unified inference and evaluation protocol. Boldface highlights results with UniRect-CoT.
Figure 4: Qualitative comparison. Each pair shows BAGEL (left) and UniRect-CoT (right).
Figure 5: Ablation of rectification strategy and efficiency Analysis. (a) Effect of the number of rectification updates K and the Greedy Selection Strategy (GSS). (b) Effect of the rectification window W at K=3 . (c) GenEval Overall versus relative latency for UniRect-CoT and competing test-time strategies. Latency is normalized to standard BAGEL inference ( 1.0× , 24.7 s/image).
Configuration
GenEval ↑
Setting
Update
ITE Target
GSS Target
Single Obj.
Two Obj.
Counting
Colors
Position
Color Attri.
Overall
Baseline
–
–
–
99.1
95.7
74.3
86.7
49.0
61.0
77.6
ITE only
Static
User
–
99.7
93.9
76.3
86.7
48.5
67.0
78.7
Static
Exp.
–
99.7
94.2
76.9
87.2
48.8
60.5
77.9
Static
Ref.
–
100.0
94.7
77.8
86.2
46.5
69.3
79.1
Dynamic
Ref.
–
100.0
95.5
80.3
87.8
46.0
69.0
79.7
Table 3: Ablation of semantic target design. “User”, “Exp.”, and “Ref.” denote the original instruction cuser , its text-only expansion, and the state-conditioned reflective target ctargett , respectively. Static targets remain fixed across the rectification window; dynamic targets are regenerated at each rectification timestep and reused for its K updates. “–” indicates that the corresponding component is disabled.
Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely decoupled understanding and generation components. This design, while effective for individual tasks, weakens the connection required for mutual enhancement, leaving the potential synergy empirically uncertain. We propose to explicitly restore this synergy by introducing Understanding-Oriented Post-Training (UNO), a lightweight framework that treats understanding not only as a distinct task, but also a direct supervisory signal to steer generative representations. By incorporating objectives that encode semantic abstraction (captioning) and structural details (visual regression), we enable effective gradient flow from understanding to generation. Extensive experiments on image generation and editing demonstrate that understanding can serve as an effective catalyst for generation.
Zeyu Liu, Zanlin Ni, Yang Yue +5
1Tsinghua University · 2Kolors Team, Kuaishou Technology
Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified capability allows UMMs to refine outputs after their initial generation, potentially extending the performance upper bound. Current UMM-based refinement methods primarily follow a refinement-via-editing (RvE) paradigm, where UMMs produce editing instructions to modify misaligned regions while preserving aligned content. However, editing instructions often describe prompt-image misalignment only coarsely, leading to incomplete refinement. Moreover, pixel-level preservation, though necessary for editing, unnecessarily restricts the effective modification space for refinement. To address these limitations, we propose Refinement via Regeneration (RvR), a novel framework that reformulates refinement as conditional image regeneration rather than editing. Instead of relying on editing instructions and enforcing strict content preservation, RvR regenerates images conditioned on the target prompt and the semantic tokens of the initial image, enabling more complete semantic alignment with a larger modification space. Extensive experiments demonstrate the effectiveness of RvR, improving Geneval from 0.78 to 0.91, DPGBench from 84.02 to 87.21, and UniGenBench++ from 61.53 to 77.41.
Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture. However, conventional training relies on image-text pairs (or sequences) whose captions are typically sparse and miss fine-grained visual details, even when they use hundreds of words to describe a simple image. We introduce Reconstruction Alignment (RECA), a resource-efficient post-training method that leverages visual understanding encoder embeddings as dense "text prompts", providing rich supervision without captions. Concretely, RECA conditions a UMM on its own visual understanding embeddings and optimizes it to reconstruct the input image with a self-supervised reconstruction loss, thereby realigning understanding and generation. Despite its simplicity, RECA is broadly applicable: across autoregressive, masked-autoregressive, and diffusion-based UMMs, it consistently improves generation and editing fidelity. With only 27 GPU hours, post-training with RECA substantially improves image generation performance on GenEval (0.73 → 0.90) and DPGBench (80.93 → 88.15), while also boosting editing benchmarks (ImgEdit 3.38 → 3.75, GEdit 6.94 → 7.27). Notably, RECA surpasses much larger open-source models and applies broadly across diverse UMM architectures, establishing it as an efficient and general post-training alignment strategy for UMMs.
Ji Xie, Trevor Darrell, Luke Zettlemoyer +1
1UC Berkeley · University of Washington · 3Duke University