Unified multimodal models can both look at and render images, so in principle they can repair their own generations: diagnose what an image gets wrong, revise it, observe the result, and diagnose again. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths, and naive RL that optimizes only the renderer or only one head leaves most of the gain untapped. We introduce UMM-Reflection, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment. Unlike single-round editing or pipelines with an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference. On BAGEL, UMM-Reflection improves GenEval by 12.05 points over SFT, and the gains transfer to WISE (+10.97), OneIG-Bench (+3.48), and T2I-CompBench++ (+4.63), none of which is used in training.
Figures & tables
Figure 1 : Native reflection before and after RL. Left: SFT and RL on the same prompt and seed. SFT already produces meaningful revisions (16 SFT rollouts contain a correct repair for 78% of failing training images; Appendix I ), but one trajectory often circles, as here (tie, dog, tie). RL concentrates the policy on revisions that reach the correct region (right; measured in Figure 5 ).
Figure 2 : UMM-Reflection RL. (1) K=16 rollouts share one detached initial image x0 . (2) Each interleaves the model’s own reflection (verbatim) with its renders for up to three rounds. (3) A frozen verifier scores every image ( q ) and trajectory ( R(τ) ). (4) Group normalization gives one advantage Ai per trajectory, (5) which updates both the text and flow heads. Green/red frames: verifier pass/fail; dashed: training only.
Figure 3 : Case studies across four benchmarks. Each row shows the prompt, the initial image, and three reflection-guided revisions (left to right), with examples from GenEval, WISE, OneIG, and CompBench. The model identifies spatial, color, material, and compositional errors through its [THINKING] output and issues targeted edits.
Figure 4 : Test-time scaling on GenEval. Macro accuracy versus reflection rounds. Other benchmarks: Appendix Figure 9 .
Figure 5 : Image trajectories in the backbone’s own correctness readout. Each point is one initially failing image, placed by two held-out linear probes for “passes the verifier”: the understanding stream at depth 20 ( x ) and the generation stream at depth 8 ( y ), the depths where each stream’s probe AUC peaks. Green filled contours mark verifier-passing reference images and red dashed contours verifier-failing ones; orange and blue contours summarize the RL and SFT images. Read the panels in pairs, from R0 to R3 for the same policy: RL (left pair) and SFT (right pair). Percentages are the share inside the dense pass region (Table 5 ).
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Backbone
BAGEL, 28-layer mixture-of-transformers (MoT)
Initialization
Reflection SFT (one epoch)
Reported RL checkpoint
1,000 committed updates, direct full weights
Direct T2I-RL control
Official Base, 1,000 RL updates, no repairs
Train/dev prompt pools
3,000 / 270
Roots per update / siblings per root
2 / 16
Appendix
Table 4 : Primary configuration and evaluation coverage.
Figure 6 : The supervised parent, recorded continuation phase. Text cross-entropy and image flow-MSE from global SFT updates 1410–2969. Faint lines show logged update values and solid lines show trailing 50-update means.
Figure 7 : Completed direct-T2I RL control. Hollow orange points denote direct T2I-RL at 1,000 updates; green points denote UMM-Reflection at 1,000 updates. Gain is the within-benchmark score difference. Each benchmark keeps its own metric on a 0 – 100 scale; scores are not averaged across benchmarks. The systems differ in initialization and training/inference compute (Appendix A ).
Figure 8 : Learning dynamics of the 1,000-update RL run. (a) Mean trajectory reward. (b) Successful repairs per edit from an incorrect state. (c) Fraction of trajectories reaching terminal exactness under the training verifier. Faint lines are individual updates; solid lines are trailing 20-update trends, with rates pooling numerators and denominators over the window.
Figure 9 : Multi-round test-time scaling on four benchmarks. Scores ( 0 – 100 ) versus reflection rounds. Blue dotted lines denote Base, pink dashed lines reflection SFT, purple dash-dotted lines UMM-Reflection -500, and green solid lines UMM-Reflection -1000.
Figure 10 : Attention over VAE image tokens is preserved after RL. Three cases show paired SFT/RL attention of THINKING and EDIT tokens onto 32×32 VAE keys, averaged over 28 layers. Maps use a shared 99th-percentile cap; cyan boxes are pre-annotated error regions.
Policy
Tokens
x0
Revision 1
Current image
SFT
THINKING
0.23 (0.20)
0.30 (0.26)
0.48 (0.41)
SFT
EDIT
0.23 (0.19)
0.31 (0.27)
0.46 (0.41)
RL
THINKING
0.26 (0.23)
0.29 (0.25)
0.45 (0.39)
RL
EDIT
0.23 (0.19)
0.31 (0.27)
0.46 (0.40)
RL, VAE keys only
THINKING
0.32 (0.27)
0.34 (0.28)
0.34 (0.26)
RL, VAE keys only
EDIT
0.32 (0.27)
0.35 (0.29)
0.33 (0.26)
Appendix
Table 6 : Share of image attention per image at the third reflection round. Mean over ten trajectories (minimum in parentheses); shares sum to 1 across the three images. Image attention is 10–13% of all attention.
RL updates
Roots
pass@1 (%)
pass@16 (%)
1–50
67
22.7
77.6
51–100
54
31.4
77.8
101–200
114
31.5
83.3
201–500
367
60.5
94.6
501–1,000
611
70.1
97.1
Appendix
Table 7 : Sibling success on incorrect initial images during RL. Roots with an incorrect initial image; 16 sibling trajectories per root.
Comparison with SFT
Wins
Losses
Net
Exact p
Initial
48
32
+16
0.093
Final
109
38
+71
<0.001
Appendix
Table 9 : Prompt-paired UMM-Reflection -1000 versus reflection SFT on GenEval. Wins and losses count discordant image-only verdicts. The exact two-sided binomial test on discordant pairs is the exact McNemar test.
Family
n
SFT final
RL initial
RL final
Single object
80
100.00
100.00
95.00
Two objects
99
87.88
83.84
95.96
Counting
80
57.50
67.50
67.50
Colors
94
87.23
84.04
90.43
Position
100
47.00
55.00
89.00
Color binding
100
51.00
48.00
65.00
Appendix
Table 10 : All six GenEval families, using image-only verdicts. Values are percentages; the macro mean appears in Table 1 .
Figure 11 : Category-level changes, with every category retained. Positive changes are green and negative changes are red. This view complements the overall gains without assuming uniform improvement.
Model
GenEval
WISE
OneIG
CompBench
Repair (%)
BAGEL-Base
0.71
0.58
0.80
0.49
–
BAGEL + GPT-5.5 critic
0.79
0.70
0.83
0.52
27.6
UMM-Reflection
0.84
0.75
0.83
0.55
64.9
Appendix
Table 15 : External GPT-5.5 critic versus UMM-Reflection . Native 0 – 1 scale. WISE in this table is scored by a separate GPT-4o run for all three rows. OneIG to three decimals: 0.825 (critic) versus 0.829 ( UMM-Reflection ). Repair is the share of initially incorrect GenEval images that end correct.
Figure 12 : Reflection trajectories of UMM-Reflection . Examples from GenEval, WISE, and CompBench show how step-by-step reflection and revision help the model produce images that match the prompt. Each row reads left to right: the image, then the unified model’s own reflection on it ( Think : what is wrong; Edit : the instruction it issues), then the image it renders from that instruction. All text is verbatim model output.
Figure 13 : Successful repairs. Each column reads top to bottom: the image, the model’s verbatim reflection, and the image it renders next. Frames mark the benchmark verdict. Cases are selected.
Figure 14 : Successful repairs (continued); layout as in Figure 13 .
Figure 15 : Successful repairs (continued); layout as in Figure 13 .
Figure 16 : Successful repairs (continued); layout as in Figure 13 .
Figure 17 : Successful repairs (continued); layout as in Figure 13 .
Figure 18 : Successful repairs (continued); layout as in Figure 13 .
Figure 19 : Failure cases. Failures caused by the reflection. Left: the first image already shows four giraffes and passes, but the reflection counts them as three and asks for one more; it repeats “three giraffes” in every later round while the image holds five. Right: the reflection states that the second largest economy is the United States, so all three edits render US dollars instead of the Chinese yuan. Frames mark the benchmark verdict.
Figure 20 : Training reward by GenEval family. Light points are the mean whole-trajectory reward of the 16 rollouts for one prompt; lines are 100-update means. Prompt draws per family follow the training pool composition; single object has six draws and its line is left open.
Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability often outperforms their generation capability. This mismatch suggests that the model's rich internal knowledge, while effective for understanding tasks, is not fully utilized during generation. To address this, we draw inspiration from the human ``Thinking-While-Drawing'' paradigm, where humans continuously reflect on intermediate results and revise them according to their understanding of the intended target. In this paper, we propose UniRect-CoT, a training-free unified rectification chain-of-thought framework. We regard the multi-step denoising process in UMMs as an intrinsic visual reasoning process, whose intermediate states provide opportunities for continuous reflection and rectification. By leveraging the UMM's inherent understanding to reflect on intermediate results in light of the target instruction and rectify them accordingly, our framework forms a reflective chain of thought within a single denoising trajectory, enabling the model's internal knowledge to guide generation throughout the process. Specifically, the UMM generates a state-conditioned reflection, whose autoregressive negative log-likelihood under the same model defines a model-native differentiable semantic loss that translates textual reflection into gradients for latent rectification. Extensive experiments demonstrate that UniRect-CoT generalizes across existing flow-based UMMs, consistently enhancing overall text-to-image generation performance.
Yibo Jiang, Tao Wu, Rui Jiang +4
School of Software Technology, Zhejiang University · College of Computer Science and Technology, Zhejiang University
Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified capability allows UMMs to refine outputs after their initial generation, potentially extending the performance upper bound. Current UMM-based refinement methods primarily follow a refinement-via-editing (RvE) paradigm, where UMMs produce editing instructions to modify misaligned regions while preserving aligned content. However, editing instructions often describe prompt-image misalignment only coarsely, leading to incomplete refinement. Moreover, pixel-level preservation, though necessary for editing, unnecessarily restricts the effective modification space for refinement. To address these limitations, we propose Refinement via Regeneration (RvR), a novel framework that reformulates refinement as conditional image regeneration rather than editing. Instead of relying on editing instructions and enforcing strict content preservation, RvR regenerates images conditioned on the target prompt and the semantic tokens of the initial image, enabling more complete semantic alignment with a larger modification space. Extensive experiments demonstrate the effectiveness of RvR, improving Geneval from 0.78 to 0.91, DPGBench from 84.02 to 87.21, and UniGenBench++ from 61.53 to 77.41.
Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture. However, conventional training relies on image-text pairs (or sequences) whose captions are typically sparse and miss fine-grained visual details, even when they use hundreds of words to describe a simple image. We introduce Reconstruction Alignment (RECA), a resource-efficient post-training method that leverages visual understanding encoder embeddings as dense "text prompts", providing rich supervision without captions. Concretely, RECA conditions a UMM on its own visual understanding embeddings and optimizes it to reconstruct the input image with a self-supervised reconstruction loss, thereby realigning understanding and generation. Despite its simplicity, RECA is broadly applicable: across autoregressive, masked-autoregressive, and diffusion-based UMMs, it consistently improves generation and editing fidelity. With only 27 GPU hours, post-training with RECA substantially improves image generation performance on GenEval (0.73 → 0.90) and DPGBench (80.93 → 88.15), while also boosting editing benchmarks (ImgEdit 3.38 → 3.75, GEdit 6.94 → 7.27). Notably, RECA surpasses much larger open-source models and applies broadly across diverse UMM architectures, establishing it as an efficient and general post-training alignment strategy for UMMs.
Ji Xie, Trevor Darrell, Luke Zettlemoyer +1
1UC Berkeley · University of Washington · 3Duke University