UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
Authors: Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, +11 more
Organizations: Stanford University · Johns Hopkins University · Independent Researcher · University of Toronto · University of Oxford · UC, Riverside · MatrAIx · University of Notre Dame · Carnegie Mellon University · UNC–Chapel Hill
Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts. The student only sees the vanilla question, while the teacher conditions on the privileged critique. Then training minimizes the per-state divergence between their denoising diffusion distributions over the student's own sampling trajectories. Experiments demonstrate that UniEvo-VL improves the image generation capabilities of multimodal models, while maintaining their sensitivity to additional reflection information. Specifically, we build on top of the open-source Qwen-image-2512 and observe a significant performance gain from 0.747 to 0.808 on GenEval and from 32.97 to 35.53 on GenEval2 Soft-TIFA. Moreover, attempts with more powerful external critics (e.g., GPT5.6-Luna) show that multimodal models with strong judge capabilities can anticipate a higher self-evolving ceiling. Last but not least, mixed text-rendering outcomes show that our self-improvements may not be uniform across different tasks. Our study aims to shed light on the current hot recursive self-improvement research line to enhance the user experience when using multimodal models without external supervision or guidance.
Figures & tables
Figure 1: From understanding to self-evolution. Generated drafts become training experience through visual critique and generator updates. The conceptual loop motivates both self-critique and an optional extension with a stronger external critic.
Figure 2: The UniEvo-VL training loop. A generated draft is inspected along semantic, text, and visual-quality axes. The resulting feedback is synthesized into a revised prompt, which acts as the privileged condition for the EMA teacher. Transition matching updates the student’s generator LoRA; the critic is fixed. At deployment, the evolved model uses the original prompt directly.
Figure 3: UniEvo-VL with revision verification across training updates. Solid curves show direct generation; dashed curves allow one inference-time critique. Lines connect evaluated checkpoints across runs; n counts plotted points.
GenEval
GenEval2
OCR
Method
Native
Atomic (%)
HumanPref
Native
Atomic (%)
HumanPref
Native
Atomic (%)
HumanPref
Base
0.747
95.30
8.48
32.58
81.69
6.65
0.771
–
9.07
UniEvo-VL
0.808
96.90
8.79
32.37
82.24
6.91
0.761
–
9.06
UniEvo-VL (GPT5.6-Luna)
0.882
98.76
8.97
35.53
83.18
6.86
0.775
–
9.03
UniEvo-VL (verification)
0.818
97.41
8.71
35.07
82.75
7.11
0.790
–
9.10
Table 1: Performance of UniEvo-VL. Parentheses identify the critic; “+ verification” denotes revision-prompt filtering. Base is averaged over different random seeds of the identical pretrained model. Bold and underlined values indicate the best and second-best for each dataset and metric.
Figure 4: Improvement and preservation on historical 500-prompt subsets, where scale is normalized to 0–100. The percentage-point changes are subset-specific, not full-test-set gains.
(a)
Base model
UniEvo-VL
Dataset
Metric
Direct
+ Reflection
Direct
+ Reflection
GenEval
Native
0.748
0.826
0.818
0.848
Atomic
95.41
98.26
97.41
98.56
HumanPref
8.47
8.92
8.71
8.87
GenEval2
Native
31.92
42.52
35.07
46.23
Atomic
81.47
84.45
82.75
86.12
Table 2: Direct generation and reflection before and after training. (a) Generation scores, where bold marks the highest score in each row. (b) Training gain is the UniEvo-VL score minus the base score under direct generation, while reflection gain is reflected minus direct within each model. Gains use unrounded scores. Native scores use 0–1 for GenEval/OCR and 0–100 for GenEval2. Atomic uses percent (gains in percentage points), and HumanPref uses 0–10.
Critic / model
Single
Two
Count
Color
Position
Attr.
Overall
Base
9.93
9.68
7.86
9.71
6.85
9.75
8.96
Qwen-VL
9.90
9.88
9.29
9.66
8.79
9.94
9.57
GPT-5.6-Luna †
9.98
9.92
9.58
9.91
9.88
9.94
9.87
Table 3: Compositional generation under different critics. All entries are Gemini compositional-fidelity ratings (0–10). (a) GenEval uses 553 prompts; Base denotes the Qwen-VL experiment’s baseline. (b) GenEval2 uses 800 prompts; Δ is Evolved minus Base. Bold indicates the highest score in each column of (a), and the higher Evolved score and larger Δ within each row of (b), including ties.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
GenEval
GenEval2
OCR
Feedback backend
Qwen3-VL
Qwen3-VL
GPT-5.6-Luna
Qwen3-VL
GPT-5.6-Luna
Critic parameters
8B
8B
-
8B
-
Generator
Qwen-Image-2512
Resolution / steps / CFG
10242 / 20 / 4.0
LoRA rank / alpha
16 / 16
Learning rate
3×10−4
1×10−4
1×10−4
3×10−4
3×10−4
Appendix
Table 4: Training configuration for acquisition without post-revision verification.
Setting
Value
Generator
Qwen-Image-2512
Feedback backend
Qwen3-VL 8B, short-panel critique
Resolution / steps / CFG
10242 / 20 / 4.0
LoRA rank / alpha
16 / 16
Optimizer / learning rate
AdamW / 3×10−4
Adam coefficients / epsilon
(0.9,0.999) / 10−8
Appendix
Table 5: Training configuration for acquisition with post-revision verification on all three tasks. Gradient accumulation operates over transition-loss evaluations.
Figure 5: EMA teacher decay on GenEval. Task and HumanPref scores across three EMA decays. Both metrics favor 0.995 and 0.999 over 0.99, with 0.995 achieving the highest scores in this sweep.
Task
Unique prompts
Images
Task-scoring protocol
GenEval
553
2,212
8,392 binary checks
GenEval2
800
800
6,012 binary checks
OCR
1,018
1,018
Edit distance
Appendix
Table 6: Full-cohort Gemini evaluation for the verified recipe. Image counts are per checkpoint and inference condition. HumanPref is evaluated separately on each corresponding image pool.
Method
Single
Two
Count
Color
Position
Attr.
Overall
SD3.5-L
0.98
0.89
0.73
0.83
0.34
0.47
0.71
HiDream-I1
1.00
0.98
0.79
0.91
0.60
0.72
0.83
Z-Image
1.00
0.94
0.78
0.93
0.62
0.77
0.84
SANA-1.5
0.99
0.93
0.86
0.84
0.59
0.65
0.81
LongCat
0.99
0.98
0.86
0.86
0.75
0.73
0.87
BAGEL
0.99
0.94
0.81
0.88
0.64
0.63
0.82
Appendix
Table 7: GenEval performance on 553 prompts. Scores use the official evaluator (0–1; higher is better). Our update-200 rows use training without post-revision verification; update-140 rows use training with post-revision verification. Each checkpoint is evaluated with direct generation and with one critique iteration. The Qwen-Image-2512 and update-200 rows use one evaluated image per prompt; update-140 rows use four images per prompt (2,212 images). Bold and underlined values indicate the highest and second-highest scores in each column, with ties at the displayed precision.
GenEval
GenEval2
OCR
Step
Task
HumanPref
Task
HumanPref
Task
HumanPref
0
8.958
8.503
6.228
6.716
9.402
9.074
40
9.128
8.526
6.277
6.726
9.307
9.025
80
9.134
8.544
6.321
6.737
9.364
9.054
120
9.316
8.671
6.356
6.824
9.301
9.042
160
9.421
8.759
6.320
6.732
9.377
9.064
Appendix
Table 8: Full-set learning curves without post-revision verification, rounded to three decimals. Each checkpoint uses 553 GenEval, 800 GenEval2, and 1,018 OCR prompts; task and HumanPref scores retain their 0–10 scale.
Figure 6: Complete GenEval2 Soft-TIFA trajectory for GPT-5.6-Luna critique, updates 0–300. The Qwen-critic baseline and update-200 endpoints are shown as separate markers. GM and AM use each run’s own baseline.
Soft-TIFA GM
Soft-TIFA AM
Critic
Base
Evolved
Δ
Base
Evolved
Δ
Qwen-VL
32.86
32.37
− 0.49
78.06
77.99
− 0.07
GPT-5.6-Luna
32.97
35.53
+2.56
78.13
79.25
+1.12
Appendix
Table 9: Paired native GenEval2 results at update 200 (800 prompts). Deltas use unrounded scores.
Figure 7: Qualitative examples of UniEvo-VL on GenEval2. Examples cover object counts, colors, materials, and patterns. Rows show the base model, UniEvo-VL direct generation, and UniEvo-VL with up to one critique. Each column shares the original prompt and seed; both UniEvo-VL rows use checkpoint 160 of the Qwen critic-confirmed variant. A KEEP decision retains the direct image.
Figure 8: Qualitative examples of UniEvo-VL on OCR. Examples illustrate text rendering across varied scenes; headings summarize the scene and quote the requested text. Rows show the base model, UniEvo-VL direct generation, and UniEvo-VL with up to one critique. Each column shares the original prompt and seed; both UniEvo-VL rows use checkpoint 130 of the Qwen critic-confirmed variant. A KEEP decision retains the direct image.