cs.CVSep 28, 2026

Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models

Authors: Wentao Zhou, Weijie Gan, Jiayun Wang

Organizations: Georgia Institute of Technology · Washington University in St. Louis

Abstract

Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We introduce MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains. MATE lets generation and understanding take turns to be challenger and solver. Given an image, the understanding branch proposes several candidate descriptions that the generation branch must turn back into similar images, and vice versa. The candidates are screened for consistency with the image or prompt they were proposed from, and the solver is trained on the candidate it handles worst. The adversary thus comes from the model's own outputs, and no separate adversary is trained. Moreover, the candidates that defeat one branch become the sources of the next challenges to the other in the next epoch, which keeps the challenges evolving with the model and turns the training into self-play in data space. On Janus-Pro-1B, MATE improves GenEval by 2.4 points, DPG-Bench by 1.7 points, and the average over nine understanding benchmarks by 0.7 points, while strengthening consistency across repeated image-text cycles.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

    Jun 25, 2026Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar +5Multimodal GenerationLarge Multimodal Models

  2. DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement

    May 25, 2026Renjie Lu, Xulong Zhang, Xiaoyang Qu +2Multimodal ModelInductive Bias

  3. UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement

    Sep 30, 2026Fang Wu, Da Xing, Yanjie Huang +16Multimodal ModelSelf-Evolution