Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models
Organizations: Georgia Institute of Technology · Washington University in St. Louis
Abstract
Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We introduce MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains. MATE lets generation and understanding take turns to be challenger and solver. Given an image, the understanding branch proposes several candidate descriptions that the generation branch must turn back into similar images, and vice versa. The candidates are screened for consistency with the image or prompt they were proposed from, and the solver is trained on the candidate it handles worst. The adversary thus comes from the model's own outputs, and no separate adversary is trained. Moreover, the candidates that defeat one branch become the sources of the next challenges to the other in the next epoch, which keeps the challenges evolving with the model and turns the training into self-play in data space. On Janus-Pro-1B, MATE improves GenEval by 2.4 points, DPG-Bench by 1.7 points, and the average over nine understanding benchmarks by 0.7 points, while strengthening consistency across repeated image-text cycles.
Figures & tables
| Method | # Params | GenEval (per category) | GenEval | DPG-Bench | |||||
| Single Obj. | Two Obj. | Counting | Colors | Position | Color Attr. | Overall | Overall | ||
| Generation only | |||||||||
| SD3-Medium ( Esser et al., 2024 ) | 2B | 99.0 | 94.0 | 72.0 | 89.0 | 33.0 | 60.0 | 74.0 | 84.1 |
| LlamaGen ( Sun et al., 2024 ) | 0.8B | 71.0 | 34.0 | 21.0 | 58.0 | 7.0 | 4.0 | 32.0 | – |
| Sana-0.6B ( Xie et al., 2025a ) | 0.6B | 99.0 | 76.0 | 64.0 | 88.0 | 18.0 | 39.0 | 64.0 | 83.6 |
| Sana-1.6B ( Xie et al., 2025a ) | 1.6B | 99.0 | 77.0 | 62.0 | 88.0 | 21.0 | 47.0 | 66.0 | 84.8 |
| Method | # Params | POPE | MME-C | MMB | SEED | MMStar | V ∗ Bench | CV-2D | CV-3D | SEED2+ | Avg. |
| Understanding only | |||||||||||
| LLaVA-v1.5 ( Liu et al., 2024 ) | 7B | 86.1 | 302.1 | 72.9 | 65.8 | 33.1 | 45.0 | 62.0 | 64.3 | 41.3 | 56.5 |
| DeepSeek-VL-1.3B ( Lu et al., 2024 ) | 1.3B | 85.9 | 225.0 | 74.5 | 66.0 | 39.9 | 42.9 | 66.2 | 62.7 | 43.7 | 56.7 |
| LLaVA-OV-0.5B ( Li et al., 2025 ) | 0.5B | 87.8 | 237.5 | 67.1 | 66.6 | 37.7 | 39.3 | 38.6 | 50.6 | 52.6 | 52.2 |
| UMMs | |||||||||||
| Janus ( Wu et al., 2025a ) | 1.3B | 84.1 | 238.6 | 70.0 | 62.9 | 37.6 | 41.4 | 49.1 | 58.7 | 38.5 | 52.5 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Other branch supplies | Tasks optimized | Training inputs | Hard cases selected for | Added module |
| SUDER ( Hong et al., 2025 ) | likelihood | , | corpus | – | – |
| GvU ( Pan et al., 2026 ) | likelihood | corpus | – | – | |
| UniRL ( Mao et al., 2025 ) | answers | , | constructed prompts | – | – |
| SILMM ( Qu et al., 2025 ) | answers | model-written prompts | – | – | |
| HermesFlow ( Yang et al., 2025 ) | answers | , | corpus | – | – |
| SRUM ( Jin et al., 2026 ) | ratings | corpus | – | – |
| Training | |||
| Candidates per input | 4 | Learning rate | |
| Rollouts per candidate | 8 | GRPO clipping range | 0.2 |
| Admission threshold | 0.5 | Reference penalty | 0.01 (low-variance KL) |
| Standard-deviation floor | Trainable parameters | all (2.1B) | |
| Epochs | 5 | Hardware | 6 NVIDIA B200 |
| Effective batch size | 48 (24 inputs per step) | Consistency weight | 0.1 |
| Und. Avg. | Gen. Avg. | |
| Janus-Pro-1B | 56.9 | 78.3 |
| MATE | 57.6 | 80.3 |
| w/o consistency term | 57.5 | 80.5 |
| All admitted candidates | 57.5 | 80.6 |
| Two separate networks | 66.3 | 64.1 |
| vs. own base models | 0.9 | 2.1 |