Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes. Existing few-step distillation methods largely focus on either image generation or text generation, making it unclear how to compress a fully discrete multimodal dLLM into a single efficient student while preserving both generation and understanding. We introduce Omni-Diffusion-Distill, a unified two-stage distillation framework that retains strong generation and understanding capabilities while substantially reducing the inference cost of a unified multimodal dLLM. Omni-Diffusion-Distill aligns the distillation of both generation and understanding, for both images and text, in the discrete token space. In the first stage, the student is trained to skip decoding steps by replaying cached teacher trajectories, and in the second stage the student is refined on intermediate states along its own rollouts. We further remedy two sources of degradation in unified distillation with a pairwise collision penalty that reduces repetition under parallel text decoding, and entropy-matched guidance that prevents entropy collapse caused by fitting the sharpened teacher distribution in image generation. Omni-Diffusion-Distill achieves state-of-the-art trade-offs between decoding efficiency and generation and understanding performance for multimodal dLLMs, reducing image generation from 128 to 8 decoding steps and multimodal understanding from 512 to 64, giving 18.2x and 21.2x wall-clock speedups. Under these budgets, it scores 0.828 on GenEval and 83.0 on DPG-Bench for text-to-image generation, while reaching GPT judge scores of 20.0 on MM-Vet and 57.2 on COCO captioning (twice the teacher's 28.4 at the same steps) for multimodal understanding.
Figures & tables
Figure 1: Few-step distillation of a unified multimodal dLLM. With up to 16× fewer decoding steps, our student dLLM unifies four tasks: text-to-image generation, image editing, controllable generation and multimodal understanding.
Figure 2: Challenges in few-step distillation of unified multimodal dLLMs. Left: Decoding more text tokens in parallel ignores their dependencies and increases repetition. Right: A student distilled toward the low-entropy CFG target ( orange ) keeps losing entropy and ends below it ( purple ). Entropy-matched guidance keeps our student’s entropy high ( green ).
Figure 3: Overview of Omni-Diffusion-Distill. Stage 1 replays cached teacher trajectories off-policy, and Stage 2 refines the student on its own rollouts. Entropy-matched guidance rescales the image teacher’s entropy, and a pairwise collision penalty reduces text repetition.
Figure 4: Qualitative comparison on T2I. All methods use the same prompts and seeds. Distilled models use 8 decoding steps, and the teacher uses 128.
Figure 5: Performance across decoding budgets. The first two columns show T2I benchmarks, and the last column shows MMU benchmarks. The teacher curve extends to the full-step budget.
Figure 6: Qualitative comparison on image editing. We compare with the teacher and the strongest distilled baseline Di[M]O. Distilled models use 8 decoding steps, and the teacher uses 128.
Figure 7: Qualitative comparison on MMU. We compare with the teacher and the strongest distilled baseline CDLM. Distilled models use 64 decoding steps, and the teacher uses 512. Red marks repeated or incorrect text, and green marks semantically correct content.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Stage 1 (off-policy)
Stage 2 (on-policy)
Optimizer
AdamW ( β1=0.9 , β2=0.95 )
Learning rate (cosine)
2×10−5→2×10−6
1×10−5→1×10−6
Weight decay
0.1
0
Warmup updates
50
25
Global batch size
256
Attention mask (text / image)
block-causal / bidirectional
Appendix
Table 6: Training Configuration. The collision penalty applies only to the text branch. The auxiliary network rψ and the student reference qref are used only in Stage 2 and are discarded after training.
Training
Wall-clock
GPU-hours
Models in memory
Teacher trajectory collection
50.5 h
1,616
pϕ
Stage 1 (off-policy replay)
29.0 h
928
qθ
Stage 2 (on-policy refinement)
31.3 h
1,002
qθ , pϕ , rψ , qref
Total
110.8 h
3,546
Appendix
Table 7: Training Cost. One-time cost on 32 A100 GPUs. Trajectory collection runs once over the training mixture, and every Stage-1 run reuses it.
Method
NFE ↓
POPE ↑
MME-P ↑
MMBench ↑
SEED ↑
MMMU ↑
Teacher (full-step)
512
86.8
1552
91.2
85.3
58.3
32 decoding steps
T3D
32
82.8
1387
60.1
79.0
41.4
Di[M]O
32
86.3
1478
57.6
76.4
36.0
CDLM
32
79.9
1318
78.5
79.3
50.3
Omni-Diffusion-Distill
32
87.6
1535
89.2
84.5
56.2
Appendix
Table 11: Quantitative Results on Discriminative MMU at Smaller Decoding Budgets. The format follows Table 3 .
Figure 8: Additional qualitative comparison on T2I. All methods use the same prompts and seeds. Distilled models use 8 decoding steps, and the teacher uses 128.
Figure 9: Additional qualitative comparison on controllable generation. All methods use the same conditioning inputs and prompts. Distilled models use 8 decoding steps, and the teacher uses 128.
Figure 10: Entropy trajectories under adaptive and fixed-temperature scaling. Entropy-matched guidance targets 0.7 times the conditional entropy along the 64-step teacher trajectory. The two fixed temperatures produce lower entropy than this target early in decoding and higher entropy near the end.
We present LLaDA2.0-Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a natively integrated framework. Its architecture combines a fully semantic discrete tokenizer, a MoE-based dLLM backbone, and a diffusion decoder. By discretizing continuous visual inputs via SigLIP-VQ, the model enables block-level masked diffusion for both text and vision inputs within the backbone, while the decoder reconstructs visual tokens into high-fidelity images. Inference efficiency is enhanced beyond parallel decoding through prefix-aware optimizations in the backbone and few-step distillation in the decoder. Supported by carefully curated large-scale data and a tailored multi-stage training pipeline, LLaDA2.0-Uni matches specialized VLMs in multimodal understanding while delivering strong performance in image generation and editing. Its native support for interleaved generation and reasoning establishes a promising and scalable paradigm for next-generation unified foundation models. Codes and models are available at https://github.com/inclusionAI/LLaDA2.0-Uni.
While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective and efficient alternatives in architectural design. Concurrently, recent studies have successfully applied discrete diffusion models to various domains, such as visual understanding and image generation, revealing their considerable potential as a promising backbone for multimodal systems. Drawing inspiration from these pioneering studies, we introduce Omni-Diffusion, the first any-to-any multimodal language model built entirely on mask-based discrete diffusion models, which unifies understanding and generation across text, speech, and images. Omni-Diffusion employs a unified mask-based discrete diffusion model to directly capture the joint distribution over discrete multimodal tokens. This approach supports not only bimodal tasks but also more complex scenarios involving multiple modalities. On a diverse set of benchmarks, our method outperforms or performs on par with existing multimodal systems that process two or more modalities, highlighting the significant promise of diffusion models in powering the next generation of multimodal foundation models. Project webpage: https://omni-diffusion.github.io.
Diffusion language models (dLLMs) can predict many tokens in parallel, but accurate generation still requires many iterative denoising steps. Few-step distillation accelerates decoding by compressing multiple teacher steps into a single student transition. However, existing methods construct supervision on off-policy trajectories. At inference, the student's early parallel commitments alter the context of later predictions, so the states it actually visits drift away from the supervised ones--precisely when step compression is most aggressive. On-policy distillation is a natural remedy for this mismatch, but it leaves open how far each transition should advance: matching only the teacher's next action limits compression, while indiscriminately merging future actions can violate intermediate dependencies. To address this limitation, we propose OPTD, On-Policy Transition Distillation with consistency-guided adaptive compression. It samples partial states from the few-step student's own trajectories, uses a frozen, question-only teacher to identify outcome-aligned future candidates, and orders them by current-state confidence. The method then selects the longest prefix whose joint commitment preserves the teacher's rollout outcome. A set-bottleneck objective promotes every verified future candidate to the decoder's release threshold, while a frozen-teacher KL anchor regularizes all other active positions. Neither target construction nor training uses a gold response. Across four mathematical reasoning and code-generation benchmarks, OPTD consistently improves the quality--efficiency trade-off and attains the strongest overall quality-constrained AUP among the evaluated few-step baselines.