Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work, we introduce a framework for reading this contextual space through natural-language interrogation. We train a lightweight bottleneck network that maps intermediate contextual tokens into the input space of a frozen Large Language Model (LLM), allowing the LLM to answer questions about the emerging image directly from these hidden representations. Our reader reveals that contextual tokens encode a rich, global representation of the emerging scene: generation-specific semantics, including attributes left underspecified by the prompt, are accessible surprisingly early in denoising, while increasingly fine-grained details become readable over time. Remarkably, this information remains decodable even when the MM-DiT receives an empty prompt, showing that contextual tokens accumulate substantial image-specific information from the evolving visual representation itself. We further find that generations with more readable contextual representations tend to receive higher human-preference scores. Building on these observations, we introduce Contextual Alignment, a training technique that explicitly reinforces the visual-semantic information encoded in the contextual tokens, improving generation quality and distributional coverage. Together, our results establish contextual tokens as both an interpretable view into the internal dynamics of MM-DiTs and an effective target for improving generative models.
Figures & tables
Figure 2: Training the Contextual Reader. At a selected denoising timestep, contextual tokens from multiple frozen MM-DiT layers are aggregated by a trainable Bottleneck Network into a compact Contextual Descriptor. A frozen LLM receives this descriptor together with a visual question and predicts an answer about the emerging image. We train only the Bottleneck Network using cross-entropy against the ground-truth answer.
Figure 3: Exemplary plain-text predictions from our conditional Contextual Reader on FLUX.2. For each prompt, we compare two random seeds that resolve underspecified attributes differently. We show the predicted clean images at 8% of the denoising trajectory for both seeds, followed by their final images, and apply the reader to the contextual tokens at the same 8% timestep to obtain the corresponding answers. Despite the visual predictions remaining ambiguous at the 8% timestep, the reader can recover seed-specific information that is not specified by the prompt.
Figure 4: Exemplary plain-text predictions from our unconditional Contextual Reader on FLUX.2. We show predicted clean images at 0% , 8% , and 20% of the denoising trajectory, together with the final generation. Colored highlights associate each intermediate timestep with the corresponding reader answer obtained from the contextual tokens at that timestep, shown below. Despite receiving no text prompt, the reader recovers increasingly specific semantic information about the eventual image as denoising progresses.
Figure 5: Contextual readability emerges early and reflects final generation quality. Left: VLM-as-a-judge evaluation across the denoising trajectory shows that contextual scores increase as generation progresses; with the input prompt they are already high from the start. With an empty prompt they rise rapidly as image-specific information enters the contextual space. Middle: Ranking seeds by reader scores separates generations with different HPS scores; Δ HPS denotes the HPS difference between the top- and bottom-ranked seeds. Right: Representative top- and bottom-ranked generations illustrate that higher contextual readability is associated with stronger final generations.
Figure 6: Contextual Alignment. A lightweight Aligner Network maps contextual tokens from the MM-DiT into a learned embedding, optimized via a negative cosine similarity loss to match the teacher embedding produced by a frozen semantic encoder from the clean input image.
Figure 7: Qualitative comparison on SD3-Medium fine-tuned on Fine-T2I. The “Reference” column shows the held-out images from Fine-T2I.
Table 1: Quantitative comparisons with representation-alignment baselines.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Visual questions used by the Contextual Reader. We manually curate 15 questions spanning diverse semantic aspects of human-centric images.
Figure 9: Reader training and validation loss. Training and validation loss over 5,000 optimization steps for the conditional and unconditional readers across both diffusion backbones. The validation set consists of 5% of the training data and is separate from the held-out data used for reader evaluation.
Figure 10: VLM-as-a-judge prompt and scoring rubric. Prometheus-Vision receives the visual question, the reader’s prediction, the reference answer, and the generated image, and assigns a score from 1 to 5 according to the rubric.
Figure 11: Comparing semantic readability with the predicted image. We compare the Contextual Reader with a Latent Reader operating on the output latent representation and an x0 +VLM baseline that answers questions directly from the predicted clean image. Results are shown for both conditional and unconditional readers across two diffusion backbones. The contextual representations provide higher readability during the early stages of denoising, while x0 +VLM achieves the highest readability from approximately 40% onward.
Figure 12: Comparing semantic readability with the prompt. We compare the Contextual Reader with a Prompt Reader using the same architecture but receiving only the input prompt. Results are shown for both conditional and unconditional readers across two diffusion backbones. The conditional Contextual Reader achieves higher readability than the Prompt Reader throughout denoising, while the unconditional Contextual Reader surpasses it for FLUX.2 and approaches it for SD3.5 as denoising progresses.
Figure 13: Generalization to unseen questions. For each question, we train a separate reader on the remaining 14 questions and evaluate it on the held-out question. We report the average semantic readability across all 15 held-out evaluations. Despite the reduced supervision, the reader maintains meaningful readability across denoising timesteps for both conditioning settings and diffusion backbones.
Figure 15: Additional plain-text predictions from our conditional Contextual Reader on FLUX.2. For each prompt, we compare two random seeds that resolve underspecified attributes differently. We show the predicted clean images at 8% of the denoising trajectory for both seeds, followed by their final images. The reader predictions are obtained from the contextual tokens at the same 8% timestep.
Figure 16: Exemplary plain-text predictions from our conditional Contextual Reader on SD3.5. For each prompt, we compare two random seeds that resolve underspecified attributes differently. We show the predicted clean images at 18% of the denoising trajectory for both seeds, followed by their final images. The reader predictions are obtained from the contextual tokens at the same 18% timestep.
Figure 17: Additional plain-text predictions from the unconditional Contextual Reader on FLUX.2. We show predicted clean images at 0% , 8% , and 20% of the denoising trajectory, together with the final generation. Colored highlights associate each intermediate timestep with the corresponding reader answer below. Despite receiving no text prompt, the reader recovers increasingly specific semantic information about the eventual image as denoising progresses.
Figure 18: Exemplary plain-text predictions from the unconditional Contextual Reader on SD3.5. We show predicted clean images at 0% , 7% , and 18% of the denoising trajectory, together with the final generation. Colored highlights associate each intermediate timestep with the corresponding reader answer below. Despite receiving no text prompt, the reader recovers increasingly specific semantic information about the eventual image as denoising progresses.
Full model training
Fine-tuning
Variant
FID ↓
CLIP ×102↑
HPS ×102↑
Prec. ↑
Rec. ↑
FID ↓
Prec. ↑
Rec. ↑
Semantic Teacher: CLIP
4.71
24.16
21.30
0.599
0.375
16.87
0.994
0.970
Semantic Teacher: DINOv2
4.67
23.96
21.03
0.604
0.392
16.92
0.993
0.964
Semantic Teacher: Qwen3-VL-2B
5.03
24.09
21.23
0.588
0.352
16.81
0.991
0.969
Contextual Layer: L3
4.54
24.20
21.30
0.607
0.388
17.04
0.992
0.968
Contextual Layer: L5
4.78
24.18
21.23
0.597
0.364
16.97
0.994
0.970
Appendix
Table 2: Ablation study conducted in both the full model training setting ( 150k steps from random initialization on MS-COCO) and the fine-tuning setting ( 30k steps on the Fine-T2I curated subset, starting from SD3). In full model training, Ours refers to REPA + CoAl; in fine-tuning, it refers to vanilla flow matching + CoAl. Bold/underline mark the best/second-best value per column within each setting.
Method
FID ↓
KID ×103↓
CLIP ×102↑
VQA ↑
HPS ×102↑
Prec. ↑
Rec. ↑
Vanilla
6.52
1.53
23.27
0.630
20.10
0.496
0.206
REG
7.33
2.10
23.79
0.613
20.81
0.486
0.268
SRA
5.89
1.23
23.49
0.637
20.37
0.515
0.261
HASTE
4.87
0.93
24.21
0.691
21.42
0.609
0.361
REPA
4.95
0.94
24.20
0.691
21.38
0.591
0.338
Vanilla + Ours
6.23 (-4.4%)
1.52 (-0.9%)
23.30 (+0.1%)
0.629 (-0.2%)
20.17 (+0.3%)
0.515 (+3.7%)
0.274 (+33.2%)
Appendix
Table 3: Extended quantitative comparison against baselines in the full model training setting ( 150k steps from random initialization on MS-COCO 2014). Baseline rows: bold / underline mark best/second-best among the baselines. CoAl rows: value colored green / red if it improves/regresses on that metric relative to its own non-CoAl counterpart, with the percent change in parentheses.
Method
FID ↓
KID ×103↓
CLIP ×102↑
VQA ↑
HPS ×102↑
Prec. ↑
Rec. ↑
Reference
–
–
26.84
0.925
25.53
–
–
Base Model
19.30
3.05
28.70
0.940
29.04
0.993
0.939
Vanilla
17.18
2.23
28.87
0.941
27.78
0.993
0.966
REG
21.13
3.68
28.43
0.923
24.18
0.976
0.961
SRA
17.48
2.45
28.87
0.940
27.71
0.991
0.964
HASTE
17.53
2.34
28.80
0.938
27.71
0.994
0.963
Appendix
Table 4: Extended quantitative comparison against baselines on the Fine-T2I curated benchmark (val5k, 30k fine-tuning steps starting from SD3-Medium). Reference refers to the real Fine-T2I validation set images, and Base Model is the pretrained SD3-Medium checkpoint before fine-tuning. Bold / underline mark the best/second-best value per column across all rows. CoAl rows are additionally colored green / red if they improve/regress on that metric relative to their own non-CoAl counterpart, with the percent change in parentheses.
Figure 19: Effect of classifier-free guidance (CFG) scale on FID and KID on the Fine-T2I curated benchmark (val5k), comparing Vanilla, REPA, HASTE, and CoAl.
Figure 20: Additional qualitative comparison against visual alignment methods over SD3-Medium fine-tuning on the Fine-T2I curated dataset.
Figure 21: Additional qualitative comparison against visual alignment methods over SD3-Medium fine-tuning on the Fine-T2I curated dataset.
Figure 22: Qualitative comparisons on Fine-T2I val5k prompts: each pair shows a baseline (left) against the same baseline with our method added (right). Each pair uses its own set of example prompts.
Modern text-to-image diffusion transformers (DiTs) generate images through joint attention, in which text and image tokens interact directly within a single sequence. In large-scale DiTs, the conditioning input contains not only the user prompt but also chat-template tokens introduced by LLM-based text encoders. Yet how these tokens participate in the denoising computation remains poorly understood. To probe this, we introduce a causal interpretability framework. Using it to separate prompt-content tokens from chat-template tokens, we find that the template tokens carry little prompt-specific information at the encoder output. Yet surprisingly, they emerge as dominant image-to-text attention sinks and causally maintain object identity inside the DiT, acting as implicit semantic registers. We show that they acquire this identity indirectly. Rather than reading the prompt tokens, they draw the identity from the image latents into which the prompt semantics have already been injected at the very first layer. We further reveal a division of labor across heads and depth in DiTs, where distinct heads route semantics or render visual structure, and identity is committed in early blocks, carried by middle blocks, and refined in late ones. As a practical payoff, this analysis yields a training-free pruning rule that removes the causally inert prompt-reading heads and cuts 20% of joint-attention FLOPs at a 1.4-point cost in GenEval accuracy. Overall, our work not only reveals that the tokens encoding semantics at the input need not be those that maintain them during generation, but also provides a causal view of internal mechanisms in diffusion transformers.
Maohua Li, Qirui Li, Yanke Zhou +10
Nanjing University · Alibaba Group · Zhejiang University
Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architecture for both generation (visual synthesis) and understanding (text generation) is to apply this framework to language modeling. We propose TextLDM, which transfers the visual latent diffusion recipe to text generation with minimal architectural modification. A Transformer-based VAE maps discrete tokens to continuous latents, enhanced by Representation Alignment (REPA) with a frozen pretrained language model to produce representations effective for conditional denoising. A standard DiT then performs flow matching in this latent space, identical in architecture to its visual counterpart. The central challenge we address is obtaining high-quality continuous text representations: we find that reconstruction fidelity alone is insufficient, and that aligning latent features with a pretrained language model via REPA is critical for downstream generation quality. Trained from scratch on OpenWebText2, TextLDM substantially outperforms prior diffusion language models and matches GPT-2 under the same settings. Our results establish that the visual DiT recipe transfers effectively to language, taking a concrete step toward unified diffusion architectures for multimodal generation and understanding.
Diffusion transformer (DiT) has been widely adopted in the generative diffusion field, advancing the denoising of query tokens through attention and Feed-Forward (\text{FFN}) layers. FFN actually acts as the key-value vocabulary for decoding visual contents where the value embeds the visual semantical knowledge. We present that focusing on critical query tokens corresponding to more complex details and encouraging the model to improve these tokens is essential for fine-grained visual generation. To this end, we propose FocusDiT, which applies a Masking scheme to focus on critical query tokens that are exclusively fed into FFN. The masked queries can retrieve visual tokens from the FFN vocabularies, and use them to decode their visual details. Extensive text-to-image experiments validate the effectiveness of token masking in enhancing generative performance.
Xueji Fang, Liyuan Ma, Jianhao Zeng +3
Zhejiang University · Westlake University · Hangzhou, China