Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semantic representations without sacrificing generation quality. SelfFlow takes a step in this direction by introducing self-supervised patch alignment into flow matching, but its main gains remain in faster convergence and improved generation. Inspired by DINO and iBOT, we extend this framework with cross-view class-token alignment to further strengthen semantic representations. Specifically, we form two independently noised, dual-timestep observations of each image and align each student class-token representation with the stop-gradient EMA-teacher target from the other observation. This objective is optimized jointly with the inherited flow-matching and local patch objectives. Notably, although the additional objective acts only on the class token, it strengthens both class-token and patch representations. Compared with a matched two-view baseline, ImageNet linear-probing accuracy improves by 9.4% using the class token and 10.1% using mean-pooled patch tokens, while frozen-backbone VOC2012 segmentation improves by 3.6 mIoU. These representation gains are achieved while maintaining comparable ImageNet generation FID. In text-to-image training, the same objective also improves generation FID, reducing it from 2.52 to 2.37 at matched checkpoints. Our results show that representation need not remain a by-product of generation or merely a tool for improving it: it can be directly optimized as a first-class capability of diffusion pretraining alongside generation.
Figures & tables
Figure 1: Class-token alignment across diffusion views. For each clean latent z , two independent noise realizations ϵ(1) and ϵ(2) produce mixed-timestep student inputs and corresponding uniformly cleaner teacher inputs. The student fθ and its exponential-moving-average (EMA) teacher fθ′ each process both views; Sv and Tv label their outputs for view v . Both variants share flow matching to the targets u(v)=z−ϵ(v) and same-view patch alignment between (S1,T1) and (S2,T2) . Our method additionally aligns class tokens across (S1,T2) and (S2,T1) ; the matched baseline sets λcls=0 . Numbered tokens indicate corresponding spatial patches. Teacher feature targets are stop-gradient, and the dashed arrow denotes the EMA update. The output boxes are schematic: alignment features may come from different blocks, and student projection heads are omitted.
Setting
ImageNet
Text-to-image
Backbone
SelfFlowDiT-XL/2, 28 blocks
Flux-style DiT, 21 feature layers
Training data
ImageNet-1K
1M text–image pairs
Resolution
256×256
256×256
Reported training step
1M
400K
Global batch size
256
512
Dual-timestep mask ratio ρ
0.50
0.75
Table 1: Training configurations. Layer numbers are one-based transformer-block indices. The baseline uses the same settings with λcls=0 .
Method
FID ↓
[CLS] linear ↑
Patch linear ↑
VOC mIoU ↑
Two-view SelfFlow
5.12
63.49
60.03
57.15
+ Class-token alignment
5.26
72.90
70.11
60.79
Difference
+0.14
+9.41
+10.08
+3.64
Table 2: ImageNet representation and generation at 1M steps. Patch linear probing and VOC segmentation use features from transformer block 20. Both models contain the same CLS token and differ only in whether its cross-view alignment objective is active.
Figure 2: ImageNet generation and representation quality throughout training. We compare the matched two-view SelfFlow baseline with our model using cross-view class-token alignment. From left to right, we report ImageNet FID, patch-token linear-probe accuracy, and dense-probe mIoU across training checkpoints. Class-token alignment yields persistent improvements in both patch-level and dense representations while maintaining comparable generation quality.
Figure 3: Text-to-image generation throughout training. We compare the matched two-view SelfFlow baseline with our model using cross-view class-token alignment. FID (left) and CLIP score (right) are evaluated across training checkpoints using the same sampling and evaluation pipeline.
Figure 4: Patch representation quality across transformer depth. The vertical markers indicate the class-token alignment block (18) and the EMA patch-target block (20). The largest gains in both ImageNet patch linear probing and VOC2012 dense probing appear at block 20 rather than at the directly supervised CLS block.
Figure 5: CLS-to-patch attention maps on ImageNet. Orange boxes indicate the target object regions in the input images. We compare the matched baseline ( λcls=0 ) with our model using class-token alignment ( λcls=0.2 ). Across these selected examples, the aligned model concentrates attention more strongly on class-relevant and discriminative image regions.
Figure 6: CLS-objective ablations at 100K ImageNet steps. The attachment-depth sweep evaluates weighted k NN at the corresponding attachment block; the loss-weight sweep fixes the attachment at block 18. Top: CLS and patch k NN accuracy. Bottom: FID. Dotted guides and bold ticks mark the selected settings (block 18 and λcls=0.2 ). Exact values are in Tables 7 and 8 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Dataset
ImageNet-1K
Resolution
256×256
Backbone
SelfFlowDiT-XL/2
Transformer blocks
28
Hidden dimension
1152
Attention heads
16
Appendix
Table 3: ImageNet training configuration. Transformer-block indices are one-based. The baseline uses the listed backbone and optimization settings with λcls=0 and no active CLS projection head.
Setting
Value
Training / validation pairs
1M / 50K
Resolution
256×256
Hidden dimension
1152
Attention heads
16
Double-stream / single-stream blocks
7 / 14
Text-conditioning dimension
7680
Appendix
Table 4: Text-to-image training configuration. Diffusion-transformer block indices are one-based. The table lists shared architecture and optimization settings, with the method’s CLS objective; the baseline uses λcls=0 . Here ρ denotes cleaner-token probability.
Model
Learning rate
Epoch
Top-1 (%)
CLS, block 18
Baseline
0.12
54
63.488
Ours
0.10
17
72.898
Mean patch, block 20
Baseline
0.18
58
60.03
Ours
0.12
57
70.11
Appendix
Table 5: Selected settings for final ImageNet linear probes. Each probe uses the same 33-rate search and 60-epoch training budget. Bold and underline mark the higher and lower accuracy within each readout, respectively.
Feature blocks
Baseline
Ours
Δ
ImageNet
20
60.052
70.084
+10.032
{14,16,18,20}
62.602
70.516
+7.914
VOC2012
20
57.180
60.912
+3.732
{14,16,18,20}
60.399
63.104
+2.705
Appendix
Table 6: Single-block and multi-block patch evaluations. ImageNet reports linear-probe top-1 accuracy; VOC reports dense-probe mIoU. Δ is the difference between ours and baseline, in percentage points. Within each dataset, bold and underline mark the highest and second-highest value in each numerical column.
CLS block
FID ↓
CLS k NN ↑
Patch k NN ↑
8
18.978
52.638
25.268
12
21.008
59.696
22.732
18
19.254
63.592
28.366
20
17.135
40.588
26.194
24
18.622
38.996
24.030
28
19.639
66.458
20.076
Appendix
Table 7: CLS attachment-depth ablation at 100K steps. λcls=0.2 and λpatch=0.8 throughout. Representation metrics use features from the corresponding attachment block. k NN accuracies are percentages.
λcls
FID ↓
CLS k NN ↑
Patch k NN ↑
0.000
17.81
42.89
25.42
0.005
18.41
43.56
24.38
0.050
19.16
53.00
25.99
0.100
20.28
58.16
24.83
0.200
19.25
63.59
28.37
0.300
19.70
64.52
29.86
Appendix
Table 8: CLS-loss-weight ablation at block 18 and 100K steps. λpatch=0.8 throughout. k NN accuracies are percentages.
Figure 7: Selected text-to-image comparisons at 400K steps. Within each pair, both models use the same prompt, random seed, and sampling configuration. Examples were selected post hoc according to per-sample CLIP improvement and illustrate individual cases rather than average performance. One pair, with prompt ID 3, is omitted because the baseline output is nearly blank. Displayed prompt labels are abbreviated; full prompts and seeds are listed in Table 9 .
ID
Exact prompt
Seed
1
Draw Lindsey Pelas as Gillian Anderson, the president of the United States, digital painting, ArtStation concept art, sharp-focus illustration art by Artgerm, H 704.
20415
2
I want an image of a magnificent taiko drummer in the dark woods.
17146
4
Make a picture of the goddess of avocados, by Donato Giancola.
49143
5
Produce an image depicting every era of Michael Jackson represented as all the members of the Jackson 5.
40777
6
Please visualize Sia Furler full body.
42874
7
I want an image of a festival poster of Jim Morrison in the Astral Plane, haunting digital art.
49751
Appendix
Table 9: Full prompts and random seeds for Figure 7 . Original prompt IDs are retained.
Representation alignment with pretrained vision models has recently shown strong potential for accelerating diffusion transformer training. By aligning intermediate diffusion features with clean-image representations from self-supervised vision encoders, existing methods improve convergence and generation quality. However, such alignment also introduces a non-trivial constraint: diffusion models operate on noisy inputs whose usable information varies across timesteps, while the reference features are extracted from clean images. In this paper, we revisit this mismatch from a token-level perspective. We find that, under full-token representation alignment, tokens with large alignment-gradient norms exhibit a stable spatial preference, suggesting that the alignment objective does not affect all tokens uniformly and may encourage the model to rely on the complete set of clean-image tokens. To address this issue, we propose MaskAlign, a token-subset representation alignment method that applies alignment to randomly sampled token subsets during training. By exposing the model to different token subsets across iterations, MaskAlign reduces the dependence of representation alignment on the complete token set and encourages alignment behavior that is more stable under token-subset perturbations. To mitigate the information loss caused by directly dropping tokens, we further introduce a lightweight pre-mask token mixing block that shares information across tokens before masking.
Lianyu Pang, Tianlin Pan, Cheng Da +5
The Hong Kong University of Science and Technology · Kuaishou Technology · University of Chinese Academy of Sciences
Representation alignment has become an effective way to accelerate diffusion transformer training and improve generation quality. Recent self-alignment methods, such as SRA and Self-Flow, further remove the dependency on external pretrained encoders by constructing alignment within the diffusion model itself. However, the mechanism behind the improvement from SRA to Self-Flow, dual-time scheduling, remains under-examined: Self-Flow attributes its gain to interactions between tokens at different noise levels, where cleaner tokens help infer noisier ones. In this work, we revisit this explanation and ask whether the gain instead comes from data augmentation along the noise dimension. To disentangle these factors, we introduce Attention Separation, which preserves the same dual-timestep input as Self-Flow while blocking attention between tokens assigned to different noise levels. Surprisingly, removing such interaction does not degrade performance and can even improve it, suggesting that the improvement from SRA to Self-Flow mainly comes from data augmentation. Furthermore,We show that Attention Separation itself provides an augmentation effect by splitting a single image into multiple effective training parts to expand the training data. Based on these observations, we combine self-representation alignment with dual-timestep and attention-separation augmentation, and demonstrate the effectiveness of this design on ImageNet.
Dengyang Jiang, Mengmeng Wang, Harry Yang +1
1The Hong Kong University of Science and Technology · 2Zhejiang University of Technology · 3Baidu Inc.
Most diffusion and flow-matching generators define the prior, probability path, and prediction target in the same representation space. Latent diffusion improves efficiency by moving this path into an autoencoder latent space, but the final sample is still produced by a separately trained decoder. This separation creates a mismatch: the generator is optimized for latent-space prediction, while final quality depends on how the decoder handles generated latents that may differ from clean encoder outputs. We introduce CrossFlow, a cross-space flow formulation that maps noisy latent inputs directly to pixel-space images. The key technical step is a velocity-free one-step objective: the latent trajectory defines the training path, but the supervised prediction is an image rather than a latent displacement. This lets one model act both as a one-step latent-to-pixel generator and as a decoder replacement for latent diffusion pipelines. On class-conditional ImageNet-1k at 256×256, CrossFlow-XL achieves 1.62 FID with one function evaluation. Ablations show that the latent encoder and pixel-space perceptual and adversarial losses are important for fidelity. These results indicate that cross-space flow objectives can combine the efficiency of latent representations with direct pixel-space supervision, without requiring a separate decoder at inference.
Xiyuan Wang, Xiao Zhang, Yang Li +4
Institute for Artificial Intelligence, Peking University · Tencent · Fudan University