Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models' own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semantic representations without sacrificing generation quality. SelfFlow takes a step in this direction by introducing self-supervised patch alignment into flow matching, but its main gains remain in faster convergence and improved generation. Inspired by DINO and iBOT, we extend this framework with cross-view class-token alignment to further strengthen semantic representations. Specifically, we form two independently noised, dual-timestep observations of each image and align each student class-token representation with the stop-gradient EMA-teacher target from the other observation. This objective is optimized jointly with the inherited flow-matching and local patch objectives. Notably, although the additional objective acts only on the class token, it strengthens both class-token and patch representations. Compared with a matched two-view baseline, ImageNet linear-probing accuracy improves by 9.4% using the class token and 10.1% using mean-pooled patch tokens, while frozen-backbone VOC2012 segmentation improves by 3.6 mIoU. These representation gains are achieved while maintaining comparable ImageNet generation FID. In text-to-image training, the same objective also improves generation FID, reducing it from 2.52 to 2.37 at matched checkpoints. Our results show that representation need not remain a by-product of generation or merely a tool for improving it: it can be directly optimized as a first-class capability of diffusion pretraining alongside generation.
Figures & tables
Figure 1: Class-token alignment across diffusion views. For each clean latent z , two independent noise realizations ϵ(1) and ϵ(2) produce mixed-timestep student inputs and corresponding uniformly cleaner teacher inputs. The student fθ and its exponential-moving-average (EMA) teacher fθ′ each process both views; Sv and Tv label their outputs for view v . Both variants share flow matching to the targets u(v)=z−ϵ(v) and same-view patch alignment between (S1,T1) and (S2,T2) . Our method additionally aligns class tokens across (S1,T2) and (S2,T1) ; the matched baseline sets λcls=0 . Numbered tokens indicate corresponding spatial patches. Teacher feature targets are stop-gradient, and the dashed arrow denotes the EMA update. The output boxes are schematic: alignment features may come from different blocks, and student projection heads are omitted.
Setting
ImageNet
Text-to-image
Backbone
SelfFlowDiT-XL/2, 28 blocks
Flux-style DiT, 21 feature layers
Training data
ImageNet-1K
1M text–image pairs
Resolution
256×256
256×256
Reported training step
1M
400K
Global batch size
256
512
Dual-timestep mask ratio ρ
0.50
0.75
Table 1: Training configurations. Layer numbers are one-based transformer-block indices. The baseline uses the same settings with λcls=0 .
Method
FID ↓
[CLS] linear ↑
Patch linear ↑
VOC mIoU ↑
Two-view SelfFlow
5.12
63.49
60.03
57.15
+ Class-token alignment
5.26
72.90
70.11
60.79
Difference
+0.14
+9.41
+10.08
+3.64
Table 2: ImageNet representation and generation at 1M steps. Patch linear probing and VOC segmentation use features from transformer block 20. Both models contain the same CLS token and differ only in whether its cross-view alignment objective is active.
Figure 2: ImageNet generation and representation quality throughout training. We compare the matched two-view SelfFlow baseline with our model using cross-view class-token alignment. From left to right, we report ImageNet FID, patch-token linear-probe accuracy, and dense-probe mIoU across training checkpoints. Class-token alignment yields persistent improvements in both patch-level and dense representations while maintaining comparable generation quality.
Figure 3: Text-to-image generation throughout training. We compare the matched two-view SelfFlow baseline with our model using cross-view class-token alignment. FID (left) and CLIP score (right) are evaluated across training checkpoints using the same sampling and evaluation pipeline.
Figure 4: Patch representation quality across transformer depth. The vertical markers indicate the class-token alignment block (18) and the EMA patch-target block (20). The largest gains in both ImageNet patch linear probing and VOC2012 dense probing appear at block 20 rather than at the directly supervised CLS block.
Figure 5: CLS-to-patch attention maps on ImageNet. Orange boxes indicate the target object regions in the input images. We compare the matched baseline ( λcls=0 ) with our model using class-token alignment ( λcls=0.2 ). Across these selected examples, the aligned model concentrates attention more strongly on class-relevant and discriminative image regions.
Figure 6: CLS-objective ablations at 100K ImageNet steps. The attachment-depth sweep evaluates weighted k NN at the corresponding attachment block; the loss-weight sweep fixes the attachment at block 18. Top: CLS and patch k NN accuracy. Bottom: FID. Dotted guides and bold ticks mark the selected settings (block 18 and λcls=0.2 ). Exact values are in Tables 7 and 8 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Dataset
ImageNet-1K
Resolution
256×256
Backbone
SelfFlowDiT-XL/2
Transformer blocks
28
Hidden dimension
1152
Attention heads
16
Appendix
Table 3: ImageNet training configuration. Transformer-block indices are one-based. The baseline uses the listed backbone and optimization settings with λcls=0 and no active CLS projection head.
Setting
Value
Training / validation pairs
1M / 50K
Resolution
256×256
Hidden dimension
1152
Attention heads
16
Double-stream / single-stream blocks
7 / 14
Text-conditioning dimension
7680
Appendix
Table 4: Text-to-image training configuration. Diffusion-transformer block indices are one-based. The table lists shared architecture and optimization settings, with the method’s CLS objective; the baseline uses λcls=0 . Here ρ denotes cleaner-token probability.
Model
Learning rate
Epoch
Top-1 (%)
CLS, block 18
Baseline
0.12
54
63.488
Ours
0.10
17
72.898
Mean patch, block 20
Baseline
0.18
58
60.03
Ours
0.12
57
70.11
Appendix
Table 5: Selected settings for final ImageNet linear probes. Each probe uses the same 33-rate search and 60-epoch training budget. Bold and underline mark the higher and lower accuracy within each readout, respectively.
Feature blocks
Baseline
Ours
Δ
ImageNet
20
60.052
70.084
+10.032
{14,16,18,20}
62.602
70.516
+7.914
VOC2012
20
57.180
60.912
+3.732
{14,16,18,20}
60.399
63.104
+2.705
Appendix
Table 6: Single-block and multi-block patch evaluations. ImageNet reports linear-probe top-1 accuracy; VOC reports dense-probe mIoU. Δ is the difference between ours and baseline, in percentage points. Within each dataset, bold and underline mark the highest and second-highest value in each numerical column.
CLS block
FID ↓
CLS k NN ↑
Patch k NN ↑
8
18.978
52.638
25.268
12
21.008
59.696
22.732
18
19.254
63.592
28.366
20
17.135
40.588
26.194
24
18.622
38.996
24.030
28
19.639
66.458
20.076
Appendix
Table 7: CLS attachment-depth ablation at 100K steps. λcls=0.2 and λpatch=0.8 throughout. Representation metrics use features from the corresponding attachment block. k NN accuracies are percentages.
λcls
FID ↓
CLS k NN ↑
Patch k NN ↑
0.000
17.81
42.89
25.42
0.005
18.41
43.56
24.38
0.050
19.16
53.00
25.99
0.100
20.28
58.16
24.83
0.200
19.25
63.59
28.37
0.300
19.70
64.52
29.86
Appendix
Table 8: CLS-loss-weight ablation at block 18 and 100K steps. λpatch=0.8 throughout. k NN accuracies are percentages.
Figure 7: Selected text-to-image comparisons at 400K steps. Within each pair, both models use the same prompt, random seed, and sampling configuration. Examples were selected post hoc according to per-sample CLIP improvement and illustrate individual cases rather than average performance. One pair, with prompt ID 3, is omitted because the baseline output is nearly blank. Displayed prompt labels are abbreviated; full prompts and seeds are listed in Table 9 .
ID
Exact prompt
Seed
1
Draw Lindsey Pelas as Gillian Anderson, the president of the United States, digital painting, ArtStation concept art, sharp-focus illustration art by Artgerm, H 704.
20415
2
I want an image of a magnificent taiko drummer in the dark woods.
17146
4
Make a picture of the goddess of avocados, by Donato Giancola.
49143
5
Produce an image depicting every era of Michael Jackson represented as all the members of the Jackson 5.
40777
6
Please visualize Sia Furler full body.
42874
7
I want an image of a festival poster of Jim Morrison in the Astral Plane, haunting digital art.
49751
Appendix
Table 9: Full prompts and random seeds for Figure 7 . Original prompt IDs are retained.