Organizations: Indian Institute of Technology Kanpur, India · Rajiv Gandhi Institute of Petroleum Technology, India · IIIT Delhi, India · Tezpur University, India · The University of Queensland, Australia
Recent work on language-model adaptation has shown that single models can obtain informative training signals by evaluating their behavior in demonstrationor feedback-augmented contexts, with the help of a teacher network, which is driven by the student's learned parameters. Inspired by this internal-reference principle, we investigate how diffusion models can identify self-referenced training signals without external demonstrations or teacher networks. We introduce SALD, a self-referenced training framework that evaluates each image-caption pair at two noise levels using the same model. The easier, lower-noise path is evaluated without gradient tracking to provide a reference, while the harder, higher-noise path provides the training gradient. Rather than directly distilling the easy-path prediction, SALD uses the difference between two path errors to adapt the hardpath objective. The proposed Advantage-Guided Diffusion (AGD) converts this relative error into a differentiable sample-level weight. Temporal Advantage Memory (TAM) accumulates relative difficulty across training and adapts the future gap between the two noise levels. Spectral Advantage Decomposition (SAD) further compares the residual power spectra of the two paths and constructs a differentiable, frequency-derived latent-element weight. All components share a single set of model parameters, requiring neither an external teacher network nor additional trainable parameters during training or inference, and no modification to the inference procedure. Experiments across multiple architectures and datasets demonstrate consistent improvements in generation quality, while component-wise ablations quantify the contributions of the proposed components.
Figures & tables
Figure 1: Images generated by SALD using the SANA 0.6B model at a resolution of 1024×1024 .
Figure 2: SALD framework. A frozen VAE encodes the input latent, processed through two paths: Low-Noise Path A (w/o gradient) and High-Noise Path B (with gradient). AGD computes an advantage for sample reweighting, while TAM stabilizes it via per-sample EMA, forming a persistent curriculum across timestep gaps. SAD compares FFT power spectra of velocity residual maps, applies a sigmoid gate, and reconstructs a spatial weight map via IFFT to focus gradients on high-frequency error regions. Together, these define LSALD for adaptive gradient allocation across samples.
Model
Method
COCO (2017)
MultiGen-20M
Flickr8k
CUB-200
Oxford 102 Flowers
FID ↓
CLIP ↑
LPIPS ↓
FID ↓
CLIP ↑
LPIPS ↓
FID ↓
CLIP ↑
LPIPS ↓
FID ↓
CLIP ↑
LPIPS ↓
FID ↓
CLIP ↑
LPIPS ↓
SANA (1.6B)
Base
26.49
26.44
0.7411
26.44
26.25
0.7541
27.22
28.49
0.7336
21.63
26.01
0.7748
33.11
26.35
0.7896
SpeeD
25.38
26.83
0.7409
26.11
26.26
0.7558
26.87
28.52
0.7324
21.53
26.28
0.7753
32.04
26.51
0.7782
Temporal Diff
26.36
26.87
0.7405
26.31
26.29
0.7539
26.85
28.85
0.7325
20.72
26.25
0.7766
33.10
26.54
0.7801
SRA
26.89
26.32
0.7461
26.05
25.89
0.7554
26.87
28.78
0.7341
20.78
26.18
0.7722
32.75
26.40
0.7781
SALD
22.76
26.88
0.7391
23.86
26.29
0.7533
26.75
28.88
0.7316
18.16
26.29
0.7537
27.37
26.57
0.7781
Table 1: Quantitative results on ( 512×512 ) resolution datasets. Best values per metric are in Blue .
Model
Method
MAGICK
DALL ⋅ E 3
FID ↓
CLIP ↑
LPIPS ↓
FID ↓
CLIP ↑
LPIPS ↓
SANA 0.6B
Base
38.59
26.68
0.7470
14.68
29.96
0.7882
SpeeD
38.27
26.65
0.7303
14.18
30.05
0.7878
Temporal Diff
37.95
26.69
0.7294
13.84
30.38
0.7854
SRA
38.15
26.70
0.7297
13.96
30.73
0.7839
SALD
36.74
26.85
0.7230
11.98
31.06
0.7826
Table 2: Quantitative results on 1024×1024 resolution images. Best values per metric within each model are in Blue .
Figure 5
Figure 5: Visual comparison of generated samples. Red boxes indicate missing details or artifacts in latest baseline models. Green boxes highlight that our method generates better high-quality details. We use the same seed for all comparisons to ensure fairness. Captions are in Blue Color .
Cases
Setting
COCO ( 512 )
MAGICK ( 1024 )
FID ↓
CLIP ↑
LPIPS ↓
FID ↓
CLIP ↑
LPIPS ↓
Case 1
AGD only
30.95
26.30
0.7462
37.75
26.73
0.7354
SAD only
31.02
26.49
0.7427
37.86
26.70
0.7314
TAM only
30.98
26.50
0.7474
37.35
26.74
0.7323
Case 2
AGD + TAM
30.78
26.60
0.7529
37.34
26.72
0.7396
AGD + SAD
30.56
26.66
0.7375
37.78
26.79
0.7400
Table 3: Ablation study of SALD on SANA 0.6B across COCO 2017 ( 512 ) and MAGICK ( 1024 ).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Params (B)
FLOPs (G)
Resolution
LR
Guidance
Infer. Steps
SANA (0.6B)
0.6
548
512×512
1e-4
5
20
SANA (1.6B)
1.6
567
512×512
1e-4
5
20
PixArt- Σ
0.6
567
512×512
2e-5
5
20
SANA (0.6B) †
0.6
1397
1024×1024
1e-4
5
20
SANA (1.6B) †
1.6
2240
1024×1024
1e-4
5
20
PixArt- Σ†
0.6
1399
1024×1024
2e-5
5
20
Appendix
Table 4: Detailed model configurations used in our experiments, including parameter count, per-forward-pass FLOPs, training resolution, learning rate, guidance scale, and inference steps. † denotes models additionally evaluated at 1024×1024 resolution.
Figure 6: Qualitative comparison across baseline models and SALD on 1024 × 1024 datasets.
Figure 7: Qualitative comparisons: samples generated by SALD and other different Baselines on 1024 × 1024 Datasets. Captions are, Column-1: A cute orange cat sleeping curled up, flat vector illustration style. Column-2: An hourglass with golden sand, dramatic cinematic lighting on a wooden table. Column-3: A woman’s face in profile with flowing blue hair made of water, digital art.
Figure 8: Qualitative comparison of inference-time generation across different sampling steps for SALD and the baseline using SANA 1.6B trained on DALL ⋅ E 3 1M.
Dataset
Model
Variant
Params (B)
Steps
FID ↓
CLIP ↑
LPIPS ↓
CUB-200
SANA
Base
1.6
12.5k
21.47
26.01
0.7709
SpeeD
1.6
12.5k
20.83
26.08
0.7670
SRA
1.6
12.5k
20.72
26.15
0.7632
Temporal Diff
1.6
12.5k
20.61
26.22
0.7593
SALD
1.6
10k
18.16
26.29
0.7537
SANA
Base
0.6
12.5k
25.54
26.19
0.7802
Appendix
Table 5: Wall-clock time and quantitative comparison on CUB-200 and MultiGen datasets. Best values in Blue
Model
Params (B)
Variant
Steps = 10
Steps = 15
Steps = 20
FID ↓
CLIP ↑
LPIPS ↓
FID ↓
CLIP ↑
LPIPS ↓
FID ↓
CLIP ↑
LPIPS ↓
SANA
1.6
Base
55.06
25.35
0.8153
37.49
25.60
0.8055
33.11
26.35
0.7896
SALD
41.56
25.48
0.8088
30.60
25.72
0.7962
27.37
26.57
0.7781
SANA
0.6
Base
44.05
24.31
0.8059
35.26
25.29
0.7971
30.64
26.43
0.7805
SALD
33.68
24.48
0.7906
25.54
25.63
0.7865
25.83
26.65
0.7728
PixArt- Σ
0.6
Base
49.12
25.10
0.8120
43.25
25.25
0.7910
37.88
25.37
0.7687
Appendix
Table 6: Quantitative comparison across different inference steps on the Oxford Flowers dataset. Light blue highlights indicate the metric values for the SALD variant.
Steps
Variant
MultiGen
Oxford 102 Flowers
SANA 1.6
SANA 0.6
SANA 1.6
SANA 0.6
FID
CLIP
LPIPS
FID
CLIP
LPIPS
FID
CLIP
LPIPS
FID
CLIP
LPIPS
10k
Baseline
26.44
26.25
0.7541
31.93
26.68
0.7495
33.11
26.35
0.7896
30.64
26.43
0.7805
SALD
23.86
26.29
0.7533
29.15
26.69
0.7472
27.37
26.57
0.7781
25.83
26.65
0.7728
20k
Baseline
26.18
26.31
0.7550
31.42
26.48
0.7605
32.45
26.30
0.7842
30.41
26.38
0.7761
SALD
23.52
26.35
0.7542
28.78
26.52
0.7598
26.80
26.52
0.7745
25.70
26.60
0.7695
Appendix
Table 7: Evaluation metrics across training steps for SANA 1.6 and SANA 0.6 on MultiGen and Oxford 102 Flowers datasets. Lower FID and LPIPS ( ↓ ) are better, while higher CLIP ( ↑ ) is better.
Figure 9: Qualitative samples generated by SALD using PixArt- Σ at 1024×1024 resolution on the DALL ⋅ E 3 1M dataset.
Figure 10: Qualitative samples generated by SALD using Flux.2 klein Base at 1024×1024 resolution on the DALL ⋅ E 3 1M dataset.
Figure 11: Qualitative samples generated by SALD across multiple datasets and architectures. Row 1: SANA 1.6B at 1024×1024 resolution trained on DALL ⋅ E 3 1M. Row 2: SANA 1.6B generations on CUB-200. Row 3: PixArt- Σ generations on Oxford-102 Flowers.