Recent advances in diffusion models highlight the importance of representation regularization for improving sample quality and training efficiency. However, commonly used regularization methods often overlook the built-in conditions (such as labels or texts) which directly determine the generation target. In this work, we demonstrate how conditioning signals affect the feature distribution and introduce the CARE (Condition-Aware REpresentation regularization). CARE is a lightweight plug-and-play regularization framework that dynamically modulates feature distribution based on condition similarity. CARE leverages built-in conditioning signals to judiciously guide the representation space, promoting tighter feature clusters for similar conditions without relying on explicit alignment losses or external supervision. Empirically, CARE consistently improves both visual fidelity and convergence stability across both class-to-image and text-to-image tasks. On ImageNet, CARE achieves a 19.08% reduction in FID in 400k training steps, leading to a 3.5× speed-up. When applied to text-to-image generation, CARE lowers FID by 16.61% in 200k iterations and improves semantic alignment between generated samples and text prompts. Moreover, CARE can be seamlessly integrated with existing regularization methods, yielding additional performance gains.
Figures & tables
Figure 1 : Overview of CARE. Left: DiT pipeline. Right: Representation and condition spaces with CARE regularization. CARE penalizes representation collapse for all sample pairs, with a similarity-dependent strength based on condition similarity.
Figure 2 : FID-50K on ImageNet 256×256 for SiT-XL/2 with and without CARE. All models are evaluated with SDE sampling, without CFG, using 250 sampling steps. CARE consistently improves FID throughout training.
Figure 3 : FID-50K versus representation alignment with text conditions, measured by a linear probe on intermediate representations. Across both settings with and without REPA, CARE consistently improves representation–condition alignment while reducing FID.
Method
Iter.
Sampler
NFEs
CFG scale
FID ↓
IS ↑
CLIP-T ↑
class-to-image generation
SiT-B/2
400k
SDE
250
1.0
33.02
43.71
-
SiT-B/2 + dispersive loss
400k
SDE
250
1.0
31.37
47.85
-
SiT-B/2 + CARE
400k
SDE
250
1.0
29.71
50.10
-
SiT-B/2 + REPA
400k
ODE
250
1.0
24.31
62.23
-
SiT-B/2 + REPA + dispersive loss
400k
ODE
250
1.0
23.36
64.79
-
Table 1 : CARE consistently improves diffusion models across class-to-image and text-to-image benchmarks.
Figure 4 : Qualitative comparison on ImageNet 256 × 256. Both models share the same noise, and use a Euler ODE sampler, 50 steps, and CFG scale 4.0 for sampling.
w/o REPA
w/ REPA
α
FID ↓
IS ↑
FID ↓
IS ↑
baseline
34.84
41.53
24.31
62.23
1.0
32.79
44.80
23.36
64.79
0.5
30.91
47.22
23.28
64.88
0.25
31.55
46.17
23.18
64.31
0.05
–
–
22.25
67.33
Table 2 : Effect of α on ImageNet with and without REPA. Evaluated using an ODE sampler for 250 steps (without CFG).
w/o REPA
w/ REPA
α
FID ↓
CLIP-T ↑
FID ↓
CLIP-T ↑
baseline
11.32
18.25
8.16
19.25
1.0
9.98
18.60
7.80
19.60
0.75
9.77
18.62
8.12
19.62
0.5
9.44
18.54
7.44
19.66
0.25
9.69
18.53
7.67
19.69
Table 3 : Effect of the α parameter on text-to-image generation, with and without REPA. All results are evaluated using an ODE sampler (50 steps) with CFG scale =2.0 .
α
s(⋅)
FID ↓
CLIP-T ↑
baseline w/o CARE
11.32
18.25
0.25
linear
9.69
18.53
softmax (τc=1)
9.76
18.62
softmax (τc=0.5)
9.77
18.55
0.5
linear
9.44
18.54
softmax (τc=1)
9.88
18.57
Table 4 : Ablation results of similarity measure in text-to-image generation. Evaluated using an ODE sampler for 50 steps with CFG scale =2.0
w/o REPA
w/ REPA
λCARE
FID ↓
IS ↑
FID ↓
IS ↑
baseline
34.84
41.53
24.31
62.23
0.25
30.91
47.22
21.64
68.50
0.5
32.27
45.25
21.57
68.35
Table 5 : Effect of the loss coefficient λCARE on ImageNet 256 × 256, with and without REPA. All results are evaluated using an ODE samplers for 250 steps (without CFG).
Layer
FID ↓
IS ↑
baseline
34.84
41.53
4
34.46
42.61
8
32.69
44.88
12
30.91
47.22
Table 6 : Effect of CARE injection layer on ImageNet 256 × 256 using SiT-B/2 (a 12-layer model). All results are evaluated using an ODE sampler for 250 steps (without CFG).
Method
Iter.
FID ↓
IS ↑
SiT-B/2 + disp. loss
400k
32.79
44.80
SiT-B/2 + disp. loss + d-sampler
400k
32.07
45.50
SiT-B/2 + CARE
400k
30.94
47.22
Table 7 : Ablation results on enforcing label distinctness within local batch ( d-sampler ). Evaluated using an ODE sampler for 250 steps without CFG
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Evolution of the ratio r=Y/X over training steps of CARE. The ratio quickly decreases to the order of 10−4 after several hundreds steps.
Figure 6 : Uncurated samples generated by SiT-XL/2 trained with CARE for 2.4M iterations. Best view zoom in. Sampling uses a 50-step ODE sampler with CFG scale 4.0.
Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text prompts or semantic maps, which require extensively annotated datasets. In this preliminary work, we explore diffusion models conditioned on representations from a pre-trained self-supervised model. The self-conditioning mechanism not only improves the quality of unconditional image generation, but also provides a representation space that can be used to control the generation. We explore this conditioning space by identifying directions of variations, and demonstrate promising properties in terms of smoothness and disentanglement.
Nithesh Chandher Karthikeyan, Jonas Unger, Gabriel Eilertsen
Text-to-image diffusion models are capable of generating high-quality images, but suboptimal pre-trained text representations often result in these images failing to align closely with the given text prompts. Classifier-free guidance (CFG) is a popular and effective technique for improving text-image alignment in the generative process. However, CFG introduces significant computational overhead. In this paper, we present DIstilling CFG by sharpening text Embeddings (DICE) that replaces CFG in the sampling process with half the computational complexity while maintaining similar generation quality. DICE distills a CFG-based text-to-image diffusion model into a CFG-free version by refining text embeddings to replicate CFG-based directions. In this way, we avoid the computational drawbacks of CFG, enabling high-quality, well-aligned image generation at a fast sampling speed. Furthermore, examining the enhancement pattern, we identify the underlying mechanism of DICE that sharpens specific components of text embeddings to preserve semantic information while enhancing fine-grained details. Extensive experiments on multiple Stable Diffusion v1.5 variants, SDXL, and PixArt-α demonstrate the effectiveness of our method. Code is available at https://github.com/zju-pi/dice.
Zhenyu Zhou, Defang Chen, Can Wang +2
Zhejiang University, State Key Laboratory of Blockchain and Data Security · University at Buffalo, State University of New York
Data availability remains a critical bottleneck in many deep learning applications. Large-scale datasets are often expensive to collect, curate and annotate, which can limit the scalability and applicability of supervised learning methods. In this work, we evaluate the classification performance of models trained on synthetic image datasets produced by generative deep learning. In particular, we use latent diffusion models conditioned on learned representations from DINOv2, DINOv3, and CLIP. Our results demonstrates that this representation-conditioned formulation significantly outperforms class-conditioned generation by a large margin (+10.76 p.p. top-1 accuracy on ImageNet100), by improving sample quality and mode coverage. Furthermore, by scaling the size of the synthetic dataset, we are able to outperform a classifier trained on the real data (+2.0 p.p top-1 accuracy). We also demonstrate how generated images can be used for augmentation purposes, outperforming classical augmentation methods, and how the conditioning space can be used for sample filtering to further improve training value. Collectively, these findings highlight that representation-conditioned diffusion models provide a promising approach for augmenting, complementing, or potentially replacing real-world datasets in large-scale visual learning tasks.
Nithesh Chandher Karthikeyan, Jonas Unger, Gabriel Eilertsen