Recent advances in diffusion models highlight the importance of representation regularization for improving sample quality and training efficiency. However, commonly used regularization methods often overlook the built-in conditions (such as labels or texts) which directly determine the generation target. In this work, we demonstrate how conditioning signals affect the feature distribution and introduce the CARE (Condition-Aware REpresentation regularization). CARE is a lightweight plug-and-play regularization framework that dynamically modulates feature distribution based on condition similarity. CARE leverages built-in conditioning signals to judiciously guide the representation space, promoting tighter feature clusters for similar conditions without relying on explicit alignment losses or external supervision. Empirically, CARE consistently improves both visual fidelity and convergence stability across both class-to-image and text-to-image tasks. On ImageNet, CARE achieves a 19.08% reduction in FID in 400k training steps, leading to a 3.5× speed-up. When applied to text-to-image generation, CARE lowers FID by 16.61% in 200k iterations and improves semantic alignment between generated samples and text prompts. Moreover, CARE can be seamlessly integrated with existing regularization methods, yielding additional performance gains.
Figures & tables
Figure 1 : Overview of CARE. Left: DiT pipeline. Right: Representation and condition spaces with CARE regularization. CARE penalizes representation collapse for all sample pairs, with a similarity-dependent strength based on condition similarity.
Figure 2 : FID-50K on ImageNet 256×256 for SiT-XL/2 with and without CARE. All models are evaluated with SDE sampling, without CFG, using 250 sampling steps. CARE consistently improves FID throughout training.
Figure 3 : FID-50K versus representation alignment with text conditions, measured by a linear probe on intermediate representations. Across both settings with and without REPA, CARE consistently improves representation–condition alignment while reducing FID.
Method
Iter.
Sampler
NFEs
CFG scale
FID ↓
IS ↑
CLIP-T ↑
class-to-image generation
SiT-B/2
400k
SDE
250
1.0
33.02
43.71
-
SiT-B/2 + dispersive loss
400k
SDE
250
1.0
31.37
47.85
-
SiT-B/2 + CARE
400k
SDE
250
1.0
29.71
50.10
-
SiT-B/2 + REPA
400k
ODE
250
1.0
24.31
62.23
-
SiT-B/2 + REPA + dispersive loss
400k
ODE
250
1.0
23.36
64.79
-
Table 1 : CARE consistently improves diffusion models across class-to-image and text-to-image benchmarks.
Figure 4 : Qualitative comparison on ImageNet 256 × 256. Both models share the same noise, and use a Euler ODE sampler, 50 steps, and CFG scale 4.0 for sampling.
w/o REPA
w/ REPA
α
FID ↓
IS ↑
FID ↓
IS ↑
baseline
34.84
41.53
24.31
62.23
1.0
32.79
44.80
23.36
64.79
0.5
30.91
47.22
23.28
64.88
0.25
31.55
46.17
23.18
64.31
0.05
–
–
22.25
67.33
Table 2 : Effect of α on ImageNet with and without REPA. Evaluated using an ODE sampler for 250 steps (without CFG).
w/o REPA
w/ REPA
α
FID ↓
CLIP-T ↑
FID ↓
CLIP-T ↑
baseline
11.32
18.25
8.16
19.25
1.0
9.98
18.60
7.80
19.60
0.75
9.77
18.62
8.12
19.62
0.5
9.44
18.54
7.44
19.66
0.25
9.69
18.53
7.67
19.69
Table 3 : Effect of the α parameter on text-to-image generation, with and without REPA. All results are evaluated using an ODE sampler (50 steps) with CFG scale =2.0 .
α
s(⋅)
FID ↓
CLIP-T ↑
baseline w/o CARE
11.32
18.25
0.25
linear
9.69
18.53
softmax (τc=1)
9.76
18.62
softmax (τc=0.5)
9.77
18.55
0.5
linear
9.44
18.54
softmax (τc=1)
9.88
18.57
Table 4 : Ablation results of similarity measure in text-to-image generation. Evaluated using an ODE sampler for 50 steps with CFG scale =2.0
w/o REPA
w/ REPA
λCARE
FID ↓
IS ↑
FID ↓
IS ↑
baseline
34.84
41.53
24.31
62.23
0.25
30.91
47.22
21.64
68.50
0.5
32.27
45.25
21.57
68.35
Table 5 : Effect of the loss coefficient λCARE on ImageNet 256 × 256, with and without REPA. All results are evaluated using an ODE samplers for 250 steps (without CFG).
Layer
FID ↓
IS ↑
baseline
34.84
41.53
4
34.46
42.61
8
32.69
44.88
12
30.91
47.22
Table 6 : Effect of CARE injection layer on ImageNet 256 × 256 using SiT-B/2 (a 12-layer model). All results are evaluated using an ODE sampler for 250 steps (without CFG).
Method
Iter.
FID ↓
IS ↑
SiT-B/2 + disp. loss
400k
32.79
44.80
SiT-B/2 + disp. loss + d-sampler
400k
32.07
45.50
SiT-B/2 + CARE
400k
30.94
47.22
Table 7 : Ablation results on enforcing label distinctness within local batch ( d-sampler ). Evaluated using an ODE sampler for 250 steps without CFG
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Evolution of the ratio r=Y/X over training steps of CARE. The ratio quickly decreases to the order of 10−4 after several hundreds steps.
Figure 6 : Uncurated samples generated by SiT-XL/2 trained with CARE for 2.4M iterations. Best view zoom in. Sampling uses a 50-step ODE sampler with CFG scale 4.0.