Although diffusion models produce high-quality images, they also reproduce and amplify demographic imbalances in their training data. Debiasing their generation process post-training w.r.t. some sensitive attribute usually relies on classifier guidance or explicit text extra-conditioning, but this reduces methods' applicability and output diversity. Conversely, methods promoting diversity alone do not ensure fair attribute representation. In this paper, we propose a method tackling fairness and diversity jointly that is generally applicable to any diffusion model and any sensitive attribute. To this end, an adapter connects the frozen diffusion model to a pretrained vision-language embedding space, enabling fairness and diversity guidance without sensitive-attribute annotations. For fairness, pairs of text prompts define attribute directions which guide batch composition towards specific proportions. For diversity, we introduce a score measuring disagreement between the semantic estimates derived from this representation. The formulation supports unconditional and text-conditional diffusion models, while requiring no prior knowledge or data of sensitive attribute. Experiments confirm that our method improves quality and diversity scores at comparable fairness levels.
Figures & tables
Figure 1: Generations examples on CelebA-HQ. Top : the original model, 6 random draws of women; two of them appear with artifacts, the third looking away from the camera and the fourth turning her head. Middle : guidance with the fairness term. Three columns turn into men and the faces sharpen, but the pose collapses, every face now frontal and centered. Bottom : guidance with the fairness and diversity terms. The woman in the second column looks older, the man in the fourth turns his head away and gains a mustache.
no attribute supervision
arbitrary proportion
arbitrary conditioning
no retraining
attribute
Fair Diffusion ( Friedrich et al., 2025 )
✓
✓
✗
✓
text
ITI-GEN ( Zhang et al., 2023 )
✗
✓
✗
✓
ref. images
Self-debiasing ( Vardhana et al., 2026 )
✓
✗
✓
✓
discovered
Latent editing ( Kwon et al., 2023 )
✓
✗
✓
✓
text, training per attribute
Latent direction ( Li et al., 2024 )
✓
✓
✗
✓
text, training per attribute
FairGen ( Jiang et al., 2025 )
✓
✓
✗
✗
text, training per attribute
Table 1: Positioning among debiasing methods for diffusion models. A method may need attribute supervision (e.g. labels, reference images), may not accept an arbitrary target proportion, may not have been applied to both unconditional and text-conditioned models, and may retrain the generator. ( ∗ ) Score Guidance also presents an alternate supervision-free version, but with significantly diminished performance compared to the supervised version, which is the one we compare with.
Figure 2: One guided sampling step. The prompts tk are set once and encoded into directions Δk . At each step, the adapter φ maps an inner layer h(xσ) of the frozen denoiser to m(xσ;σ) for every image of the batch, where the fairness and diversity terms are computed.
Figure 3: CelebA-HQ, P2 backbone, six shared seeds per column. Top: unguided samples. Second row: Balancing Act steered towards a dark-skinned woman with its labelled predictor. Third row: Debias Anything on the same target, given as two sentences. Bottom: Debias Anything on a second attribute with no CelebA-HQ label, red hair, from the same seeds.
Table 2: CelebA-HQ, P2 backbone, protocol of Balancing Act ( 10000 images, one seed); all rows are our runs. FD is CLEAM-corrected except on race, which has no label to correct against. Guided rows are shaded from best (dark, bold) to worst (light); the bracket marks the rows matched on discrepancy.
Gender
Age
Race
Method
FD↓
FID↓
FD↓
FID↓
FD↓
FID↓
Original
0.564
120.1
0.752
120.1
0.558
120.1
Attribute-specific
Latent editing ( Kwon et al., 2023 )
0.408
166.1
0.682
200.9
0.524
153.1
h -distribution ( Parihar et al., 2024 )
0.222
151.7
0.506
147.7
0.544
126.9
Latent direction ( Li et al., 2024 )
0.305
129.4
0.052
113.8
0.175
128.3
Table 3: Stable Diffusion 1.5, protocol of Shi et al. (2025) , whose reported numbers fill the baseline rows bar ICM. All baselines incorporate attribute knowledge (labeled data, attribute-specific training); ours only gets two sentences. Guided rows are shaded from best (dark, bold) to worst (light).
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Data
Noise levels
Optimisation
Epochs
Time
EDM, CelebA
162,770 training images
σ log-uniform on [0.002,80]
AdamW, batch 512, lr 10−3
50
2.9 h
P2, CelebA-HQ
28,000 ( +2,000 val.)
t uniform on the 49 guided steps
AdamW, batch 32, lr 10−3 , wd 10−4
47 of 60
4.2 h
SD 1.5
49,000 generated ( +1,000 val.)
σ uniform on [0.03,14.6]
AdamW, batch 64, lr 5⋅10−4 , wd 0.05 , warm-up and cosine
47 of 50
11 h
Appendix
Table 4: Adapter training, on one A100 GPU.
Figure 4: What the CelebA adapter reads of the attribute as the noise grows. AUC of the logit S(xσ;σ) along the direction Δ of Equation 15 , against the CelebA label, on 500 images per class of the test split, for gender ( a photo of a man / a photo of a woman , left) and eyeglasses ( a photo of a person wearing eyeglasses / a photo of a person , right). Circles use the adapter output φψ(h(xσ);σ) , squares the encoder on the Tweedie estimate, Eϑimg(x^0) ; the dotted line is the encoder on the clean image, Eϑimg(x0) . Over the guidance window σ≤10 the adapter matches the Tweedie estimate on gender and exceeds it on eyeglasses, at no decoding cost.
Figure 5: Steering an attribute that has no label. Top: unguided samples from a diffusion model trained on CelebA. Bottom: the same eight initial noises, guided towards dark-skinned female with sunglasses . The attribute is handed to the sampler as this sentence and a contrasting one describing its absence. CelebA has no label for it, and neither the generator nor any classifier was trained.
Figure 6: Same compound target, dark-skinned female with sunglasses , three ways of reading the guidance term. Top: from φ(h(xσ);σ) . Middle: from Eϑimg(x^0) at the same sampling budget. Bottom: from Eϑimg(x^0) with more than twice the budget. Eight samples per row.
h-space guidance
pixel guidance
w
MIND
FD
FID
MIND
FD
FID
Eyeglasses attribute ( FD=0.614 unguided)
0
19.13
0.614
16.23
19.13
0.614
16.23
50
15.14
0.534
13.81
19.69
0.611
16.92
100
11.81
0.441
12.14
19.88
0.598
17.66
150
9.97
0.377
11.60
20.38
0.587
18.56
Appendix
Table 5: h -space against pixel-space fairness term at equal weight w ; FD on the steered attribute; both share the w=0 run, 5000 images, one seed. Bold: the better MIND of the two on each row.
Figure 7: FD– MIND trade-off, h-space versus pixel-space fairness term; each polyline follows its own sweep in order of increasing guidance weight w , and the white square is the shared unguided model. Left: on eyeglasses the two curves do not overlap, pixel guidance stays pinned in the top-right corner, drifting rightwards as w grows, while h-space sweeps down and to the left, then turns back up past w=200 as fidelity starts to cost, so no pixel setting matches any h-space setting on both axes. Right: on gender the two sweeps trace the same short hook, pixel running marginally inside h-space; both bottom out in MIND around w=5 and then climb steeply as FD is driven to zero.
Method
wdiv
MIND↓
FD↓
FID↓
Vendi ↑
Prec. ↑
Rec. ↑
Cov. ↑
Dens. ↑
avgKNN
0
10.76
0.336
19.21
6.47
0.814
0.679
0.893
1.218
2
10.56
0.332
19.21
6.52
0.812
0.689
0.893
1.208
5 ⋆
10.44
0.330
19.39
6.58
0.805
0.690
0.893
1.190
10
10.46
0.330
19.78
6.69
0.795
0.697
0.880
1.159
20
11.24
0.317
20.77
6.88
0.778
0.692
0.870
1.131
40
14.72
0.309
23.80
7.31
0.765
0.696
0.835
1.035
Appendix
Table 6: Diversity-weight sweep on eyeglasses, fairness weight fixed at 200 . wdiv=0 is the fairness-only baseline, 2000 images, one seed; ⋆ marks the setting carried to Table 8 .
Method
wdiv
MIND↓
FD↓
FID↓
Vendi ↑
Prec. ↑
Rec. ↑
Cov. ↑
Dens. ↑
avgKNN
0
2.76
0.0104
12.41
6.18
0.829
0.723
0.970
1.133
2
2.65
0.0075
12.40
6.24
0.828
0.723
0.969
1.124
5 ⋆
2.54
0.0039
12.39
6.32
0.832
0.732
0.973
1.113
10
2.56
0.0017
12.54
6.45
0.820
0.734
0.963
1.094
20
3.42
0.0019
13.28
6.72
0.824
0.732
0.967
1.060
40
8.34
0.0121
16.81
7.26
0.792
0.749
0.902
0.936
Appendix
Table 7: Diversity-weight sweep, on gender, fairness weight fixed at 10 . wdiv=0 is the fairness-only baseline, 2000 images, one seed; ⋆ marks the setting carried to Table 8 .
time (s)
FD↓
FID↓
MIND↓
Vendi ↑
Gender, wfair=5
Debias Anything
1613±665
0.031±0.004
6.59±0.04
2.52±0.13
6.26±0.03
+ SGMS
1380±5
0.023±0.005
6.46±0.05
2.05±0.14
6.46±0.02
+ avgKNN
1635±713
0.021±0.003
6.40±0.11
2.06±0.15
6.42±0.03
+ SigLIPMS
2179±962
0.011±0.004
6.48±0.07
2.13±0.17
6.41±0.04
+ PerturbProj
1476±592
0.006±0.003
6.44±0.10
2.03±0.15
6.43±0.04
Appendix
Table 8: Diversity terms on top of the fairness term at the selected weights, 5000 images, 5 seeds; time is wall-clock generation of one run and is not ranked.
time (s)
FD↓
FID↓
MIND↓
Vendi ↑
Gender, wfair=5
Debias Anything
11065
0.024
2.48
2.14
6.33
+ SGMS
15747
0.018
2.24
1.61
6.55
+ avgKNN
10942
0.015
2.22
1.63
6.50
+ SigLIPMS
15560
0.008
2.29
1.72
6.48
+ PerturbProj
11417
0.002
2.23
1.59
6.52
Appendix
Table 9: As Table 8 , with 50000 images per configuration and one seed.
Method
Setting
MIND↓
Prec. ↑
Rec. ↑
Cov. ↑
Dens. ↑
Vendi ↑
Original
—
124.10
0.972
0.389
0.809
3.010
5.94
avgKNN
κ=50,wdiv=30
62.71
0.954
0.425
0.874
2.875
7.33
SGMS
σs=0.8,wdiv=40
100.96
0.955
0.468
0.841
2.783
6.63
PerturbProj
σs=0.4,wdiv=250
92.34
0.958
0.386
0.837
2.909
6.56
SigLIPMS
wdiv=250
71.17
0.953
0.332
0.865
3.076
6.96
Appendix
Table 10: Each diversity term alone, at its lowest- MIND setting with precision ≥0.95 in Tables 12 , 14 , 14 and 12 .
Figure 8: MIND against the diversity guidance weight wdiv , one curve per hyperparameter configuration. All four panels share the MIND axis. avgKNN is nearly insensitive to κ over two orders of magnitude (seven superimposed curves, κ=1 to κ=100 ). For SGMS and PerturbProj the perturbation scale is non-monotone: σs=0.2 is markedly worse than 0.4 , so the useful range is σs∈[0.4,0.6] . SGMS is the only method reaching MIND below 60 , and it does so at a precision below 0.9 ( Table 14 ). SigLIPMS turns around: MIND bottoms out at wdiv=250 and rises again beyond it.
Table 19Table 20
Figure 9: Gender on CelebA-HQ, six shared seeds: the unguided model, Balancing Act, and Debias Anything at matched discrepancy.
Figure 10: Sensitivity to the prompt pair, gender on CelebA-HQ, wfair=20 , six shared seeds. Top: unguided. Then, one pair per row: a photo of a woman / a photo of a man ; a female person / a male person ; a girl’s portrait / a boy’s portrait ; a photo of a lady / a photo of a gentleman ; a woman’s close-up photo / a man’s close-up photo .
Figure 11: Stable Diffusion 1.5, prompts A face of a doctor (top pair) and A face of a firefighter (bottom pair), six shared seeds each: the unguided model and Debias Anything at a gender target of one half. Two doctor columns and three firefighter columns change gender; clothing, pose, background and expression stay.
Occupation
FD gender ↓
FD age ↓
FD race ↓
doctor
0.021
0.092
0.208
firefighter
0.104
0.276
0.193
nurse
0.296
0.077
0.206
receptionist
0.378
0.381
0.174
mean ( Table 3 )
0.200
0.206
0.195
Appendix
Table 15: Debias Anything on Stable Diffusion 1.5, discrepancy per occupation. Each entry is the distance of the mean FairFace output to the uniform target on the 500 images of one prompt; the last row is their average, the FD of Table 3 .
Gender
Age
Race
Method
Attribute
FD↓
FID↓
CLIP-I ↑
CLIP-T ↑
FD↓
FID↓
CLIP-I ↑
CLIP-T ↑
FD↓
FID↓
CLIP-I ↑
CLIP-T ↑
Original
—
0.564
120.06
–
0.6155
0.752
120.06
–
0.6155
0.558
120.06
–
0.6155
Feature or weight editing
Latent editing ( Kwon et al., 2023 )
text, training per attribute
0.408
166.11
0.8253
0.6005
0.682
200.90
0.8527
0.6122
0.524
153.05
0.8804
0.6086
Latent direction ( Li et al., 2024 )
text, training per attribute
0.305
129.37
0.8058
0.6091
0.052
113.81
0.8151
0.6067
0.175
128.30
0.8211
0.6132
Fine-tuning ( Shen et al., 2024 )
attribute classifier
0.050
161.47
0.8779
0.6095
0.746
161.47
0.8779
0.6095
0.198
161.47
0.8779
0.6095
Appendix
Table 16: Stable Diffusion 1.5 ( Rombach et al., 2022 ) under the protocol of DiffLens ( Shi et al., 2025 ) , all four metrics. Baseline rows other than ICM ( Zaleska et al., 2026 ) are as reported by Shi et al. (2025) , which gives fine-tuning ( Shen et al., 2024 ) a single FID , CLIP-I and CLIP-T for the three attributes. Guided rows are shaded from best (dark, bold) to worst (light); equal values share a shade. The unguided row has no CLIP-I by definition.