Ensuring fairness is essential as machine learning increasingly informs consequential decisions. However, many fairness-aware methods focus on the outputs of individual predictors, without directly controlling sensitive information retained in the underlying representations. We propose Deep Fair Learning (DFL), which combines distance covariance regularization with predictive loss to jointly learn representations and downstream predictors, promoting fairness at both levels while preserving task-relevant information. Its marginal and class-conditional formulations target independence and separation, respectively. Under suitable regularity conditions, we establish non-asymptotic joint excess-risk rates and convergence of the learned representation up to natural invariances. We further derive fairness-inheritance bounds linking representation-level dependence to downstream disparities over suitable predictor classes, extending fairness guarantees beyond the jointly trained predictor. Experiments on tabular, text, and image benchmarks show that DFL achieves lower fairness gaps than competing methods in many evaluated settings while maintaining competitive predictive accuracy, with fairness gains largely preserved after downstream retraining.
Figures & tables
Method
Dependence measure
Level
Joint
Sep.
Guarantee
AdvDebias ( Zhang et al., 2018 )
adversarial (implicit)
output
—
✓
—
FairMixup ( Chuang and Mroueh, 2021 )
path-wise DP surrogate
output
—
✓
—
DRAlign ( Li et al., 2023 )
distributional alignment
output
—
✓
—
DiffMCDP ( Jin et al., 2024 )
smoothed MCDP surrogate
output
—
—
—
INLP ( Ravfogel et al., 2020 )
linear predictability
repr. (post-hoc)
—
—
—
RLACE ( Ravfogel et al., 2022a )
linear adversarial game
repr. (post-hoc)
—
—
linear-erasure
Table 1: Comparison with related fairness methods. “Joint” denotes encoder–head training driven by the task loss; “Sep.” denotes support for class-conditional fairness. Guarantees describe the published formulations. The experiments evaluate the subset listed in Section 5.1 .
Method
Accuracy ↑
TPR ↓
MCDP ↓
STD
77.67
3.55
12.46
AdvDebias
76.49(0.65)
3.27(0.28)
29.32(2.31)
FairMixup
74.48(0.43)
3.14(0.23)
24.91(1.58)
DRAlign
75.01(0.41)
3.66(0.17)
22.04(1.22)
DiffMCDP
73.93(0.31)
3.28(0.38)
11.50(1.09)
INLP
68.27(0.57)
3.96(0.44)
8.72(1.31)
Table 2: Predictive accuracy and fairness on Adult and Bank. Parentheses give standard deviations. Among fairness-aware methods, best in bold , second best underlined .
Method
Accuracy ↑
TPR ↓
MCDP ↓
STD
79.79
15.59
6.67
AdvDebias
77.51(1.32)
13.51(1.24)
10.89(1.33)
FairMixup
75.01(0.83)
12.58(1.43)
11.72(1.29)
DRAlign
74.68(1.15)
11.36(1.31)
9.27(0.83)
DiffMCDP
75.18(1.30)
10.04(1.03)
6.42(0.76)
INLP
71.72(1.04)
9.68(0.44)
6.70(0.56)
Table 3: Predictive accuracy and fairness on BIOS and MOJI. Parentheses give standard deviations. Bold and underline mark the best and second-best fairness-aware means.
Y
Metric
STD
DFL
DiffMCDP
RLACE
CFair-EO
LAFTR-EO-CE
Young
Accuracy ↑
85.01
79.48(0.92)
79.15(0.74)
81.91(1.22)
82.10(0.81)
82.08(0.43)
TPR Gap ↓
26.70
2.43(2.01)
12.29(1.14)
6.59(2.45)
2.70(1.43)
2.22(1.74)
MCDP Gap ↓
49.74
11.07(2.04)
20.79(2.38)
26.47(2.78)
17.73(5.54)
11.92(3.36)
Attractive
Accuracy ↑
78.32
75.53(1.87)
74.79(0.85)
75.43(0.95)
74.59(0.48)
74.67(0.49)
TPR Gap ↓
40.50
2.89(1.31)
18.66(1.45)
15.15(3.12)
7.84(2.34)
8.38(2.96)
MCDP Gap ↓
57.20
21.15(2.59)
23.91(2.67)
30.42(3.15)
27.11(2.08)
28.39(2.44)
Table 4: Predictive accuracy and fairness on CelebA with Z= gender. Parentheses give standard deviations. Best and second-best fairness-aware means are bold and underlined. Other sensitive attributes are in Appendix C .
Accuracy ↑
TPR gap ↓
MCDP gap ↓
Dataset
STD
DFL head
fresh head ( δ )
STD
DFL head
fresh head ( δ )
STD
DFL head
fresh head ( δ )
Adult
77.67
79.24
79.32 ( +0.08 )
3.55
3.43
3.47 ( +0.04 )
12.46
7.31
7.52 ( +0.21 )
Bank
90.69
90.55
90.13 ( −0.42 )
2.42
2.25
2.73 ( +0.48 )
11.31
8.23
8.66 ( +0.43 )
BIOS
79.79
77.94
77.83 ( −0.11 )
15.59
9.01
9.52 ( +0.51 )
6.67
6.24
6.83 ( +0.59 )
MOJI
71.19
74.59
74.16 ( −0.43 )
38.95
10.47
10.93 ( +0.46 )
39.86
8.41
9.12 ( +0.71 )
CelebA
83.52
80.10
79.45 ( −0.65 )
29.14
3.49
5.49 ( +2.00 )
48.13
18.04
19.54 ( +1.50 )
Table 5: Persistence of fairness gains after downstream retraining. “DFL head” uses the jointly trained classifier fϕ^ ; “fresh head” discards it and trains an unconstrained one-hidden-layer classifier on the frozen representation gθ^ . δ is (fresh head − DFL head). The CelebA row uses Z= gender, Y= Wavy Hair; the remaining CelebA cells are in Appendix C . Means over 20 replications.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
p
K
Growth Rate
Depth
Reduction Factor
α
Adult
101
2
20
10
0.2
0.3
Bank
62
2
20
10
0.2
0.1
BIOS
768
28
64
10
0.2
0.7
MOJI
2304
2
64
10
0.2
0.2
CelebA
512
2
64
10
0.2
0.5
Appendix
Table B.1: DFL benchmark configurations.
Figure C.1: DFL fine-tuning performance on Adult and Bank
Figure C.2: DFL fine-tuning performance on BIOS and MOJI.
Target Attribute Y
Metric
Standard
DFL
DiffMCDP
RLACE
CFair-EO
LAFTR-EO-CE
Young
Accuracy ↑
85.01
81.68(1.41)
81.38(0.78)
82.33(1.11)
82.55(0.23)
82.38(0.45)
TPR Gap ↓
19.50
8.39(2.47)
8.97(0.95)
12.10(1.67)
8.72(1.37)
8.55(1.12)
MCDP Gap ↓
11.40
6.18(0.83)
8.72(0.87)
9.07(1.31)
5.46(1.41)
4.73(0.98)
Attractive
Accuracy ↑
78.32
77.01(1.09)
76.85(0.88)
76.50(1.25)
76.24(0.71)
77.15(0.63)
TPR Gap ↓
5.83
5.78(0.43)
5.68(0.74)
4.62(1.44)
5.36(0.17)
5.58(0.50)
MCDP Gap ↓
5.63
3.13(0.66)
3.35(0.65)
4.99(1.04)
3.34(0.47)
3.19(0.46)
Appendix
Table C.1: CelebA results with Black Hair as the sensitive attribute. Parentheses give standard deviations. Best and second-best fairness-aware means are bold and underlined.
Target Attribute Y
Metric
Standard
DFL
DiffMCDP
RLACE
CFair-EO
LAFTR-EO-CE
Young
Accuracy ↑
85.01
84.53(0.44)
83.20(0.75)
81.88(1.10)
84.23(0.50)
84.15(0.57)
TPR Gap ↓
8.07
6.62(1.32)
7.72(0.68)
7.01(1.23)
6.20(0.86)
6.80(0.69)
MCDP Gap ↓
18.58
9.94(0.91)
9.77(1.04)
9.89(1.88)
8.74(1.81)
7.34(0.91)
Attractive
Accuracy ↑
78.32
77.51(0.53)
77.96(0.81)
77.48(1.14)
75.81(0.99)
75.99(0.94)
TPR Gap ↓
13.90
9.13(1.65)
10.40(0.92)
10.63(1.90)
8.05(1.40)
8.49(2.24)
MCDP Gap ↓
23.10
18.93(0.76)
16.65(1.39)
18.29(2.05)
17.18(1.11)
19.08(1.56)
Appendix
Table C.2: CelebA results with Pale Skin as the sensitive attribute. Parentheses give standard deviations. Best and second-best fairness-aware means are bold and underlined.
Y
Young
Attractive
Smiling
Wavy Hair
Metric
Z
STD
DFL
STD
DFL
STD
DFL
STD
DFL
Accuracy ↑
Male
85.01
79.12(1.08)
78.32
74.85(1.65)
80.51
77.06(1.34)
83.52
79.45(0.85)
Black Hair
85.01
81.43(0.99)
78.32
76.54(1.03)
80.51
78.81(1.85)
83.52
81.69(1.56)
Pale Skin
85.01
83.98(0.67)
78.32
77.15(0.89)
80.51
79.35(1.42)
83.52
82.91(0.68)
TPR Gap ↓
Male
26.70
4.85(2.45)
40.50
5.91(1.74)
7.65
6.14(2.03)
29.14
5.49(2.78)
Black Hair
19.50
9.81(2.93)
5.83
7.02(0.87)
0.57
1.06(1.04)
9.75
4.15(1.06)
Appendix
Table C.3: Results for the CelebA dataset with a single sensitive attribute, using the frozen DFL encoder gθ^ with a freshly trained unconstrained head (the head-replacement setting of Section 5.3 ).
Y
Young
Attractive
Smiling
Wavy Hair
Metric
Z
STD
DFL
STD
DFL
STD
DFL
STD
DFL
Accuracy ↑
Male
85.01
81.36(1.12)
78.32
75.14(1.45)
80.51
79.54(1.02)
83.52
82.11(1.06)
Black Hair
85.01
81.36(1.12)
78.32
75.14(1.45)
80.51
79.54(1.02)
83.52
82.11(1.06)
Pale Skin
85.01
81.36(1.12)
78.32
75.14(1.45)
80.51
79.54(1.02)
83.52
82.11(1.06)
TPR Gap ↓
Male
26.73
5.94(3.74)
40.50
15.12(3.11)
7.64
7.35(2.72)
29.13
13.24(2.85)
Black Hair
19.52
15.02(1.87)
5.84
5.94(1.64)
0.57
1.15(1.43)
9.75
5.31(1.84)
Appendix
Table C.4: Results for the CelebA dataset with multiple sensitive attributes, using the frozen DFL encoder gθ^ with a freshly trained unconstrained head.
Y
Young
Attractive
Smiling
Wavy Hair
Metric
Z
STD
DFL
STD
DFL
STD
DFL
STD
DFL
Accuracy ↑
Male
85.01
81.79(0.83)
78.32
75.87(0.90)
80.51
80.31(0.66)
83.52
82.70(0.46)
Black Hair
85.01
81.79(0.83)
78.32
75.87(0.90)
80.51
80.31(0.66)
83.52
82.70(0.46)
Pale Skin
85.01
81.79(0.83)
78.32
75.87(0.90)
80.51
80.31(0.66)
83.52
82.70(0.46)
TPR Gap ↓
Male
26.73
3.94(3.41)
40.50
13.45(4.79)
7.64
6.93(2.12)
29.13
12.08(4.39)
Black Hair
19.52
12.89(1.67)
5.84
6.03(1.09)
0.57
1.65(1.19)
9.75
4.68(1.63)
Appendix
Table C.5: Results for the CelebA dataset with multiple sensitive attributes using DFL. Mean (standard deviation) over 20 replications. Marginal per-coordinate gaps; see Appendix A.2 for why these are not intersectional.
Objective
Composition
Accuracy ↑
TPR gap ↓
MCDP gap ↓
L1
α(T1−T2)+(1−α)T3 (full objective)
78.09 (0.42)
9.08 (0.65)
6.21 (0.81)
L2
αT1+(1−α)T3 (no informativeness term)
77.85 (0.38)
8.97 (0.74)
6.52 (0.71)
L3
α(T1−T2)+(1−α)(T3+T4)
61.46 (0.33)
7.35 (0.69)
6.99 (0.66)
L4
αT1+(1−α)(T3+T4)
60.52 (0.49)
4.28 (0.75)
8.52 (0.77)
L5
αT1+(1−α)T4 (prediction-level DC only)
64.78 (0.36)
5.12 (0.70)
7.76 (0.62)
L6
−αT2+(1−α)T3 (no task loss)
74.79 (0.51)
11.01 (0.62)
7.83 (0.74)
Appendix
Table C.6: Effect of the objective components on BIOS. Mean (standard deviation) over 5 replications, all other settings fixed. L2 versus L5 isolates representation-level against prediction-level dependence control; L1 versus L2 isolates the informativeness term T2 .
α
Accuracy ↑
TPR gap ↓
MCDP gap ↓
0.1
48.53 (0.39)
4.77 (0.51)
5.38 (0.63)
0.3
62.33 (0.48)
8.76 (0.66)
6.31 (0.61)
0.5
71.67 (0.37)
8.52 (0.72)
6.04 (0.68)
0.7
77.82 (0.52)
9.44 (0.62)
6.42 (0.69)
0.9
78.76 (0.47)
12.02 (0.73)
6.83 (0.71)
Appendix
Table C.7: Sensitivity to the fairness–utility weight on BIOS. Mean (standard deviation) over 5 replications; depth 10, growth rate 64 throughout.
Alpha
Depth
Growth Rate
Parameters
Acc ↑
TPR Gap ↓
MCDP Gap ↓
0.7
10
32
132139
77.45 (0.44)
9.93 (0.64)
6.51 (0.76)
0.7
10
64
295030
77.67 (0.37)
9.52 (0.72)
6.04 (0.68)
0.7
10
96
510973
77.82 (0.33)
9.49 (0.71)
5.53 (0.70)
0.7
20
32
206187
78.42 (0.36)
9.41 (0.67)
5.84 (0.63)
0.7
20
64
529775
78.61 (0.41)
10.21 (0.75)
6.11 (0.53)
0.7
20
96
991071
78.86 (0.39)
10.03 (0.69)
5.63 (0.66)
Appendix
Table C.8: Performance of DFL on BIOS dataset under different DenseNet architecture (growth rates and depths). Values are reported as mean (std) for 5 replications.
We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it remains stable under such shifts. Under this perspective, several classical notions of fairness arise as stability with respect to specific perturbations, with the associated unfairness gap given by a Lipschitz constant of a prediction-rate functional. This formulation also yields guarantees that hold uniformly over a range of demographic compositions at test time, without requiring knowledge of the deployment distribution. It leads to a learning procedure based on convex combinations of reweighted predictors, formulated as a second-order cone program, for which we establish generalization bounds. Experiments on standard benchmarks illustrate the approach.
Gayane Taturyan, Charlotte Laclau, Stephan Clémencon
LTCI, Télécom Paris, Institut Polytechnique de Paris, Palaiseau, France
Federated learning (FL) allows collaborative training of machine learning models across multiple parties without sharing raw data. However, heterogeneous data can cause some clients to have disproportionate influence on the global model, leading to disparities in their performance. Fairness, understood as reducing these disparities, is therefore a crucial concern in FL and has been addressed in various ways. We studied performance equitable fairness in FL, where the goal is to minimize performance disparities across clients. We evaluated several existing fairness-aware methods and introduce here a new gradient-variance-regularized method, implemented in two variants: FairGrad (approximate) and FairGrad* (exact). We theoretically characterize the connections between these methods and, empirically, on heterogeneous benchmarks, show that FairGrad and FairGrad* consistently improve fairness by reducing variance in client accuracies, while maintaining competitive or improved mean performance compared to existing fairness-aware baselines.
Zahra Kharaghani, Ali Dadras, Tommy Löfstedt
Department of Computing Science, Umeå University, Umeå, Sweden · Department of Mathematics and Mathematical Statistics, Umeå University, Umeå, Sweden
Self-supervised learning methods learn high-quality visual representations, yet recent studies show that these representations often capture demographic biases present in the training data. Existing fairness-aware methods address this by redesigning the self-supervised objective itself, limiting portability across the rapidly evolving landscape of self-supervised learning (SSL) frameworks. We propose ProtoFair, a fairness-aware contrastive loss designed to work alongside existing SSL objectives without modifying them. ProtoFair leverages unsupervised prototype clustering to identify pseudo-counterfactual pairs: samples sharing the same cluster assignment but belonging to different sensitive groups. By pulling these content-matched, cross-group samples together in the embedding space, ProtoFair encourages the encoder to learn representations that are invariant to the sensitive attribute. The method requires only sensitive attribute annotations, no target labels, and integrates seamlessly with both SimCLR and SupCon. Experiments on CelebA and UTKFace demonstrate consistent fairness improvements while maintaining competitive accuracy.