stat.MLApr 8, 2025

Deep Fair Learning: Task-Aware Fair Representations via Joint Distance-Covariance Regularization

Authors: Enze Shi, Yiqun Xiao, Linglong Kong, Bei Jiang

Organizations: Department of Mathematical and Statistical Science, University of Alberta.

Abstract

Ensuring fairness is essential as machine learning increasingly informs consequential decisions. However, many fairness-aware methods focus on the outputs of individual predictors, without directly controlling sensitive information retained in the underlying representations. We propose Deep Fair Learning (DFL), which combines distance covariance regularization with predictive loss to jointly learn representations and downstream predictors, promoting fairness at both levels while preserving task-relevant information. Its marginal and class-conditional formulations target independence and separation, respectively. Under suitable regularity conditions, we establish non-asymptotic joint excess-risk rates and convergence of the learned representation up to natural invariances. We further derive fairness-inheritance bounds linking representation-level dependence to downstream disparities over suitable predictor classes, extending fairness guarantees beyond the jointly trained predictor. Experiments on tabular, text, and image benchmarks show that DFL achieves lower fairness gaps than competing methods in many evaluated settings while maintaining competitive predictive accuracy, with fairness gains largely preserved after downstream retraining.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Sep 29, 2026cs.LG

A Comprehensive View of Fairness through Distributional Stability

We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it remains stable under such shifts. Under this perspective, several classical notions of fairness arise as stability with respect to specific perturbations, with the associated unfairness gap given by a Lipschitz constant of a prediction-rate functional. This formulation also yields guarantees that hold uniformly over a range of demographic compositions at test time, without requiring knowledge of the deployment distribution. It leads to a learning procedure based on convex combinations of reweighted predictors, formulated as a second-order cone program, for which we establish generalization bounds. Experiments on standard benchmarks illustrate the approach.
Aug 16, 2025cs.LG

FAIRVAR: Fair Federated Learning via Variance Regularization

Federated learning (FL) allows collaborative training of machine learning models across multiple parties without sharing raw data. However, heterogeneous data can cause some clients to have disproportionate influence on the global model, leading to disparities in their performance. Fairness, understood as reducing these disparities, is therefore a crucial concern in FL and has been addressed in various ways. We studied performance equitable fairness in FL, where the goal is to minimize performance disparities across clients. We evaluated several existing fairness-aware methods and introduce here a new gradient-variance-regularized method, implemented in two variants: FairGrad (approximate) and FairGrad* (exact). We theoretically characterize the connections between these methods and, empirically, on heterogeneous benchmarks, show that FairGrad and FairGrad* consistently improve fairness by reducing variance in client accuracies, while maintaining competitive or improved mean performance compared to existing fairness-aware baselines.
May 3, 2026cs.CV

ProtoFair: Fair Self-Supervised Contrastive Learning via Pseudo-Counterfactual Pairs

Self-supervised learning methods learn high-quality visual representations, yet recent studies show that these representations often capture demographic biases present in the training data. Existing fairness-aware methods address this by redesigning the self-supervised objective itself, limiting portability across the rapidly evolving landscape of self-supervised learning (SSL) frameworks. We propose ProtoFair, a fairness-aware contrastive loss designed to work alongside existing SSL objectives without modifying them. ProtoFair leverages unsupervised prototype clustering to identify pseudo-counterfactual pairs: samples sharing the same cluster assignment but belonging to different sensitive groups. By pulling these content-matched, cross-group samples together in the embedding space, ProtoFair encourages the encoder to learn representations that are invariant to the sensitive attribute. The method requires only sensitive attribute annotations, no target labels, and integrates seamlessly with both SimCLR and SupCon. Experiments on CelebA and UTKFace demonstrate consistent fairness improvements while maintaining competitive accuracy.