cs.LGJun 1, 2026

Fair Finetuning Mitigates Distribution Inference Attacks

Authors: Rakshit Naidu

Organizations: Georgia Institute of Technology

Abstract

Machine learning models trained on sensitive data can inadvertently leak population-level information about their training distributions -- a threat known as distribution inference attack (DIA). An adversary with black-box access can infer sensitive demographic properties, such as subgroup proportions, without observing any training data directly. While defenses such as differential privacy and property unlearning have been proposed, the link between fairness constraints and distributional leakage remains unexplored. We propose Fair Fine-tuning (FFt): a trained model is fine-tuned on samples from the complementary distribution under an Equalized Odds (EO) constraint. We provide a complete theoretical characterization, proving the tight bound Adv(A,Mf)ΔEOW\text{Adv}(\mathcal{A},M_f) \le Δ_{\text{EO}} \cdot W, where WW quantifies how distinguishable the two training distributions are by their sensitive-attribute composition. We also establish a necessary condition for FFt to reduce adversarial advantage and prove tightness of the bound. We evaluate across six datasets spanning tabular (ACS Income, COMPAS, German Credit), image (UTKFaces), and NLP (Bias in Bios) modalities. Rehearsal-based FFt consistently reduces the adversarial accuracy gap below the detection threshold τ!=!0.1τ!=!0.1 across all settings; on ACS Income, the gap falls from !15\sim!15% to under 44%. Our work provides the first formal bound connecting a model's measured EO disparity directly to its adversarial advantage in the DIA game, opening a new avenue for unified fairness-and-privacy defenses.

Explore similar work

May 16, 2026cs.LG

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning

Differentiable optimization layers are traditionally integrated in predict-then-optimize frameworks where a neural model estimates parameters that subsequently serve as fixed inputs to downstream decision-making optimization problems. In this work, we introduce the concept of a "fairness layer": a differentiable optimization layer appended to a model's output layer that guarantees a chosen notion of output parity is satisfied when integrated into a neural network. Additionally, we introduce an online primal-dual inference algorithm that provides provable aggregate fairness guarantees for streaming predictions with arbitrarily small batch sizes, where traditional per-batch constraints become overly restrictive. Numerical experiments demonstrate the effectiveness of the fairness layer and associated algorithm, and theoretical analysis characterizes the layer's differentiability and stability properties during model training and backpropagation. Our code for these experiments is publicly available on GitHub (https://github.com/dtroxell19/FairDL-ICML-2026.git) and our public Python package documentation can be found online: https://dtroxell19.github.io/fairness_training/.
David Troxell, Noah Roemer, Guido Montúfar
Feb 28, 2025cs.CR

Approaching the Harm of Gradient Attacks While Only Flipping Labels

Machine learning systems deployed in distributed or federated environments are highly susceptible to adversarial manipulations, particularly availability attacks -- rendering the trained model unavailable. Prior research in distributed ML has demonstrated such adversarial effects through the injection of gradients or data poisoning. In this work, we ask whether comparable degradation is still possible under a substantially more constrained action space: the adversary may only flip a limited number of labels of existing training examples, without modifying features, injecting samples, or directly controlling gradients. We analyze the extent of damage caused by constrained label flipping attacks against distributed learning under mean aggregation -- the dominant baseline in research and production. Focusing on classification problems, (1) we propose a novel formalization of label flipping attacks as a per-round constrained optimization problem, derive a greedy label-selection rule for logistic regression, and empirically evaluate it beyond its derivation setting, including on MLPs and robust aggregators. The rule is provably per-epoch optimal for the attacker under the mean aggregator. (2) Empirically, we show that optimized label flipping can cause substantial accuracy degradation while outperforming random label flipping under similar budgets. (3) We shed light on an interesting interplay between what the attacker gains from more write-access versus what they gain from more flipping budget. (4) Finally, although the attack is derived for mean aggregation, we find that it can transfer empirically to the coordinate-wise median and trimmed mean aggregators, where its effectiveness approaches that of the Little-is-Enough gradient attack. This demonstrates that even highly constrained label-flipping adversaries can pose a significant availability threat to distributed learning.
Abdessamad El-Kabid, El-Mahdi El-Mhamdi
Jul 16, 2026cs.LG

Auditing Fairness-Privacy Trade-offs: Subpopulation-Level Effects of Fairness-Enhancing Algorithms

Machine learning (ML) models deployed in sensitive domains such as healthcare, law enforcement, and finance must satisfy not only utility requirements but also fairness and privacy guarantees. While prior work has largely examined how privacy-preserving techniques affect fairness, the inverse question-how fairness-enhancing algorithms influence privacy leakage-remains underexplored. We present the first comprehensive study of how fairness interventions affect membership inference privacy risks at the subpopulation level. By adapting the Likelihood Ratio Attack (LiRA) for subgroup auditing, we uncover privacy disparities that aggregate evaluations obscure. We further analyze how Differential Privacy (DP) interacts with fairness-enhancing methods across different categories, showing that DP's privacy benefits and utility costs are unevenly distributed across subpopulations. Our results demonstrate that fairness interventions do not uniformly increase privacy risk; their impact depends on model architecture, subgroup size, and mitigation strategy. These findings reveal that fairness, privacy, and utility must be jointly evaluated at the subpopulation level, and we introduce the first unified empirical framework to support such auditing in practice.
Umid Suleymanov, Ilhama Novruzova, Khalid Mammadov +2