Emerging legislation requires large language models (LLMs) to be audited for compliance with regulatory standards, particularly fairness. Such black-box audits typically assume a single auditor with access to a large, representative set of queries. In practice, it can be difficult for an auditor to obtain such a query set, but multiple auditors can together cover the relevant demographic groups by auditing the LLM collaboratively with their individual query sets. However, relying on multiple auditors raises a fundamental trust problem, as they may act on behalf of the LLM provider to portray a misleading appearance of fairness, i.e., fairwashing. We propose Auditopus, a novel approach for robust decentralized fairness auditing. In Auditopus, auditing proceeds in rounds without a central server. In each round, every auditor issues a fixed number of queries to the LLM, and sends only cumulative statistics vectors of its query results to other auditors instead of sensitive queries in clear. The fairness of the audited LLM is then estimated by aggregating all the vectors. We show theoretically and empirically that even a single adversarial auditor in the network can steer this estimate by fabricating the vectors it sends, making an unfair LLM appear fair. To address this threat, Auditopus has each honest auditor locally down-weight any auditor whose cumulative statistics vectors are statistically inconsistent with previous ones. We implement Auditopus and compare it to robust aggregation baselines on two datasets with two pre-trained LLMs. Against an attacker that optimizes the vectors it sends to make the LLM appear fair, Auditopus reduces audit error by up to 78% on average relative to no defense and at least 62% relative to the robust aggregation baselines. Even when 49% of the auditors are adversarial, Auditopus never lets a very unfair or moderately unfair LLM pass as fair.
Figures & tables
Fig. 1: A collaborative fairness audit of a LLM ( LLM ).
Fig. 2: Local DP estimations of N=20 honest auditors on the Civil Comments dataset as data heterogeneity grows (lower α = more heterogeneous). The solid horizontal line indicates the true DP and the dotted lines indicate the 5th and 95th percentiles. Heterogeneity of private query sets makes local fairness estimates unreliable.
Fig. 3: One-shot auditing on the Civil Comments dataset (Christian vs. Muslim, true DP −0.228 ), with N=20 honest auditors holding heterogeneous data ( α=1 ). Each adversary submits fabricated counts as large as the largest honest auditor’s, all non-toxic comments about group B . Mean ± SD over 20 seeds.
Fig. 4: The workflow of Auditopus , from the perspective of an honest auditor i .
Fig. 5: Estimate of \DP with K=1,2,3 adversaries and no defense, on Civil Comments (Christian vs. Muslim), with N=20 honest auditors, α=1 and n=300 (20 seeds).
Bias in Bios Llama-3.1-8B-Instruct , zero-shot occupation prediction (one-vs-rest)
Very unfair
Gender
male vs. female
Healthcare
nurse
45,742
48,985
0.028
0.300
−0.272
TABLE I: Summary of the six audit cases and the true DP of the audited LLM . Case : very unfair, moderate or near-fair, depending on the magnitude of the true DP; Population : the inputs the auditors query, i . e ., the comments mentioning group A or B (Civil Comments) or the biographies in the given sector (Bias in Bios); Positive : the prediction defined as positive; for Bias in Bios, the multi-class occupation classifier is turned into a one-vs-rest classifier for the target occupation ( Appendix C ). ∣A∣ , ∣B∣ : number of inputs per group; Rate : proportion of positive predictions per group; True DP = Rate A− Rate B over the whole population.
Attack success rate (%)
Average DP estimation error
LLM fairness
No-defense
Median
Trimmed mean
Auditopus
No-defense
Median
Trimmed mean
Auditopus
Very unfair
95
5
0
0
0.267
0.137
0.079
0.060
Moderate
97
46
71
0
0.144
0.189
0.155
0.034
Near-fair
100
20
53
9
0.062
0.079
0.037
0.006
All three cases
97
24
41
3
0.158
0.135
0.090
0.034
TABLE II: The ASR ( ASR ) and mean absolute DP estimation error on the Bias in Bios dataset, over every run with 1 – 19 adversaries ( 5 – 49% of all auditors; 380 runs per case). We provide similar results for Civil Comments in Table IV .
Fig. 6: The attack success rate (top row) and DP estimation error (in the bottom row) for Bias in Bios, while varying the number of adversaries K from 0 to 19 . We consider a very unfair (left column), moderate (middle column) and near-fair LLM (right column). The dashed horizontal line (in the bottom row) corresponds to a DP estimate of 0. Results for the Civil Comments dataset are provided in Section D-A .
Fig. 7: The DP estimation error on the Bias in Bios dataset, for Auditopus and baselines, on the moderate LLM . We vary the number of honest auditors ( N ) and heterogeneity level ( α ). Lower is better.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Meaning
Defined in
Protocol
f
black-box LLM under audit
Sec. II-A
A , B
the two demographic groups
Sec. II-A
τ
half-width of the fair band [−τ,τ]
Sec. II-A
N , K
number of honest auditors and of adversaries
Sec. III-B
V
set of all auditors, ∣V∣=N+K
Sec. III-B
Appendix
TABLE III: Notations used throughout the paper.
Fig. 8: Prompt used to classify each comment in Civil Comments with Qwen2.5-7B-Instruct .
Fig. 9: Prompt used to classify each biography in Bias in Bios with Llama-3.1-8B-Instruct . { biography } is replaced by the biography text. The answer is constrained to the 25 occupation names in the list.
Attack success rate (%)
Average DP estimation error
LLM fairness
No-defense
Median
Trimmed mean
Auditopus
No-defense
Median
Trimmed mean
Auditopus
Very unfair
95
42
63
0
0.221
0.226
0.221
0.050
Moderate
99
33
35
0
0.186
0.235
0.219
0.022
Near-fair
97
68
80
78
0.061
0.081
0.073
0.028
All three cases
97
48
59
26
0.156
0.180
0.171
0.034
Appendix
TABLE IV: The ASR ( ASR ) and mean absolute DP estimation error on the Civil Comments dataset, over every run with 1 – 19 adversaries ( 5 – 49% of all auditors; 380 runs per case; as Table II ). Bold: the lowest in each row.
Fig. 10: The attack success rate (top row) and DP estimation error (in the bottom row) for Civil Comments, while varying the number of adversaries K from 0 to 19 . We consider a very unfair (left column), moderate (middle column) and near-fair LLM (right column). The dashed horizontal line in Figure 10 (in the bottom row) corresponds to a DP value of 0.
Fig. 11: Honest auditors only ( N=20 , no adversary), querying without and with replacement (mean ± SD over 20 runs; top rows: Civil Comments, bottom rows: Bias in Bios). Without replacement the estimate converges to the true DP; with replacement it stays where it started, off the true DP by an amount that grows with heterogeneity.
Fig. 12: The median and the 20% trimmed mean applied to each source’s local DP estimate, against no defense and Auditopus , in the setting of RQ1 (top rows: Civil Comments; bottom rows: Bias in Bios).
Fig. 13: The trimmed mean of the sources’ group rates at trimming fractions β=10 , 20 , 30 and 40% , against Auditopus , in the setting of RQ1.
Fig. 14: Attack success as the number of queries per auditor per round grows from 100 to 3,000 , with K=6 adversaries next to N=20 honest auditors ( 20 runs per point); Civil Comments (top) and Bias in Bios (bottom).
Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM applications, but suffers from resource-intensive query access. We conceptualise auditing as uncertainty estimation over a target fairness metric and introduce BAFA, the Bounded Active Fairness Auditor for query-efficient auditing of black-box LLMs. BAFA maintains a version space of surrogate models consistent with queried scores and computes uncertainty intervals for fairness metrics (e.g., Δ AUC) via constrained empirical risk minimisation. Active query selection narrows these intervals to reduce estimation error. We evaluate BAFA on two standard fairness dataset case studies: \textsc{CivilComments} and \textsc{Bias-in-Bios}, comparing against stratified sampling, power sampling, and ablations. BAFA achieves target error thresholds with up to 40× fewer queries than stratified sampling (e.g., 144 vs 5,956 queries at ε=0.02 for \textsc{CivilComments}) for tight thresholds, demonstrates substantially better performance over time, and shows lower variance across runs. These results suggest that active sampling can reduce resources needed for independent fairness auditing with LLMs, supporting continuous model evaluations.
David Hartmann, Lena Pohlmann, Lelia Hanslik +3
Weizenbaum Institut Berlin · Technische Universität Berlin · FIZ Karlsruhe +2
External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often have limited access to deployed models and must rely on query-based interactions. Most existing fairness evaluation methods assume static datasets and fixed-sample statistical tests, making them poorly suited to real-world auditing scenarios in which evidence must be collected sequentially under query constraints. In this work, we formulate fairness auditing as a tolerance-aware sequential hypothesis-testing problem under limited model output access. We develop a sequential generalized likelihood-ratio framework that allows auditors to accumulate evidence from a finite audit pool and stop once sufficient support for compliance or violation has been obtained. The framework is instantiated for decision-based Statistical Parity and Equal Opportunity audits, and extended to score- and logit-based proxy audits when richer observables are available. Our results show that both the fairness metric and the level of model access significantly affect audit efficiency, and that the benefits of richer output information are not uniform across auditing settings. In particular, richer outputs can substantially reduce the number of queries required for some fairness metrics and operating regimes, while offering limited gains in near-threshold cases. This work provides a practical statistical framework for sequential fairness auditing under realistic deployment constraints.
Ioannis Pitsiorlas, Martha V. Sourla, Marios Kountouris
EURECOM, France · DaSCI, University of Granada, Spain
Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators. Current audits of large language models (LLMs) primarily rely on black-box evaluations that assess model behavior only through input-output testing. These methods are limited to tests constructed in the input space, often generated by heuristics. In addition, many socially relevant model properties (e.g., gender bias) are abstract and difficult to measure through text-based inputs alone. To address these limitations, we propose a white-box sensitivity auditing framework for LLMs that leverages activation steering to conduct more rigorous assessments through model internals. Our auditing method conducts internal sensitivity tests by manipulating key concepts relevant to the model's intended function for the task. We demonstrate its application to bias audits in four simulated high-stakes LLM decision tasks. Our method consistently indicates substantial dependence on protected attributes in model predictions, even in settings where standard black-box evaluations suggest little or no bias. Our code is openly available at https://github.com/hannahxchen/llm-steering-audit