Latent space bias directions in LLMs capture confidence, not fairness
Authors: Stephanie Buttigieg, Maeve Madigan, Parameswaran Kamalaruban, Stuart Burrell
Organizations: Institute of Astronomy & Kavli Institute for Cosmology, University of Cambridge, UK · Risk and Security AI Lab, Visa Inc., Cambridge, UK
Activation steering has gained popularity as a lightweight inference-time debiasing technique for large language models. However, prior work reports that steering vectors generalise poorly, with unintended effects on model performance and limited transfer to new datasets. Our work analyses what the debiasing direction used for activation steering actually encodes, in order to shed light on its inconsistent performance. We study the linear debiasing direction obtained by contrasting the activations of anti-biased and biased prompts, and evaluate it as a steering intervention across bias and general knowledge benchmarks. We find that this direction is dominated by model confidence, pointing from regions of high to low-probability tokens in activation space rather than encoding a meaningful representation of model bias. Steering along it does reduce measured bias, but this is a consequence of reducing model confidence: on QA benchmarks we find that this steering drives the model to abstain from answering, with a side effect of improving fairness metrics. Our experiments show that model confidence is the dominant separating factor between biased and anti-biased prompts in hidden space, indicating that isolating a linear representation of bias which is disentangled from model confidence is difficult and steering-based debiasing results should be interpreted with care. In short, steering appears to reduce bias, not by correcting the model's underlying preferences, but by making it less confident, even on tasks unrelated to bias.
Figures & tables
Figure 1. An illustration of the paper’s two central claims, for Llama-3.1-8B-Instruct. Left : distribution of the linear classifier bias score P(biased) , trained on ambiguous BBQ contexts, evaluated on BBQ disambiguated contexts and labelled as biased (pink) or anti-biased (teal) answer type. The linear classifier poorly separates the biased and anti-biased prompts, achieving an AUROC of only 0.57. Centre : Same as the left panel but labelled by whether the scored answer is the model’s highest- (purple) or lowest- (green) probability choice. The classifier achieves a higher separability on confidence than on bias, with an AUROC of 0.94. Right : mean probability assigned to biased (pink), anti-biased (teal) and neutral/abstain (grey) answers on BBQ (ambig.) as a function of activation-steering strength α along the debiasing direction. Steering increases the rate of abstention rather than correcting the underlying biased/anti-biased preference, consistent with the direction being dominated by confidence rather than bias. Bias-probe scores separate by confidence far better than by bias class; steering increases abstention. Three-panel figure. Left: histograms of the probe's bias score for biased versus anti-biased answers on BBQ disambiguated questions – the two distributions overlap almost completely (AUROC 0.57). Middle: the same scores split instead by whether the answer was the model's highest- or lowest-probability choice – the distributions separate almost completely (AUROC 0.94). Right: mean answer probability against steering strength α from −1 to 2 for biased, anti-biased, and neutral (abstain) answers – the biased and anti-biased curves decline together from about 0.5 to 0.2, while the neutral curve rises from near 0 to about 0.65 and comes to dominate.
Figure 2. Model-assigned probability distributions for each answer option in the test set of BBQ prompts, coloured by answer type: biased (pink), anti-biased (teal), and neutral (grey). Ambiguous and disambiguated contexts are shown in the left and right panels respectively. Vertical dashed lines mark the mean of each distribution. The inset shows the proportion of greedy-decoding responses falling into each answer type. Biased answers get systematically higher probability than anti-biased answers in both contexts. Two-panel figure comparing the model's answer probability, P(answer∣prompt), for biased, anti-biased, and neutral answers under ambiguous (left) and disambiguous (right) BBQ contexts, with dashed vertical lines marking each distribution's mean and inset bar charts showing how often each answer type is selected. In both panels, the biased (pink) distribution sits slightly to the right of the anti-biased (teal) distribution, and the two otherwise overlap heavily. In the ambiguous panel, the mean probability is 0.41 for biased versus 0.36 for anti-biased, and the model selects the biased answer 48.6% of the time versus 36.2% for anti-biased. In the disambiguous panel the gap persists – mean probability 0.44 for biased versus 0.41 for anti-biased, selected 49.8% versus 45.8% of the time – even though disambiguation should make the two answers equally easy to justify. The neutral (gray) answer is a distinct, lower-probability third option in both contexts (mean 0.228 ambiguous, 0.147 disambiguous), and is selected far less often after disambiguation (15.2% down to 4.4%).
Figure 3. AUROC score of a linear classifier trained on ambiguous biased/anti-biased BBQ prompts as a function of layer number for a selection of models. Base models are shown as the solid lines, while the instruction-tuned version of each model is shown in the same colour but in dashed lines. The circular marker indicates the best layer with the highest AUROC value. AUROC by layer rises sharply from chance to a peak, then declines, but the layer of onset and peak varies by model. Line plot of linear-classifier AUROC as a function of transformer layer (x-axis), for four model families – Falcon3-7B (olive), Llama-3.1-8B (pink), Ministral-3-8B (purple), and Qwen3.5-9B (teal) – each shown as a solid line for the base model and a dashed line for its instruction-tuned counterpart, with a dot marking each curve's peak (best) layer. Every curve starts flat at chance level, marked by a dotted horizontal reference line, through the earliest layers, then rises sharply over a short span of layers before flattening into a peak and declining gradually and somewhat unevenly toward the final layers. The layer at which this rise begins, and the layer at which each curve peaks, differs noticeably across families: for Falcon3-7B, Llama-3.1-8B, and Ministral-3-8B, both the base and instruction-tuned variants rise early and peak within roughly the first third of their layers with the instruction-tuned variant peaking at higher AUROC values and remaining at higher values than the base model towards later layers. Qwen3.5-9B breaks this pattern: its instruction-tuned variant rises and peaks early like the other families, but its base variant instead stays flat at chance well past where every other curve has already peaked, rising only gradually and in an irregular, step-like fashion, and reaching its own peak at the latest layers.
Figure 4. Scatter plots of model-assigned answer probability P(answer∣prompt) against the linear classifier bias score P(biased) indicate a correlation between bias and confidence for the test set of BBQ prompts. The left two panels show ambiguous contexts while the right two panels show disambiguous contexts. The first and third panels (in pink) correspond to the biased prompts while the second and fourth panels (in teal) correspond to the anti-biased prompts. The black line shows a linear fit in each case. The corresponding Spearman rank correlation coefficient between the two quantities is given in Table 1 . Model answer probability correlates with bias score. Four scatter plots side by side, each plotting individual answers' bias score P(biased) (x-axis) against the model's assigned answer probability P(answer∣prompt) (y-axis), with a dashed vertical reference line at P(biased)=0.5 and a solid linear regression line fit through the points. The first two panels cover ambiguous-context answers – biased (pink) then anti-biased (teal) – and the last two cover disambiguous-context answers, again biased then anti-biased. In both ambiguous panels, points are scattered broadly across nearly the full range of answer probability regardless of bias score, and the regression line has a shallow positive slope, indicating a weak relationship between an answer's bias score and how likely the model actually is to give it. In both disambiguous panels, the point cloud instead forms a wedge shape: probabilities stay low and tightly clustered at low bias scores, then spread out and climb steadily as the bias score increases, and the regression line's slope is visibly much steeper than in the ambiguous panels. This shift from a weak to a strong relationship holds for both biased and anti-biased answers, showing that disambiguating the context substantially strengthens the link between the linear classifier's bias score and the model's actual answer probability.
BBQ (ambig.)
BBQ (disambig.)
StereoSet
CrowS-Pairs
Model
Biased
Anti-biased
Biased
Anti-biased
Biased
Anti-biased
Biased
Anti-biased
Llama-3.1-8B
0.21
0.22
0.64
0.62
0.53
0.52
0.47
0.49
Llama-3.1-8B-Instruct
0.27
0.10
0.77
0.80
0.54
0.64
0.36
0.48
Falcon3-7B-Base
0.10
0.21
0.57
0.59
0.19
0.29
0.28
0.47
Falcon3-7B-Instruct
0.10
0.23
0.73
0.76
0.33
0.38
0.29
0.31
Qwen3.5-9B-Base
0.44
0.45
0.61
0.59
0.41
0.44
0.38
0.46
Table 1. Correlations are found between the linear classifier bias score P(biased) and the model-assigned answer probability P(answer∣prompt) as measured by Spearman rank correlation coefficients rs , across models, for BBQ (ambiguous and disambiguated context conditions, as shown in Fig. 4 for the Llama-3.1-8B model) and for StereoSet and CrowS-Pairs (single-context-condition datasets with no ambiguous/disambiguated distinction), evaluated on the held-out test split. Columns correspond to combinations of dataset/context condition and answer type (biased, anti-biased).
Model
Train
BBQ (disambig.)
MMLU
Open BookQA
Model
Train
BBQ (disambig.)
MMLU
Open BookQA
Llama-3.1-8B
BBQ (ambig.)
0.81
0.71
0.79
Qwen3.5-9B- Base
BBQ (ambig.)
0.94
0.80
0.87
StereoSet
0.33
0.72
0.81
StereoSet
0.62
0.60
0.73
CrowS-Pairs
0.80
0.88
0.90
CrowS-Pairs
0.70
0.61
0.74
Llama-3.1-8B- Instruct
BBQ (ambig.)
0.94
0.72
0.88
Qwen3.5-9B
BBQ (ambig.)
0.95
0.63
0.72
StereoSet
0.61
0.78
0.83
StereoSet
0.80
0.63
0.76
CrowS-Pairs
0.35
0.80
0.88
CrowS-Pairs
0.74
0.81
0.88
Table 2. A linear classifier trained on biased vs. anti-biased prompts achieves a high AUROC when evaluated on the highest vs. lowest probability prompts, even for datasets which do not contain societal bias (MMLU and OpenBookQA). AUROC for highest vs. lowest probability answer embeddings across evaluation datasets (columns), using linear classifiers trained on biased/anti-biased prompts from different training datasets (rows).
Figure 5. Steering along the debiasing direction does not correct bias in a targeted way: it instead reduces model confidence, causing an increase in abstention where an abstain option is available, and an equilibration of the remaining answer probabilities where it is not. Figure shows the mean probability across prompts split according to answer type; biased (pink), anti-biased (teal), neutral (or abstain; grey) and unrelated (only for StereoSet; olive). The columns show results for the different datasets: BBQ (ambiguous), BBQ (disambiguous), CrowS-Pairs and StereoSet from left to right. The top two rows show the datasets with an option to abstain while the bottom two rows show the same datasets without an option to abstain. Thick borders indicate datasets in their original format. The solid lines in the first and third rows show results for Llama-3.1-8B, while the dashed lines in the second and fourth rows show results for Llama-3.1-8B-Instruct. The star marker shows the natural magnitude of the CAA steering vector. Steering in the debiasing direction raises abstention rates when allowed, and equilibrates answer probabilities when abstaining is not an option. A 4x4 grid of line plots showing mean answer probability as a function of steering strength α (x-axis, from −1 to 2, with a dashed vertical line at α=0), for four datasets in columns – BBQ (ambig.), BBQ (disambig.), CrowS-Pairs, and StereoSet – and, in rows, two settings stacked as blocks: an "Abstain option" block on top and a "No abstain option" block below, each block containing a base-model sub-row (solid lines) above an instruction-tuned sub-row (dashed lines). Lines are coloured by answer type – biased (pink), anti-biased (teal), neutral/abstain (gray, only in the abstain-option rows), and unrelated (olive, only in the StereoSet column) – and a star on each curve marks its value at that model/dataset's natural steering magnitude. Panels drawn with a heavier black border mark each dataset's native evaluation format: BBQ's native format includes the abstain option, so its two columns are bordered in the top block, while CrowS-Pairs and StereoSet have no native abstain option, so their columns are instead bordered in the bottom block. In the abstain-option rows, increasing α steadily raises the neutral/abstain curve while the biased and anti-biased curves decline together; this effect is mild for base models but dramatic for instruction-tuned models, where the abstain curve rises to dominate almost completely while the biased and anti-biased curves collapse toward each other and toward zero. In the no-abstain-option rows, the same steering instead redistributes probability directly between biased and anti-biased, with anti-biased rising as biased falls; again the shift is stronger for instruction-tuned models. The StereoSet unrelated curve stays low and comparatively flat in every row, rising slightly in the no-abstain variant of the dataset.
BBQ
StereoSet
CrowS-Pairs
Model
MMLU
OBQA
MMLU
OBQA
MMLU
OBQA
Llama-3.1-8B
0.53
0.56
0.49
0.55
0.38
0.39
Llama-3.1-8B-Instruct
0.78
0.74
0.75
0.74
0.57
0.56
Falcon3-7B-Base
0.58
0.61
0.36
0.51
0.30
0.37
Falcon3-7B-Instruct
0.73
0.76
0.47
0.61
0.29
0.35
Qwen3.5-9B-Base
0.91
0.81
0.68
0.68
0.61
0.74
Table 3. The linear directions in activation space representing model confidence and bias are found to be highly aligned. Cosine similarity between the CAA debiasing direction (Eq. ( 4 ), computed from BBQ (ambig.), CrowS-Pairs, or StereoSet at that dataset’s best-AUROC layer) and the confidence direction computed out-of-distribution, from MMLU and OpenBookQA at the same layer. Bold marks the highest value per column.
Acc.
Bias
Ent.
Acc.
Bias
Ent.
Model
Dataset
0
α~
0
α~
0
α~
Model
Dataset
0
α~
0
α~
0
α~
Llama-3.1-8B
BBQ (amb.)
0.13
0.25 ↑
0.12
0.10 ↓
0.97
1.03
Qwen3.5-9B- Base
BBQ (amb.)
0.56
0.56
0.18
0.18
0.91
0.91
BBQ (dis.)
0.85
0.82 ↓
0.03
0.03
0.75
0.91
BBQ (dis.)
0.94
0.94
0.01
0.01
0.37
0.37
StereoSet
–
–
70.09
62.66 ↓
0.66
0.77 ↑
StereoSet
–
–
72.31
72.47
0.57
0.57
CrowS-Pairs
–
–
75.82
64.71 ↓
0.52
0.59 ↑
CrowS-Pairs
–
–
74.51
73.20
0.53
0.53
Llama-3.1- 8B-Instruct
BBQ (amb.)
0.76
0.99 ↑
0.08
0.01 ↓
0.49
0.29
Qwen3.5-9B
BBQ (amb.)
0.85
0.86 ↑
0.07
0.06 ↓
0.53
0.53
Table 4. Accuracy / Bias / Entropy, α=0 vs. α~ , native dataset format, where α~ is the natural CAA magnitude αnat when αnat>1 , and α=1 otherwise. Bias is samb / sdis for BBQ but a bias-answer percentage (0–100) for CrowS-Pairs/StereoSet. Note that the two are not on the same scale. ↑ / ↓ on the α~ column mark a consistent direction of change across (nearly) all eight models; unmarked cells show no consistent trend or are flat.
Baseline
BBQ (ambig.)
StereoSet
CrowS-Pairs
MMLU
OBQA
MMLU
OBQA
MMLU
OBQA
MMLU
OBQA
Model
Acc
Ent
Acc
Ent
Acc ↓
Ent ↑
Acc ↓
Ent ↑
Acc ↓
Ent ↑
Acc ↓
Ent ↑
Acc ↓
Ent ↑
Acc ↓
Ent ↑
Llama-3.1-8B
0.64
0.86
0.75
0.78
0.63
0.97
0.74
0.92
0.63
0.95
0.72
0.91
0.62
0.94
0.71
0.91
Llama-3.1-8B-Instruct
0.63
0.55
0.76
0.33
0.59
0.80
0.70
0.57
0.60
0.74
0.72
0.51
0.61
0.73
0.73
0.49
Falcon3-7B-Base
0.67
0.70
0.74
0.69
0.68
0.71
0.74
0.71
0.67
0.75
0.73
0.75
0.65
0.80
0.70
0.79
Falcon3-7B-Instruct
0.69
0.29
0.80
0.19
0.69
0.30
0.80
0.20
0.69
0.31
0.79
0.21
0.68
0.34
0.79
0.23
Table 5. MMLU / OpenBookQA accuracy and entropy: baseline ( α=0 , direction-independent) vs. α~ (the natural CAA magnitude αnat when αnat>1 , and α=1 otherwise) for each of the three classifier-trained steering directions. ↓ / ↑ on the Acc/Ent headers mark the consistent direction of change from baseline across (nearly) all eight models and all three directions.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6. Example BBQ prompt (ambiguous context). Options are shuffled per question; A / B / C do not consistently map to a fixed answer type. A BBQ prompt example.
Figure 7. Example StereoSet prompt (intrasentence). Options are shuffled per question; A / B / C do not consistently map to a fixed answer type. Like CrowS-Pairs, there is a single context condition; unlike CrowS-Pairs, a neutral (unrelated) option is present. A StereoSet prompt example.
Figure 8. Example CrowS-Pairs prompt. Options are shuffled per question; A / B do not consistently map to a fixed answer type. CrowS-Pairs has no ambiguous/disambiguated distinction and no neutral option. A CrowS-Pairs prompt example.
BBQ (ambig.)
StereoSet
CrowS-Pairs
Model
Layer
AUROC
Layer
AUROC
Layer
AUROC
Llama-3.1-8B
13
0.70
13
0.92
14
0.94
Llama-3.1-8B-Instruct
14
0.75
14
0.91
14
0.94
Falcon3-7B-Base
17
0.67
15
0.87
16
0.90
Falcon3-7B-Instruct
17
0.70
15
0.90
16
0.92
Qwen3.5-9B-Base
30
0.68
26
0.76
26
0.82
Appendix
Table 6. Highest AUROC value and the corresponding layer number for the linear classifier trained to separate biased from anti-biased hidden states, for each model and each of the three bias datasets (BBQ ambiguous, StereoSet and CrowS-Pairs). (Sec. 4 ); Fig. 3 shows the full per-layer AUROC curve for the BBQ-trained classifiers.
Model
BBQ (ambig.)
StereoSet
CrowS-Pairs
Llama-3.1-8B
0.10
0.66
1.09
Llama-3.1-8B-Instruct
0.18
0.79
0.90
Falcon3-7B-Base
2.19
11.89
20.27
Falcon3-7B-Instruct
2.87
10.20
12.08
Qwen3.5-9B-Base
3.61
3.14
3.19
Qwen3.5-9B
0.97
4.56
3.49
Appendix
Table 7. Natural alpha (unscaled CAA direction magnitude) per model / dataset.
BBQ (amb.)
BBQ (dis.)
StereoSet
StereoSet (unrel.)
CrowS-Pairs
Model
0
α~
0
α~
0
α~
0
α~
0
α~
Llama-3.1-8B
0.61
0.87
0.30
0.53
1.29
1.40
0.28
0.32
0.86
0.99
Llama-3.1-8B-Instruct
5.51
21.90
0.22
1.82
0.60
2.50
0.19
0.26
0.98
2.24
Falcon3-7B-Base
1.17
1.25
0.39
0.43
0.77
0.92
0.29
0.31
1.02
1.33
Falcon3-7B-Instruct
7.83
8.78
0.18
0.20
1.36
1.78
0.10
0.11
2.49
3.37
Qwen3.5-9B-Base
1.90
1.89
0.21
0.21
1.38
1.37
0.28
0.28
1.35
1.35
Appendix
Table 8. When the model is allowed to abstain, steering increases the probability mass concentrated on the abstain option. Ratio of the mean probability of the neutral (abstain) answer - and, for StereoSet, also the unrelated answer - to that of the anti-biased answer, by dataset and model, in the abstain-option format. α~ is the natural CAA magnitude αnat when αnat>1 , and α=1 otherwise.
BBQ (amb.)
BBQ (dis.)
StereoSet
CrowS-Pairs
Model
0
α~
0
α~
0
α~
0
α~
Llama-3.1-8B
1.13
1.09
1.07
1.05
1.95
1.50
1.73
1.21
Llama-3.1-8B-Instruct
1.58
1.38
1.05
1.07
2.45
2.13
1.94
1.22
Falcon3-7B-Base
1.16
1.15
1.07
1.07
1.82
1.53
1.97
1.32
Falcon3-7B-Instruct
1.85
1.86
1.04
1.04
2.90
2.70
2.94
2.23
Qwen3.5-9B-Base
1.49
1.49
1.04
1.04
2.51
2.51
2.02
2.02
Appendix
Table 9. When the model is allowed to abstain, steering increases the probability mass concentrated on the abstain option and, in most cases, reduces the probability of the biased and anti-biased options. Ratio of the mean probability of the biased answer to that of the anti-biased answer, by dataset and model, in the abstain-option format. α~ is the natural CAA magnitude αnat when αnat>1 , and α=1 otherwise. The ratio is either left unchanged, as in BBQ (dis.) or is reduced as in BBQ (ambig.), StereoSet and CrowS-Pairs.
BBQ (amb.)
BBQ (dis.)
StereoSet
StereoSet (unrel.)
CrowS-Pairs
Model
0
α~
0
α~
0
α~
0
α~
0
α~
Llama-3.1-8B
1.14
1.09
1.08
1.05
1.81
1.45
0.29
0.34
1.91
1.43
Llama-3.1-8B-Instruct
1.35
1.22
1.05
1.06
2.30
1.80
0.19
0.31
2.04
1.27
Falcon3-7B-Base
1.17
1.16
1.07
1.07
1.85
1.57
0.25
0.28
2.05
1.31
Falcon3-7B-Instruct
1.49
1.48
1.03
1.03
2.53
2.29
0.12
0.14
2.73
2.11
Qwen3.5-9B-Base
1.40
1.40
1.04
1.04
2.36
2.36
0.31
0.31
1.85
1.86
Appendix
Table 10. When the model is not allowed to abstain, steering equilibrates the biased and anti-biased probabilities, reducing their ratio towards a value of 1. Ratio of the mean probability of the biased answer - and, for StereoSet, also the unrelated answer - to that of the anti-biased answer, by dataset and model, in the no-abstain format. α~ is the natural CAA magnitude αnat when αnat>1 , and α=1 otherwise.