Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable
Organizations: Lossfunk
Abstract
Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap across five open-weight families and three closed APIs. A single verified-source note endorsing a wrong answer flips 45-88% of baseline-correct responses in seven of eight models, and compliance rises with how authoritative the note sounds. Source deference and user agreement are not behaviorally interchangeable inside the model: on matched items with the same wrong answer, causal interventions can selectively suppress one without equally affecting the other. In three open-weight families, removing a fitted source direction lowers source compliance by 65-80 percentage points while removing a user or assistant direction has far smaller effects, and removing the user direction shows the reverse preference. A separately fitted intervention derived from source-versus-user cue activations moves compliance in both directions while leaving the prompt text unchanged. An authority direction fitted on trivia also transfers to PIQA and multi-turn SYCON dialogues without refitting, and removing it lowers wrong-source compliance by tens of percentage points in four of five families with no detected change in MMLU-Pro or GSM8K accuracy at our evaluation sizes. Source deference and user agreement therefore need separate evaluation.
Figures & tables
| Model | Extraction | Position | Layer |
|---|---|---|---|
| Qwen3.5-27B | 1808 | endorsed answer | L5 |
| GPT-OSS-20B | 791 | endorsement end | L18 |
| OLMo-3.1-32B | 311 | endorsement start | L5 |
| OLMo-2-32B | 275 | answer position | L16 |
| Gemma-4-26B-A4B | 842 | endorsement span | L22 |
| Open-weight | Closed API | |||||||
|---|---|---|---|---|---|---|---|---|
| Experiment | Qwen3.5 | GPT-OSS | OLMo-2 | OLMo-3.1 | Gemma-4 | GPT-5.4 | Grok-4.20 | Gemini |
| Behavioral override (Fig. 1 A) | ||||||||
| Source-vs-user (Fig. 1 B) | ||||||||
| Authority gradient | — | — | — | |||||
| Forward steering (trivia + PIQA) | inert | — | — | — | ||||
| Source/user causal | — | — | — | — | ||||
| Model | Layer | Matched flip | |
|---|---|---|---|
| OLMo-3.1-32B | L15 | 1.0 | 58.8% |
| GPT-OSS-20B | L16 | 1.0 | 32.7% |
| OLMo-2-32B | L16 | 0.5 | 19.0% |
| Qwen3.5-27B | L2 | 1.0 | 16.1% |
| Gemma-4-26B-A4B | L15 | 1.0 | 10.6% |
| Model (layer) | Removed vector | Source-wrong pp [95% CI] | User-wrong pp [95% CI] |
|---|---|---|---|
| Qwen3.5 (L5) | source | [ ] | [ ] |
| user | [ ] | [ ] | |
| assistant | [ ] | [ ] | |
| GPT-OSS (L16) | source | [ ] | [ ] |
| user | [ ] | [ ] | |
| assistant | [ ] | [ ] |
| Model (layer) | Baseline gap | Source user | User source |
|---|---|---|---|
| Qwen3.5 (L5) | 20.1 pp | [ ] | [ ] |
| GPT-OSS (L16) | 50.0 pp | [ ] | [ ] |
| OLMo-3.1 (L22) | 34.0 pp | [ ] | [ ] |
| Model | Capitulation rate | pp [95% CI] |
|---|---|---|
| GPT-OSS-20B | [ ] | |
| Qwen3.5-27B | [ ] | |
| OLMo-3.1-32B | [ ] | |
| OLMo-2-32B | [ ] | |
| Gemma-4-26B-A4B | [ ] |
| Model | Trivia | PIQA |
|---|---|---|
| Qwen3.5 | ||
| GPT-OSS | ||
| OLMo-3.1 | ||
| OLMo-2 | ||
| Gemma-4 |
| Model | Layer, multiplier | Wrong-source | Wrong-user |
|---|---|---|---|
| Qwen3.5 | L49, | ||
| GPT-OSS | L16, | ||
| OLMo-3.1 | L29, |
| Model | MMLU-Pro pp [95% CI] | GSM8K pp [95% CI] |
|---|---|---|
| Qwen3.5 | [ ] | [ ] |
| GPT-OSS | [ ] | [ ] |
| OLMo-3.1 | [ ] | [ ] |
| OLMo-2 | [ ] | [ ] |
| Gemma-4 | [ ] | [ ] |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Direction | Fitted from | Used for |
|---|---|---|
| Correct- and wrong-source prompts vs. neutral prompts (Eq. 1 ) | Steering, SYCON, PIQA transfer | |
| with its assistant-axis component removed | Authority removal (mitigation) | |
| Wrong-source prompts vs. neutral prompts | Source/user removal (Table 4 ) | |
| Wrong-user prompts vs. neutral prompts | Source/user removal (Table 4 ) | |
| Default-assistant vs. role-play activations [ 19 ] | Control | |
| Source-minus-user shift over the cue tokens, shared part removed (Eq. 3 ) | Attribution patch |
| Model | Fit | Steer | Causal | Affect | Mitigation | SYCON |
|---|---|---|---|---|---|---|
| Qwen3.5 | L5 | L2 | L5 | L5 | L5 | L10 |
| GPT-OSS | L18 | L16 | L16 | L18 | L16 | L12 |
| OLMo-2 | L16 | L16 | L22 | L16 | L10 | L22 |
| OLMo-3.1 | L5 | L15 | L22 | L5 | L15 | L22 |
| Gemma-4 | L22 | L15 | — | — | L20 | L24 |
| Family | Dataset | Layer | Reference MF% | Held-out MF% | pp | ||
|---|---|---|---|---|---|---|---|
| Qwen3.5 | Trivia | 200 | L5 | 0.5 | 32.7 | 31.0 | |
| Qwen3.5 | PIQA | 125 | L5 | 0.7 | 20.6 | 20.0 | |
| GPT-OSS | Trivia | 222 | L16 | 1.0 | 32.7 | 31.5 | |
| GPT-OSS | PIQA | 100 | L16 | 1.0 | 18.8 | 19.0 | |
| OLMo-2 | Trivia | 39 | L10 | 1.0 | 19.0 | 23.1 | |
| OLMo-2 | PIQA | 143 | L16 | 0.5 | 10.4 | 10.5 |
| Family | Removal | Source-cued (ref/held/ ) | User-cued (ref/held/ ) | ||||
|---|---|---|---|---|---|---|---|
| Qwen3.5 | source | 16.8 | 14.0 | 42.7 | 49.0 | ||
| Qwen3.5 | user | 76.5 | 78.0 | 16.9 | 18.5 | ||
| Qwen3.5 | assistant | 87.2 | 86.0 | 59.3 | 58.5 | ||
| GPT-OSS | source | 18.2 | 18.0 | 42.4 | 41.0 | ||
| GPT-OSS | user | 86.1 | 87.0 | 17.8 | 18.0 | ||
| GPT-OSS | assistant | 95.7 | 95.0 | 35.4 | 35.0 | ||
| Family | Variant | Trivia (ref/held/ ) | PIQA (ref/held/ ) | ||||
|---|---|---|---|---|---|---|---|
| Qwen3.5 | authority | 5.2 | 6.6 | 4.5 | 6.0 | ||
| Qwen3.5 | assistant | 43.0 | 43.4 | 43.4 | 44.0 | ||
| GPT-OSS | authority | 43.2 | 44.0 | 6.2 | 7.0 | ||
| GPT-OSS | assistant | 57.8 | 57.0 | 57.5 | 58.0 | ||
| OLMo-2 | authority | 29.5 | 31.9 | 28.5 | 30.0 | ||
| OLMo-2 | assistant | 36.8 | 37.5 | 35.2 | 36.0 | ||
| Model | Open | N parse% | N wrong% | W parse% | W wrong% | Flip rate |
|---|---|---|---|---|---|---|
| Qwen3.5-27B | yes | 99.9 | 25.1 | 99.9 | 50.3 | 44.9% |
| GPT-OSS-20B | yes | 98.5 | 22.9 | 98.0 | 71.9 | 65.4% |
| OLMo-2-32B | yes | 91.9 | 18.2 | 91.6 | 66.0 | 82.8% |
| OLMo-3.1-32B | yes | 95.0 | 53.2 | 97.4 | 90.5 | 82.9% |
| Gemma-4-26B-A4B | yes | 97.2 | 10.0 | 97.8 | 66.8 | 62.9% |
| GPT-5.4 | no | 98.5 | 7.9 | 98.5 | 49.4 | 44.7% |
| Model | Layer | VA-plane fraction | ||
|---|---|---|---|---|
| Qwen3.5-27B | L5 | |||
| GPT-OSS-20B | L18 | |||
| OLMo-2-32B | L16 | |||
| OLMo-3.1-32B | L15 | |||
| Gemma-4-26B-A4B | L15 |
| Intervention | Qwen3.5 (L5) | GPT-OSS (L18) | OLMo-2 (L16) | OLMo-3.1 (L5) |
|---|---|---|---|---|
| None (baseline) | ||||
| Affect plane only | ||||
| Authority affect | ||||
| Authority affect |
| Model | Layer | Baseline | Source removal | Without assistant component | |
|---|---|---|---|---|---|
| Qwen3.5 | L5 | 244 | 86.5% | 13.9% | 21.3% |
| GPT-OSS | L16 | 244 | 95.9% | 18.4% | 26.6% |
| OLMo-3.1 | L22 | 71 | 88.7% | 23.9% | 36.6% |
| OLMo-2 | L22 | 40 | 87.5% | 30.0% | 35.0% |
| Model | Layer | Source user | User source |
|---|---|---|---|
| Qwen3.5 | L5 | ||
| GPT-OSS | L16 | ||
| OLMo-3.1 | L22 |
| Model | Layers | Source user | User source |
|---|---|---|---|
| Qwen3.5 | L3 / L4 / L5 / L6 / L7 | ||
| GPT-OSS | L14 / L15 / L16 / L17 / L18 | ||
| OLMo-3.1 | L20 / L21 / L22 / L23 / L24 |
| Model | Wording | Baseline gap | Source user | Gap closed | User source |
|---|---|---|---|---|---|
| Qwen3.5 | Direct | 34.8 | 54.9% | ||
| Reported speech | 20.1 | 54.7% | |||
| Structured fields | 15.6 | 55.1% | |||
| GPT-OSS | Direct | 86.5 | 60.6% | ||
| Reported speech | 50.0 | 61.0% | |||
| Structured fields | 38.9 | 59.9% |
| Model | Cue position | Baseline gap | Source user | User source | Gap closed |
|---|---|---|---|---|---|
| Qwen3.5 | After options | 20.1 | 55% | ||
| Before question | 19.0 | 55% | |||
| GPT-OSS | After options | 50.0 | 61% | ||
| Before question | 47.8 | 60% | |||
| OLMo-3.1 | After options | 34.0 | 57% | ||
| Before question | 32.3 | 56% |
| Model | Layer | |||
|---|---|---|---|---|
| Qwen3.5 | L5 | |||
| GPT-OSS | L16 | |||
| GPT-OSS | L18 | |||
| OLMo-2 | L10 | |||
| OLMo-2 | L16 | |||
| OLMo-2 | L22 |
| Model | L | flip 0 % | flip 0.3 % | flip 0.5 % | traj 0 | traj 0.3 | traj 0.5 |
|---|---|---|---|---|---|---|---|
| Qwen3.5 | 2 | 43.5 | 50.6 | 40.7 | 3.20 | 2.97 | 3.33 |
| Qwen3.5 | 5 | 44.7 | 47.0 | 31.0 | 3.20 | 3.12 | 3.41 |
| Qwen3.5 | 10 | 44.7 | 47.1 | 63.3 | 3.18 | 3.16 | 2.48 |
| GPT-OSS | 12 | 38.8 | 24.5 | 91.4 | 1.95 | 2.07 | 0.73 |
| GPT-OSS | 16 | 38.5 | 45.8 | 66.7 | 2.09 | 1.85 | 0.97 |
| GPT-OSS | 18 | 34.7 | 44.7 | 65.2 | 2.01 | 1.75 | 1.47 |
| Run | agree | agree% | Cohen | DS=1% | Gemini=1% | |
|---|---|---|---|---|---|---|
| Qwen3.5 | 4498 | 3838 | 85.3% | 0.61 | 79.4% | 71.6% |
| GPT-OSS | 6000 | 5014 | 83.6% | 0.67 | 52.5% | 55.6% |
| Gemma-4 | 4500 | 3696 | 82.1% | 0.56 | 77.1% | 67.0% |
| OLMo-2 | 4500 | 3925 | 87.2% | 0.72 | 65.4% | 64.0% |
| OLMo-3.1 | 4500 | 3920 | 87.1% | 0.67 | 75.9% | 71.6% |
| Model | Cell (L, ) | DS baseline flip | DS pp [95% CI] | Gemini pp |
|---|---|---|---|---|
| GPT-OSS-20B | L12, 0.5 | [ ] | ||
| Qwen3.5-27B | L10, 0.5 | [ ] | ||
| OLMo-3.1-32B | L22, 0.5 | [ ] | ||
| OLMo-2-32B | L22, 0.5 | [ ] | ||
| Gemma-4-26B-A4B | L24, 0.5 | [ ] |
| Model | Cue surface | baseline | source | user | assistant |
|---|---|---|---|---|---|
| Qwen3.5-27B | inserted note | ||||
| system prompt | |||||
| RAG document | |||||
| GPT-OSS-20B | inserted note | ||||
| system prompt | |||||
| RAG document |
| Model | Layer | MMLU-Pro | GSM8K | ||
|---|---|---|---|---|---|
| Qwen3.5-27B (auth) | L5 | 60.7 59.6 | 46.5 42.5 | ||
| Qwen3.5-27B (asst) | L5 | 60.7 60.7 | 46.5 46.5 | ||
| GPT-OSS-20B (auth) | L16 | 52.4 50.5 | 88.5 89.0 | ||
| OLMo-2-32B (auth) | L10 | 34.7 34.7 | 12.0 12.5 | ||
| OLMo-2-32B (asst) | L10 | 34.7 34.0 | 12.0 11.5 | ||
| OLMo-3.1-32B (auth) | L15 | 44.3 44.7 | 26.5 26.0 |
| Model | Task | Authority removal pp [95% CI] |
|---|---|---|
| Qwen3.5 | Trivia | [ ] |
| Qwen3.5 | PIQA | [ ] |
| OLMo-2 | Trivia | [ ] |
| OLMo-2 | PIQA | [ ] |