Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it. We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capital of Spain?"). To understand and address this brittleness, we mechanistically examine how models operate under negation. Our main finding is that specialized attention heads and MLP neurons jointly implement negation by (1) suppressing retrieval of the original answer (e.g., "Madrid") while (2) promoting a favored candidate within the answer category (e.g., "Paris"). This contrasts with accounts of human negation processing, in which information about the original answer helps to determine what should be excluded. Furthermore, we find that this difference from human processing is a key source of negation failures: the model's mechanism relies on suppressing the original answer rather than using it to determine what to exclude, so the model can repeat the original answer when suppression is too weak or when a bias toward particular answers prevents it from selecting an alternative. To address this weakness in the model's negation mechanism, we propose a training objective that requires larger shifts in answer preference for more confident original predictions, and show that it reduces negation failures with less degradation of general capabilities than standard fine-tuning baselines. Together, our results demonstrate how mechanistic analysis can reveal why a linguistic capability fails and guide training that targets the underlying limitation.
Figures & tables
Figure 1: Overview of LLMs’ negation mechanism. Left: The negation signal triggers three downstream changes: (A) suppressing the original answer, (B) promoting answer category information, and (C) promoting a category-specific alternative. Right: These changes preserve the answer category subspace while shifting preference to an alternative answer (“Paris”).
Factual Association
Logical Reasoning
Visual Question Answering
Average
PopQA
RippleEdits
PhantomWiki
SynthWorlds
GQA
PTR
Negated
Model
Ans. change / Cat. pres.
Ans. change / Cat. pres.
Ans. change / Cat. pres.
Ans. change / Cat. pres.
Ans. change / Cat. pres.
Ans. change / Cat. pres.
Ans. change / Cat. pres.
Original acc.
Open-source models
Gemma 3-4B-IT
40.8 / 98.5
14.1 / 95.7
66.4 / 90.8
35.0 / 90.3
46.3 / 93.0
31.8 / 97.3
39.1 / 94.3
32.6
Gemma 3-12B-IT
58.1 / 99.7
31.1 / 95.9
50.4 / 96.2
42.3 / 95.2
58.3 / 94.9
44.4 / 99.1
47.4 / 96.8
41.9
Gemma 3-27B-IT
68.2 / 99.9
44.8 / 98.8
70.1 / 99.6
42.9 / 95.4
77.5 / 94.1
72.5 / 98.9
62.6 / 97.8
44.8
Table 1: Answer change and category preservation rates on the test split of our benchmark (%). Averages are computed over six benchmarks for multimodal models and four for text-only models.
Figure 2: Left: Answer change rates decrease with confidence in the original answer (mean token log-probability; results averaged across six open-source models). Right: Answer change rates are also lower when the original answer is the model’s most common response under negation.
Figure 3: Removing the negation direction in Gemma 3-12B-IT. (a) Layerwise effects of steering with the negation signal ( β=1 ), overlaid with negation-delta similarity. (b) Causal effects of steering with the negation signal at (L24). (c) Cosine similarity between attention and MLP output changes induced by natural negation and by steering ( β=3 ). Effects are averaged over six benchmarks. Shading indicates 95% bootstrap intervals.
Figure 4: Components involved in forming alternatives in Gemma 3-12B-IT. (a) Cumulative head and neuron scores. (b) Effects of restoring their non-negated values, averaged across six benchmarks; dots show matched random controls. (c) Unfiltered logit lens top/bottom five for an illustrative example.
Figure 5: Negation largely preserves the shared answer category structure (a–b), while redirecting candidate-space representations toward a specific direction (c–d). Arrows show independently computed activations without centering or normalization.
In-domain (ID)
Out-of-domain (OOD)
General Capability
Method
Answer change
Category preservation
Answer diversity
Original accuracy
Answer change
Category preservation
Answer diversity
Original accuracy
Avg.
Base
58.14
99.71
90.90
21.90
46.93
96.64
80.05
37.15
78.04
ALiT (ours)
96.55
99.86
86.54
22.10
84.96
97.63
76.26
37.40
77.68
Unlikelihood
99.95 ( +3.40 )
53.97 ( −45.89 )
75.21 ( −11.33 )
21.16 ( −0.94 )
89.41 ( +4.45 )
57.68 ( −39.95 )
84.46 ( +8.20 )
37.40 ( 0.00 )
77.95 ( +0.27 )
SFT
96.84 ( +0.29 )
100.00 ( +0.14 )
8.03 ( −78.51 )
21.64 ( −0.46 )
86.52 ( +1.56 )
99.06 ( +1.43 )
39.11 ( −37.15 )
36.64 ( −0.76 )
76.98 ( −0.70 )
DPO
100.00 ( +3.45 )
42.76 ( −57.10 )
73.61 ( −12.93 )
21.97 ( −0.13 )
99.23 ( +14.27 )
52.63 ( −45.00 )
83.71 ( +7.45 )
37.16 ( −0.24 )
77.36 ( −0.32 )
Table 2: Training results for Gemma 3-12B-IT (%). ID denotes PopQA, and OOD denotes five other benchmarks.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Form
Prompt
PopQA Original answer: New Delhi
Original
What is the capital of India?
not
What is not the capital of India?
n’t
What isn’t the capital of India?
never
What is never the capital of India?
by no means
What is by no means the capital of India?
Appendix
Table 3: Examples of all four negation forms from each source benchmark, drawn from the training split. Bold text marks the negation expression. Original answers are source references, not targets for the negated prompts. Underlines indicate completion slots for display. Supporting documents are omitted; the corresponding GQA and PTR images appear in Figure 6 .
Figure 6: Images accompanying the visual examples in Table 3 . The same image is used for the original prompt and all four negated variants.
Validation Criterion
PASS
FAIL
FAIL Rate (%)
Grammaticality
125,527
1,229
0.97
Unambiguity
120,329
6,427
5.07
Negation Scope Fidelity
105,330
21,426
16.90
Overall
105,048
21,708
17.13
Appendix
Table 4: Benchmark validation results for the negation benchmark dataset. Overall rejection indicates that a record was rejected if it failed at least one validation criterion.
Benchmark
Original
PASS
PASS Rate (%)
PopQA
45,476
37,426
82.30
RippleEdits
18,488
14,150
76.54
PhantomWiki
6,660
5,668
85.11
SynthWorlds
1,120
840
75.00
GQA
21,112
15,912
75.37
PTR
33,900
31,052
91.60
Appendix
Table 5: Benchmark validation results by source benchmark. Original denotes the number of examples in the original benchmark, while PASS denotes the number of examples that passed our filtering process. We also report the corresponding pass rate.
Validation Criterion
Unanimous PASS
Unanimous FAIL
Unanimous Agreement (%)
Grammaticality
237 (98.75%)
2 (0.83%)
99.58%
Unambiguity
237 (98.75%)
0 (0.00%)
98.75%
Negation Scope Fidelity
174 (72.50%)
11 (4.58%)
77.08%
Overall
648 (90.00%)
13 (1.81%)
91.81%
Appendix
Table 6: Human–human agreement for benchmark validation by validation criterion. Percentages are calculated over all criterion-level decisions within each validation criterion.
Validation Criterion
Agreed PASS
Agreed FAIL
Human–LLM Agreement (%)
Grammaticality
230 (95.83%)
2 (0.83%)
96.67%
Unambiguity
217 (90.42%)
0 (0.00%)
90.42%
Negation Scope Fidelity
185 (77.08%)
18 (7.50%)
84.58%
Overall
632 (87.78%)
20 (2.78%)
90.56%
Appendix
Table 7: Human–LLM agreement for benchmark validation by validation criterion. Agreed PASS and Agreed FAIL indicate decisions for which the human majority label and the LLM label were identical. Percentages are calculated over all decisions within each criterion.
Evaluation Criterion
Unanimous PASS
Unanimous FAIL
Unanimous Agreement (%)
Original Answer Correctness
91 (37.92%)
135 (56.25%)
94.17%
Repetition Under Negation
95 (39.58%)
132 (55.00%)
94.58%
Negated Answer Category Validity
179 (74.58%)
23 (9.58%)
84.17%
Overall
365 (50.69%)
290 (40.28%)
90.97%
Appendix
Table 8: Human–human agreement for negation evaluation by evaluation criterion. Percentages are calculated over all criterion-level decisions within each evaluation criterion.
Evaluation Criterion
Agreed PASS
Agreed FAIL
Human–LLM Agreement (%)
Original Answer Correctness
96 (40.00%)
137 (57.08%)
97.08%
Repetition Under Negation
62 (60.78%)
38 (37.25%)
98.04%
Negated Answer Category Validity
33 (84.62%)
0 (0.00%)
84.62%
Overall
191 (50.13%)
175 (45.93%)
96.06%
Appendix
Table 9: Human–LLM agreement for negation evaluation by evaluation criterion. Agreed PASS and Agreed FAIL indicate decisions for which the human majority label and the LLM label were identical. Percentages are calculated over non-null human–LLM decision pairs within each criterion.
Figure 7: Answer change by confidence percentile for each model on the test split. Confidence is measured by mean token log-probability of the original answer, with percentiles computed within each model. We include only originally correct cases and average negation forms within each question. Error bars show 95% question-bootstrap intervals.
Figure 8: Answer change by model and benchmark on the test split. The most frequent negated answer is identified within each model’s relation/type group using its test outputs. Rates are computed over originally correct pairs with a judged repetition outcome. Not evaluated indicates unavailable model–benchmark results.
Figure 9: Information flow in Gemma 3-12B-IT. Left: from negation tokens to other token groups and Right: from each group to the final prompt token.
Figure 10: Information flow in Gemma 3-4B-IT (top) and Llama 3.1-8B-Instruct (bottom). Left: paths from negation tokens. Right: paths to the final prompt token. Curves show normalized log-probability gap shift; shading shows 95% question-bootstrap intervals.
Figure 11: Information flow in Qwen3.5-9B, including softmax and linear attention. Final-position interventions continue during generation. Effects equally weight six benchmarks; shading shows 95% question-bootstrap intervals.
Model
Pairs
Layer
Match β=0
Match β=1
Gap shift β=1
Gemma 3-4B-IT
925
L18
13.8
40.1 [34.5, 46.4]
22.1 [18.3, 26.0]
Gemma 3-12B-IT
1540
L24
4.6
41.8 [37.2, 46.9]
24.6 [21.5, 28.0]
Llama 3.1-8B-Instruct
423
L24
19.6
19.7 [12.3, 27.8]
-1.2 [-4.9, 2.5]
Qwen3.5-9B
2304
L24
1.1
11.6 [8.2, 15.4]
0.5 [-0.7, 1.6]
Appendix
Table 10: Prompt-only negation-direction removal. Match denotes original-answer match rate; gap shift denotes normalized log-probability gap shift, both in percent. Brackets are 95% question-bootstrap intervals. Results equally weight six benchmarks for Gemma and Qwen, and four text benchmarks for Llama. The fixed dose layers are not uniformly the layerwise maxima; see the full sweeps.
Figure 12: Negation-direction removal in Gemma 3-4B-IT (top) and Gemma 3-12B-IT (bottom). Dose layers are L18 and L24; downstream alignment uses β=3 and β=2 , respectively. Effects equally weight six benchmarks. Shading shows available 95% question-bootstrap intervals; Gemma 12B downstream cosine is shown without a pooled interval.
Figure 13: Negation-direction removal in Llama 3.1-8B-Instruct (top; four text benchmarks) and Qwen3.5-9B (bottom; six benchmarks). Both displayed dose experiments use L24. Downstream alignment uses β=1 and β=2 , respectively.
Figure 14: Prompt-only negation-direction removal at L15 in Llama 3.1-8B-Instruct.
Intervention
Δs(a)
Δs(b)
Normalized gap shift (%)
Restore L31H13 answer-token contribution
+0.129
−0.028
0.8
Restore five head outputs
+3.438
−2.336
21.4
Restore 256 neuron activations
+4.053
−2.467
24.7
Restore both sets
+4.937
−4.548
35.2
Double 256 neuron differences
−8.259
−1.057
−28.2
Appendix
Table 11: Component interventions in Gemma 3-12B-IT, averaged across six benchmarks. Δs(a) and Δs(b) are changes in the original and alternative answers’ log-probabilities (nats); gap shift is the normalized log-probability gap shift toward the original prompt.
Model
Restored components
Match (%)
Gap shift (%)
Random gap shift (%)
Gemma 3-4B-IT
Five heads
52.4 [46.1, 58.7]
29.5 [26.2, 32.8]
12.0
Gemma 3-4B-IT
256 neurons
41.8 [36.6, 47.8]
22.0 [19.2, 24.9]
-0.3
Gemma 3-4B-IT
Both
59.6 [52.8, 66.3]
36.5 [33.3, 39.8]
12.0
Gemma 3-12B-IT
Five heads
37.4 [32.8, 43.0]
21.4 [18.5, 24.4]
2.1
Gemma 3-12B-IT
256 neurons
43.6 [39.0, 48.5]
24.7 [22.4, 26.8]
0.1
Gemma 3-12B-IT
Both
56.3 [50.7, 62.0]
35.2 [32.5, 38.0]
2.2
Appendix
Table 12: Prompt-only restoration of non-negated component values in negated prompts. Metrics and intervals follow Table 10 .
Model
Layer
Pairs
Match (%) Removal → restore
Gap shift (%) Removal → restore
Match difference (pp) Selected − random
Gemma 3-4B-IT
L18
290
50.4 → 34.3
29.4 → 11.6
-12.3 [-19.3, -6.8]
Gemma 3-12B-IT
L24
612
56.6 → 17.4
32.3 → 9.2
-39.3 [-44.5, -34.1]
Llama 3.1-8B-Instruct
L14
368
21.3 → 13.6
21.0 → 9.3
-7.0 [-11.3, -3.4]
Qwen3.5-9B
L16
642
11.1 → 2.6
5.0 → 1.0
-7.7 [-10.0, -5.4]
Appendix
Table 13: Signal removal followed by joint restoration of selected heads and neurons at their natural negated values ( β=1 ).
Figure 15: Signal removal followed by restoration of natural negated component values ( β=1 ). Rows show (a) Gemma 3-4B-IT, (b) Gemma 3-12B-IT, (c) Llama 3.1-8B-Instruct, and (d) Qwen3.5-9B.
Model
Category
Layer
Pairs
Learned
Random
Full
Gemma 3-4B-IT
City
L23
64
29.9
0.1
35.0
Gemma 3-4B-IT
Person Name
L27
583
30.1
0.2
91.9
Gemma 3-4B-IT
Religion
L26
33
36.9
0.1
87.9
Gemma 3-4B-IT
Sport
L31
40
38.7
0.2
91.1
Gemma 3-12B-IT
City
L38
64
32.2
0.1
69.8
Gemma 3-12B-IT
Person Name
L42
583
32.4
0.2
74.2
Appendix
Table 14: Rank-4 candidate-space interventions between non-negated PopQA prompts with different answers in the same category.
Model
Layer
Pairs
Learned
Random
Full
Within-category answer gap
Gemma 3-4B-IT
L19
720
80.1 [69.9, 89.1]
0.2
70.3
4.2
Gemma 3-12B-IT
L27
720
74.3 [32.5, 104.2]
-1.7
108.8
4.6
Llama 3.1-8B-Instruct
L15
720
141.6 [58.0, 280.0]
0.4
62.2
3.8
Qwen3.5-9B
L19
720
103.5 [96.7, 111.6]
0.0
67.9
1.0
Appendix
Table 15: Rank-4 answer-category-space interventions between non-negated PopQA prompts.
Model
Pairs/ questions
Heads
Neurons
Both
Both − random (pp)
Gemma 3-4B-IT
604/216
84.2
75.8
66.7
-30.2 [-34.9, -25.5]
Gemma 3-12B-IT
606/268
94.9
58.6
57.6
-40.9 [-45.7, -36.5]
Llama 3.1-8B-Instruct
473/202
97.2
63.7
64.0
-33.1 [-38.8, -27.7]
Qwen3.5-9B
527/206
91.7
60.0
54.3
-43.7 [-49.4, -38.4]
Appendix
Table 16: Original-answer match rate (%) on current-baseline repetition failures.
Figure 16: Interventions on repeated-answer failures in (a) Gemma 3-4B-IT, (b) Gemma 3-12B-IT, (c) Llama 3.1-8B-Instruct, and (d) Qwen3.5-9B. Left: original-answer match rate. Right: change in the original answer’s full-sequence log-probability relative to its unmodified negated baseline. Interventions continue during generation; error bars show 95% question-bootstrap intervals.
Model
Pairs
Success
Failure
Failure − success
Gemma 3-4B-IT
966
1.187
0.276
-0.910 [-1.675, -0.241]
Gemma 3-12B-IT
1312
0.084
0.047
-0.038 [-0.124, 0.053]
Llama 3.1-8B-Instruct
1381
0.079
0.028
-0.051 [-0.120, 0.015]
Qwen3.5-9B
1206
0.148
0.085
-0.064 [-0.189, 0.031]
Appendix
Table 17: Support for the original answer supplied by the selected answer-token paths in prompts without negation.
Method
MMLU-Pro
HellaSwag
GSM8K
IFEval
General Avg.
Base
58.25
83.78
89.92
80.22
78.04
ALiT (ours)
57.65
83.44
89.23
80.41
77.68
Unlikelihood
57.46
83.23
89.61
81.52
77.95
SFT
56.06
82.95
88.86
80.04
76.98
DPO
56.53
83.42
89.08
80.41
77.36
Appendix
Table 18: general-capability results for Gemma-3-12B-IT checkpoints (%).
Figure 17: Answer change for Base, ALiT, and SFT, grouped by whether the original answer matches the most frequent negated answer. Arrows show percentage-point changes from Base within each group.