cs.CLOct 7, 2026

How Do LLMs Change Predictions Under Negation?

Authors: Jongwook Yoon, Jongwon Lim, Sungjib Lim, Woojin Cho, Yohan Jo

Organizations: Graduate School of Data Science, Seoul National University · Department of Computer Science and Engineering, Seoul National University

Abstract

Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it. We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capital of Spain?"). To understand and address this brittleness, we mechanistically examine how models operate under negation. Our main finding is that specialized attention heads and MLP neurons jointly implement negation by (1) suppressing retrieval of the original answer (e.g., "Madrid") while (2) promoting a favored candidate within the answer category (e.g., "Paris"). This contrasts with accounts of human negation processing, in which information about the original answer helps to determine what should be excluded. Furthermore, we find that this difference from human processing is a key source of negation failures: the model's mechanism relies on suppressing the original answer rather than using it to determine what to exclude, so the model can repeat the original answer when suppression is too weak or when a bias toward particular answers prevents it from selecting an alternative. To address this weakness in the model's negation mechanism, we propose a training objective that requires larger shifts in answer preference for more confident original predictions, and show that it reduces negation failures with less degradation of general capabilities than standard fine-tuning baselines. Together, our results demonstrate how mechanistic analysis can reveal why a linguistic capability fails and guide training that targets the underlying limitation.

Figures & tables

Appendix figures & tables28 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. How Language Models Process Negation

    May 4, 2026Zhejian Zhou, Tianyi Zhou, Robin Jia +1NegationLarge Language Models Fail

  2. Negation Neglect: When models fail to learn negations in training

    May 13, 2026Harry Mayne, Lev McKinney, Jan Dubiński +3NegationArtificial Intelligence Safety

  3. Revisiting the Systematicity in Negation in the Era of In-Context Learning

    Jun 15, 2026Hitomi Yanaka, Taisei YamamotoNegationLarge Language Models Fail