Abstract
Fallacy-detection benchmarks pair fallacy classes with a single "valid" or "none" class that takes everything data collection did not label as a fallacy. A detector has two jobs, deciding whether an argument is fallacious and naming which fallacy it commits, and the false-positive rate is meant to measure the first. We show that what these benchmarks actually score is scheme recognition, the ability behind the second job. Their own test sets already show it: when a classifier misses a fallacy, the error lands on "none" rather than on another fallacy type, so detection is failing while classification holds. The reason is what the valid class lacks. The negatives that separate the two jobs are correct arguments using the same argumentation scheme as a fallacy, and they are scarce: nearly absent from the four benchmarks we examined, and rare even under deliberate search. A detector is therefore never tested where recognizing a scheme and judging its use come apart, and can pass on recognition alone. We construct the missing arguments, together with a control condition from the same pipeline that differs only in scheme, so whatever generation contributes, it contributes to both. The classifier labels the scheme-matched negatives as the source fallacy, and labels the wrong-scheme negatives as the scheme they actually use 85.9% of the time and as the source type 0.4%. The classifier has learned which scheme an argument uses, not whether it uses it correctly. The over-flagging follows: a model that scores 16.6% on CoCoLoFa's own valid class flags 58.9% of the constructed arguments. The same dissociation appears in three zero-shot LLM detectors that never saw these benchmarks. We release the items as Scheme Foils. A reported false-positive rate should not be trusted as a measure of detection until the valid class has been audited for scheme-matched coverage.
Explore similar work
Jun 20, 2026cs.AI
Current evaluations of Large Language Models (LLMs) on logical fallacy detection focus on predicted labels, but do not establish whether those labels are supported by the reasoning the models provide. We propose ForEx (Formal Verification for Explainable Reasoning), a framework that translates LLM-generated explanations into Lean4 and verifies whether the translated rationale is derivable under encoded premises, not the logical validity of the original natural language argument. To distinguish prediction outcomes from the formal status of the supporting reasoning, we introduce the LLM Argument Verification Matrix, which separates label consistency from formal verification status. Experiments on LOGIC-Climate show that over 90% of LLM outputs can be translated into formal reasoning chains that pass verification, while agreement with human annotations remains around 20%. These results expose a systematic gap between formal derivability and label agreement, a distinction invisible to prediction-based metrics. ForEx moves LLM evaluation beyond label correctness toward machine-checkable analysis of formalized reasoning chains.
Pei-Cing Huang, Chienyu Liu, Chan Hsu +3
Jun 30, 2026cs.CL
Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies remains underexplored. Prior work has primarily examined whether LLMs can identify or classify fallacies, leaving their robustness against fallacious persuasion insufficiently studied. To address this gap, we introduce LoFa (Logical Fallacy), a comprehensive benchmark for evaluating LLM robustness against fallacies. LoFa is constructed through a multi-agent pipeline that pairs factual questions with fallacious arguments, and is accompanied by a multi-round debate framework for assessing model resilience under sustained adversarial persuasion. To disentangle fallacy robustness from a model's inherent knowledge limitations, we further propose Logical Fallacy Resistance at k (LFR@k), a metric that quantifies resistance to fallacious attacks. Experiments show that LLMs exhibit varying levels of robustness across different fallacy types, revealing distinct vulnerability profiles among models.
Xudong Shen, Li Yuan, Ye Chen +3
Jun 25, 2026cs.CL
In today's fast-paced information era, logical fallacies, defined as defective patterns of reasoning, inevitably contribute to the growth of information disorder. However, often fallacies appear in nuanced forms that complicate automated classification. In this study, we investigate whether merging abstract logical structures with context-level linguistic cues proves beneficial for fallacy classification, developing a framework that inductively extracts such patterns from fallacious examples and their explanations using Large Language Models (LLMs). We evaluate the impact of these patterns across different LLMs and experimental zero- and one-shot configurations, showing statistically significant improvements over zero-shot baselines and outperforming competing approaches. Cross-dataset experiments validate generalization, establishing data-driven pattern extraction as an effective method for generating logical representations.
Eleni Papadopulos, Firoj Alam, Giovanni Da San Martino