Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to obscure logical leaps, hallucinated execution steps, and internal inconsistencies behind plausible prose. Our manual audit reveals that approximately 60% of correct vulnerability verdicts are accompanied by fabricated or unverifiable claims, and free-form explanations allow reasoning errors to evade LLM-as-a-judge evaluation. We present Vulnerability Explanation Reasoning Auditor (VERA), an automated framework for auditing LLM vulnerability reasoning. Rather than accepting free-form text, VERA asks models to output a Structured Reasoning Record (SRR) encoding tracked pointers, memory operations, and state transitions in machine-readable fields. A multi-stage judge audits each SRR against eight reasoning failure modes using deterministic checks, with LLM calls reserved for semantic interpretation. The standardized SRR schema also enables automated mutation testing to benchmark judges at scale without human annotation. Our evaluation shows reasoning flaws occur in correct verdicts just as frequently as incorrect ones, and VERA exposes 87% of reasoning errors that free-form LLM-as-judge systematically miss.
Figures & tables
Figure 1. Overview of VERA with a motivating example of correct verdict, flawed reasoning .
Verdict Correctness Metrics
Ext. Knowledge
Reasoning
Model
Binary Acc.
CWE recall
FP
Both V% a
Seen b / Wrong%
Vague% c
GPT-5.5 ( Singh and others, 2026 )
60%
16%
49%
44%
2% / —
19%
GPT-5.6-sol ( Singh and others, 2026 )
59%
13%
38%
33%
0% / —
28%
Gemma-4-26b ( Team, 2026 )
46%
3%
74%
58%
0% / —
22%
Sonnet-5 ( Anthropic, 2026 )
59%
13%
46%
42%
44% / 61%
34%
Table 1. Detection reliability and reasoning quality of free-form CoT (64 CWE-416 pairs, SVEN).
Figure 3
Figure 2 . Structured Reasoning Record (SRR) for the code snippet in Fig. 1(c) . The four blocks organize the detection model’s reasoning, including ungrounded assumptions and incorrect state transitions, into named, typed fields that VERA audits against the failure taxonomy. Red indicates errors detected by VERA.
Name
Explanation
Example
Lexical and evidence integrity
F1 Scope Error
Incorrect list of cwe-relevant variables and/or state-changing operations, and incorrect initial state of the variables.
Model overlooked e as a pointer.
F2 External Assumption
Unsupported claims about caller/callee behavior, or project conventions treated as an established fact.
Model infers that cache_remove releases e based on its name or comments.
State correctness
F3 State Transition Error
Transition applied to the wrong variable or the wrong operation; the state change does not follow the operation; correct operations but incorrectly sequenced relative to source
Model records free(e) but notes the transition to ALLOC
F4 Alias Error
Missing or hallucinated alias, or states not correctly propagated to its co-referencing aliases.
Model marks p1 as freed but leaves alias p2 as valid , then concludes p2 is safe to dereference.
Table 2. Taxonomy of reasoning errors in state-based vulnerability analysis.
Reasoning Failure
Mutation Operators
F1 scope error
M1a : add a nonexistent variable or remove a random variable; M1b : add a nonexistent operation or remove a random operation; M1c : introduce incorrect initial state
F2 ungrounded assumptions
M2 : inject ungrounded external assumption into reasoning prose
F3 incorrect state transitions
M3 : invert state-transition ordering or, insert spurious free step on leaked variable (included for CWE-401)
F4 incorrect aliases
M4 : add or erase the trace steps identifying aliases
F5 incomplete trace
M5a : make trace step references undeclared path ID; M5b : add or drop one random trace step
F6 fabricated path
M6a : intentionally make path-activation condition vague; M6b : make trace step reference undeclared path ID; M6c : make path activation reference undeclared path ID
Table 3 . Mapping between reasoning failure types and each mutation operator.
GPT
Mistral
DeepSeek
Failure category / Mutation operator
Representative error labels
416
476
415
401
772
416
476
415
401
772
416
476
415
401
772
F1 Scope Error
565
807
517
558
328
365
508
295
338
182
1056
1465
1007
1074
580
M1a b Add/remove identifier
phantom/missing_variable
479
678
454
433
254
311
454
241
272
149
675
949
639
625
324
M1b b Add/remove op
phantom_op , missing_memory_op
85
125
62
123
74
54
54
54
66
33
369
481
353
433
250
M1c Wrong initial state
incorrect_initial_state
1
4
1
2
0
0
0
0
0
0
12
35
15
16
6
F2 External Assumption
77
94
84
100
66
202
291
149
166
66
627
878
583
586
298
Table 4 . RQ1: error labels obtained with open-ended error enumeration mapped to our error taxonomy and mutation operators. Error label occurrence counts obtained with three judge models across 3000 reasoning records produced by six detection models. A record may receive multiple labels. The bottom section lists the 21 instances ( < 0.05%) not assignable to any class; the eight categories together cover > 99.95% of all observed error instances.
Stage
CWE-416 ( N=711 )
CWE-476 ( N=1006 )
CWE-415 ( N=332 )
CWE-401 ( N=541 )
CWE-772 ( N=178 )
GPT
DS
Mis
GPT
DS
Mis
GPT
DS
Mis
GPT
DS
Mis
GPT
DS
Mis
S1 Scope
VERA
.53/ .00
.56/ .00
.68/ .00
.61/ .00
.68/ .00
B3
.62/.13
.69/.30
.98/.01
.64/.04
.77/.19
1.00 / .00
.53/.15
.73/.20
1.00 / .00
.75/.11
.74/.24
1.00 / .00
.63/.03
.79/.17
1.00 /.02
S2 Evidence
VERA
.33/.06
.41/.11
.97/.07
.44/.09
.49/.14
1.00 /.06
.47/.08
.44/.10
.95/.12
.25/.04
.42/.09
.98/.08
.20/.06
.55/.13
.90/ .00
B3
.34/.05
.32/.11
.96/.09
.43/.12
.61/.15
1.00 /.07
.47/.03
.47/.15
.95/.08
.25/.08
.45/.15
.96/.08
.25/.08
.50/.09
.90/ .00
S3 Path ID
VERA
1.00 / .00
1.00 / .00
1.00 / .00
1.00 / .00
1.00 / .00
Table 5 . Per-stage mutation detection (recall / FP) across five CWE classes, data from different datasets aggregated. VERA vs. B3 (SRR + structured LLM judge). Judge models: GPT = GPT-4.1, DS = DeepSeek, Mis = Mistral. N : number of mutants; -: no mutant generated. Bold : recall = 1.00 or FP = 0.00. Gray cells : LLM-dependent (all B3 cells; VERA S2, S7 which are primarily LLM).
Table 6 . RQ3 comparing VERA structured framework with free-form baselines.
Large language models are increasingly used to reason about software vulnerabilities, but their outputs can silently violate domain knowledge, limiting their reliability in safety-critical settings such as medical devices. Prior work either treats that output as a prediction to be scored or constrains it to walks within a single knowledge graph; neither checks whether reasoning over a binary is consistent with an independent body of domain knowledge. We present EntailLLM, which validates each LLM-proposed analyst path by entailment: the path is a traversal of the binary's function call graph, the domain knowledge is represented in a separate graph, and verification aligns the two under temporal annotated logic. Across three CWE classes, four LLMs, three prompting strategies, and seven binaries varying in size from 405 to 12,696 function call-graph nodes, domain knowledge raises pooled entailment from 78% to 98%, with entailment decreasing in only 3% of the experiments. EntailLLM is deployed end-to-end on real medical-device binaries, reaching 98% pooled entailment without per-device tuning. Our system inherits the formal guarantees of generalized annotated logic, providing logical verification of LLM output that is both explainable and grounded in well-defined semantics.
Automated vulnerability detection is a fundamental task in software security, yet existing learning-based methods still struggle to capture the structural dependencies, domain-specific vulnerability knowledge, and complex program semantics required for accurate detection. Recent Large Language Models (LLMs) have shown strong code understanding ability, but directly prompting them with raw source code often leads to missed vulnerabilities or false alarms, especially when vulnerable and benign functions differ only in subtle semantic details. To address this, we propose VulTriage, a triple-path context augmentation framework for LLM-based vulnerability detection. VulTriage enhances the LLM input through three complementary paths: a Control Path that extracts and verbalizes AST, CFG, and DFG information to expose control and data dependencies; a Knowledge Path that retrieves relevant CWE-derived vulnerability patterns and examples through hybrid dense--sparse retrieval; and a Semantic Path that summarizes the functional behavior of the code before the final judgment. These contexts are integrated into a unified instruction to guide the LLM toward more reliable vulnerability reasoning. Experiments on the PrimeVul pair test set show that VulTriage achieves state-of-the-art performance, outperforming existing deep learning and LLM-based baselines on key pair-wise and classification metrics. Further ablation studies verify the effectiveness of each path, and additional experiments on the Kotlin dataset demonstrate the generalization ability of VulTriage under low-resource and class-imbalanced settings. Our code is available at https://github.com/vinsontang1/VulTriage
Wenxin Tang, Xiang Zhang, Junliang Liu +11
Tsinghua University · Henan University · Dalian Maritime University +8
Large language models (LLMs) can detect software vulnerabilities, but how do they actually identify vulnerable code? We address this question using mechanistic interpretability; analyzing the internal computations of a neural network to understand its reasoning process.Using Circuit Tracer on Gemma-2-2b, we trace the computational pathways activated when the model classifies 472 C/C++ code samples as vulnerable or safe. Our analysis reveals a surprising finding: the model primarily relies on safety detectors, attention heads that recognize safe coding patterns, rather than directly detecting vulnerability signatures. When these safety detectors fail to activate, the model classifies code as vulnerable. We identify the critical neural components: specific attention heads in early layers (L5, L7) that focus on safety patterns, and Multilayer Perceptron (MLP) neurons in Layer 7 that encode vulnerability-related features. Ablation experiments confirm their causal role; removing Layer 11 drops vulnerability detection accuracy from 100% to 6%, while ablating just 20 neurons in Layer 7 reduces it by 50%.Our findings show that LLM vulnerability detection uses sparse, interpretable circuits (only 16% of model capacity), enabling circuit-level explanations for security predictions and targeted improvements to detection systems.