Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to obscure logical leaps, hallucinated execution steps, and internal inconsistencies behind plausible prose. Our manual audit reveals that approximately 60% of correct vulnerability verdicts are accompanied by fabricated or unverifiable claims, and free-form explanations allow reasoning errors to evade LLM-as-a-judge evaluation. We present Vulnerability Explanation Reasoning Auditor (VERA), an automated framework for auditing LLM vulnerability reasoning. Rather than accepting free-form text, VERA asks models to output a Structured Reasoning Record (SRR) encoding tracked pointers, memory operations, and state transitions in machine-readable fields. A multi-stage judge audits each SRR against eight reasoning failure modes using deterministic checks, with LLM calls reserved for semantic interpretation. The standardized SRR schema also enables automated mutation testing to benchmark judges at scale without human annotation. Our evaluation shows reasoning flaws occur in correct verdicts just as frequently as incorrect ones, and VERA exposes 87% of reasoning errors that free-form LLM-as-judge systematically miss.
Figures & tables
Figure 1. Overview of VERA with a motivating example of correct verdict, flawed reasoning .
Verdict Correctness Metrics
Ext. Knowledge
Reasoning
Model
Binary Acc.
CWE recall
FP
Both V% a
Seen b / Wrong%
Vague% c
GPT-5.5 ( Singh and others, 2026 )
60%
16%
49%
44%
2% / —
19%
GPT-5.6-sol ( Singh and others, 2026 )
59%
13%
38%
33%
0% / —
28%
Gemma-4-26b ( Team, 2026 )
46%
3%
74%
58%
0% / —
22%
Sonnet-5 ( Anthropic, 2026 )
59%
13%
46%
42%
44% / 61%
34%
Table 1. Detection reliability and reasoning quality of free-form CoT (64 CWE-416 pairs, SVEN).
Figure 3
Figure 2 . Structured Reasoning Record (SRR) for the code snippet in Fig. 1(c) . The four blocks organize the detection model’s reasoning, including ungrounded assumptions and incorrect state transitions, into named, typed fields that VERA audits against the failure taxonomy. Red indicates errors detected by VERA.
Name
Explanation
Example
Lexical and evidence integrity
F1 Scope Error
Incorrect list of cwe-relevant variables and/or state-changing operations, and incorrect initial state of the variables.
Model overlooked e as a pointer.
F2 External Assumption
Unsupported claims about caller/callee behavior, or project conventions treated as an established fact.
Model infers that cache_remove releases e based on its name or comments.
State correctness
F3 State Transition Error
Transition applied to the wrong variable or the wrong operation; the state change does not follow the operation; correct operations but incorrectly sequenced relative to source
Model records free(e) but notes the transition to ALLOC
F4 Alias Error
Missing or hallucinated alias, or states not correctly propagated to its co-referencing aliases.
Model marks p1 as freed but leaves alias p2 as valid , then concludes p2 is safe to dereference.
Table 2. Taxonomy of reasoning errors in state-based vulnerability analysis.
Reasoning Failure
Mutation Operators
F1 scope error
M1a : add a nonexistent variable or remove a random variable; M1b : add a nonexistent operation or remove a random operation; M1c : introduce incorrect initial state
F2 ungrounded assumptions
M2 : inject ungrounded external assumption into reasoning prose
F3 incorrect state transitions
M3 : invert state-transition ordering or, insert spurious free step on leaked variable (included for CWE-401)
F4 incorrect aliases
M4 : add or erase the trace steps identifying aliases
F5 incomplete trace
M5a : make trace step references undeclared path ID; M5b : add or drop one random trace step
F6 fabricated path
M6a : intentionally make path-activation condition vague; M6b : make trace step reference undeclared path ID; M6c : make path activation reference undeclared path ID
Table 3 . Mapping between reasoning failure types and each mutation operator.
GPT
Mistral
DeepSeek
Failure category / Mutation operator
Representative error labels
416
476
415
401
772
416
476
415
401
772
416
476
415
401
772
F1 Scope Error
565
807
517
558
328
365
508
295
338
182
1056
1465
1007
1074
580
M1a b Add/remove identifier
phantom/missing_variable
479
678
454
433
254
311
454
241
272
149
675
949
639
625
324
M1b b Add/remove op
phantom_op , missing_memory_op
85
125
62
123
74
54
54
54
66
33
369
481
353
433
250
M1c Wrong initial state
incorrect_initial_state
1
4
1
2
0
0
0
0
0
0
12
35
15
16
6
F2 External Assumption
77
94
84
100
66
202
291
149
166
66
627
878
583
586
298
Table 4 . RQ1: error labels obtained with open-ended error enumeration mapped to our error taxonomy and mutation operators. Error label occurrence counts obtained with three judge models across 3000 reasoning records produced by six detection models. A record may receive multiple labels. The bottom section lists the 21 instances ( < 0.05%) not assignable to any class; the eight categories together cover > 99.95% of all observed error instances.
Stage
CWE-416 ( N=711 )
CWE-476 ( N=1006 )
CWE-415 ( N=332 )
CWE-401 ( N=541 )
CWE-772 ( N=178 )
GPT
DS
Mis
GPT
DS
Mis
GPT
DS
Mis
GPT
DS
Mis
GPT
DS
Mis
S1 Scope
VERA
.53/ .00
.56/ .00
.68/ .00
.61/ .00
.68/ .00
B3
.62/.13
.69/.30
.98/.01
.64/.04
.77/.19
1.00 / .00
.53/.15
.73/.20
1.00 / .00
.75/.11
.74/.24
1.00 / .00
.63/.03
.79/.17
1.00 /.02
S2 Evidence
VERA
.33/.06
.41/.11
.97/.07
.44/.09
.49/.14
1.00 /.06
.47/.08
.44/.10
.95/.12
.25/.04
.42/.09
.98/.08
.20/.06
.55/.13
.90/ .00
B3
.34/.05
.32/.11
.96/.09
.43/.12
.61/.15
1.00 /.07
.47/.03
.47/.15
.95/.08
.25/.08
.45/.15
.96/.08
.25/.08
.50/.09
.90/ .00
S3 Path ID
VERA
1.00 / .00
1.00 / .00
1.00 / .00
1.00 / .00
1.00 / .00
Table 5 . Per-stage mutation detection (recall / FP) across five CWE classes, data from different datasets aggregated. VERA vs. B3 (SRR + structured LLM judge). Judge models: GPT = GPT-4.1, DS = DeepSeek, Mis = Mistral. N : number of mutants; -: no mutant generated. Bold : recall = 1.00 or FP = 0.00. Gray cells : LLM-dependent (all B3 cells; VERA S2, S7 which are primarily LLM).
Table 6 . RQ3 comparing VERA structured framework with free-form baselines.