cs.CROct 5, 2026

Correct Verdicts, Flawed Reasoning: Structured Auditing of LLM-based Vulnerability Reasoning

Authors: Boyue Caroline Hu, Kaivalya Ahir, Ronghao Ni, Limin Jia

Organizations: Carnegie Mellon University, Pittsburgh, Pennsilvanya, USA

Abstract

Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to obscure logical leaps, hallucinated execution steps, and internal inconsistencies behind plausible prose. Our manual audit reveals that approximately 60% of correct vulnerability verdicts are accompanied by fabricated or unverifiable claims, and free-form explanations allow reasoning errors to evade LLM-as-a-judge evaluation. We present Vulnerability Explanation Reasoning Auditor (VERA), an automated framework for auditing LLM vulnerability reasoning. Rather than accepting free-form text, VERA asks models to output a Structured Reasoning Record (SRR) encoding tracked pointers, memory operations, and state transitions in machine-readable fields. A multi-stage judge audits each SRR against eight reasoning failure modes using deterministic checks, with LLM calls reserved for semantic interpretation. The standardized SRR schema also enables automated mutation testing to benchmark judges at scale without human annotation. Our evaluation shows reasoning flaws occur in correct verdicts just as frequently as incorrect ones, and VERA exposes 87% of reasoning errors that free-form LLM-as-judge systematically miss.

Figures & tables

Explore similar work

CardsList
  1. EntailLLM: Verifying LLM-Generated Vulnerability Discovery Paths with Domain Knowledge via Logic Programming

    Aug 3, 2026Kaustuv Mukherji, Jaikrishna Manojkumar Patil, Colton Payne +4Model VulnerabilitiesGraph Representations

  2. VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection

    May 10, 2026Wenxin Tang, Xiang Zhang, Junliang Liu +11Voltage

  3. Dissecting the Black Box: Circuit-Level Analysis of LLM Vulnerability Detection

    May 28, 2026Syafiq Al Atiiq, Chun Zhou, Christian GehrmannVulnerable CodeLarge Language Model Safety