cs.CROct 4, 2026

VulValidate: Auditing Function-Level Vulnerability Labels with Executable Evidence

Authors: Leizhen Zhang, Sheng Chen

Organizations: University of Louisiana at Lafayette, USA

Abstract

Reliable learning-based vulnerability detection requires high-quality labels, yet datasets built from vulnerability-fixing commits may label functions as vulnerable simply because they were changed by a security patch. We present VulValidate, a framework that uses LLM agents to coordinate dynamic analysis tools and construct vulnerability-triggering experiments from runtime feedback. Given a labeled function and its fixing patch, VulValidate reconstructs vulnerable and fixed revisions, selects suitable tools and execution paths, refines triggering inputs, and compares runtime behavior to assess function-level attribution. We audit all 35,849 instances originally labeled vulnerable in BigVul, PrimeVul, and DiverseVul. We confirm 20,510 (57.2%), correct 6,819 labels (19.0%), leave 7,981 attacked but undecided (22.3%), and cannot successfully measure 539 (1.5%). After conflict resolution and byte-exact deduplication, the corrected release contains 15,890 distinct confirmed vulnerable function bodies. In a blinded review of 581 sampled decisions, expert consensus supports 90.0%--92.0% of confirmations and 92.6%--99.0% of label corrections. With model parameters fixed, corrected evaluation lowers F1 for all five tested detectors on both BigVul and DiverseVul; retraining with corrected labels improves F1 for four of five detectors on each dataset. We also release a reusable VulValidate skill, corrected datasets, and reproducible evidence for future vulnerability-detection research.

Figures & tables

Explore similar work

CardsList
  1. VulTriage: Triple-Path Context Augmentation for LLM-Based Vulnerability Detection

    May 10, 2026Wenxin Tang, Xiang Zhang, Junliang Liu +11Software Vulnerability DetectionStatic Code Analysis

  2. ANVIL: Anomaly-based Vulnerability Identification without Labelled Training Data

    Aug 28, 2024Weizhou Wang, Eric Liu, Xiangyu Guo +3Software Vulnerability DetectionLLMs for Cybersecurity

  3. VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

    Aug 12, 2026Jin Lu, Xuening Han, Yang Zhong +4Software Vulnerability DetectionSoftware Engineering Benchmarks