VulValidate: Auditing Function-Level Vulnerability Labels with Executable Evidence
Organizations: University of Louisiana at Lafayette, USA
Abstract
Reliable learning-based vulnerability detection requires high-quality labels, yet datasets built from vulnerability-fixing commits may label functions as vulnerable simply because they were changed by a security patch. We present VulValidate, a framework that uses LLM agents to coordinate dynamic analysis tools and construct vulnerability-triggering experiments from runtime feedback. Given a labeled function and its fixing patch, VulValidate reconstructs vulnerable and fixed revisions, selects suitable tools and execution paths, refines triggering inputs, and compares runtime behavior to assess function-level attribution. We audit all 35,849 instances originally labeled vulnerable in BigVul, PrimeVul, and DiverseVul. We confirm 20,510 (57.2%), correct 6,819 labels (19.0%), leave 7,981 attacked but undecided (22.3%), and cannot successfully measure 539 (1.5%). After conflict resolution and byte-exact deduplication, the corrected release contains 15,890 distinct confirmed vulnerable function bodies. In a blinded review of 581 sampled decisions, expert consensus supports 90.0%--92.0% of confirmations and 92.6%--99.0% of label corrections. With model parameters fixed, corrected evaluation lowers F1 for all five tested detectors on both BigVul and DiverseVul; retraining with corrected labels improves F1 for four of five detectors on each dataset. We also release a reusable VulValidate skill, corrected datasets, and reproducible evidence for future vulnerability-detection research.
Figures & tables
| Check | Required evidence | |
|---|---|---|
| 1 | Measured execution | A replay script, actual integer exit codes, and complete logs, including both sides for a two-sided claim |
| 2 | Function execution | Trace, coverage, or debugger evidence of labelled-function execution |
| 3 | Attack input | A saved input, failure condition, or schedule targeting the mechanism, with construction and purpose documented |
| 4 | Oracle relevance | An oracle capable of observing the suspected defect and a mechanism-specific rationale ( Barr et al., 2015 ) |
| 5 | Search scope | Tested entry points, value ranges, corpus or schedules, duration or iteration bounds, and stopping reason ( Klees et al., 2018 ) |
| 6 | Source completeness | Comparison of the tested instance with the complete function at the verified upstream parent revision |
| Labelled function | BigVul index | Role in the fault | Outcome and deciding evidence |
| ZSTD_buildCTable | 182843 | frame #0 | Confirmed. ASan WRITE of size 1 past a 192-byte buffer at the disputed store; exit codes 1 against 0 |
| ZSTD_compressSequences_internal | 182844 | frame #1, a caller | Label correction. gcov: 12 calls per side, disputed line 3 times per side, no fault |
| ZSTD_encodeSequences and …_body | 182845, 182846 | absent from the stack | Label correction. 3 calls per side, disputed lines executed on both sides, no fault |
| Outcome | BigVul | PrimeVul | DiverseVul | Total | Distinct total |
|---|---|---|---|---|---|
| Confirmed | 4,412 (40.48) | 5,537 (92.22) | 10,561 (55.74) | 20,510 (57.21) | 15,890 |
| Label correction | 4,730 (43.39) | 87 (1.45) | 2,002 (10.57) | 6,819 (19.02) | 6,774 |
| Attacked but undecided | 1,480 (13.58) | 374 (6.23) | 6,127 (32.34) | 7,981 (22.26) | 7,777 |
| Not successfully measured | 278 (2.55) | 6 (0.10) | 255 (1.35) | 539 (1.51) | 539 |
| Overall total | 10,900 (100.00) | 6,004 (100.00) | 18,945 (100.00) | 35,849 (100.00) | 30,980 |
| Projects: all instances | CWE: all instances | CWE: vulnerable instances | ||||||
| Project | Count | % | CWE | Count | % | CWE | Count | % |
| linux | 123,137 | 22.91 | 119 | 49,574 | 9.22 | 119 | 1,973 | 12.42 |
| Chrome | 74,643 | 13.88 | 787 | 43,230 | 8.04 | 125 | 1,674 | 10.53 |
| qemu | 10,760 | 2.00 | 20 | 43,098 | 8.02 | 787 | 1,673 | 10.53 |
| php-src | 10,602 | 1.97 | 125 | 34,985 | 6.51 | 20 | 1,272 | 8.01 |
| linux-2.6 | 10,197 | 1.90 | 416 | 31,061 | 5.78 | 200 | 830 | 5.22 |
| Model | Train | Test | BigVul | DiverseVul | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ACC | Prec. | Rec. | F1 | MCC | AUROC | ACC | Prec. | Rec. | F1 | MCC | AUROC | |||
| CodeT5 | O | O | .9878 | .9440 | .8303 | .8835 | .8791 | .9827 | .8962 | .2570 | .4369 | .3236 | .2827 | .8257 |
| C | C | .9832 | .7086 | .7007 | .7047 | .6961 | .9800 | .9170 | .1607 | .3658 | .2233 | .2042 | .8162 | |
| O | C | .9697 | .4817 | .8309 | .6098 | .6193 | .9719 | .9064 | .1623 | .4490 | .2385 | .2299 | .8301 | |
| O | E | .9891 | .8040 | .8309 | .8172 | .8117 | .9840 | .9083 | .1667 | .4490 | .2432 | .2340 | .8319 | |
| C | E | .9886 | .8850 | .7007 | .7822 | .7820 | .9846 | .9172 | .1623 | .3658 | .2248 | .2054 | .8166 | |
| Model | ACC | Prec. | Rec. | F1 | MCC | AUROC |
|---|---|---|---|---|---|---|
| CodeT5 | .9625 | .4471 | .3966 | .4203 | .4018 | .9125 |
| LineVul | .9654 | .4888 | .1688 | .2509 | .2734 | .8423 |
| DeepDFA | .9236 | .1563 | .2541 | .1936 | .1609 | .7559 |
| ReVeal | .9342 | .1840 | .2422 | .2091 | .1773 | .7858 |
| Devign | .9462 | .2452 | .2399 | .2425 | .2146 | .7371 |