Reliable learning-based vulnerability detection requires high-quality labels, yet datasets built from vulnerability-fixing commits may label functions as vulnerable simply because they were changed by a security patch. We present VulValidate, a framework that uses LLM agents to coordinate dynamic analysis tools and construct vulnerability-triggering experiments from runtime feedback. Given a labeled function and its fixing patch, VulValidate reconstructs vulnerable and fixed revisions, selects suitable tools and execution paths, refines triggering inputs, and compares runtime behavior to assess function-level attribution. We audit all 35,849 instances originally labeled vulnerable in BigVul, PrimeVul, and DiverseVul. We confirm 20,510 (57.2%), correct 6,819 labels (19.0%), leave 7,981 attacked but undecided (22.3%), and cannot successfully measure 539 (1.5%). After conflict resolution and byte-exact deduplication, the corrected release contains 15,890 distinct confirmed vulnerable function bodies. In a blinded review of 581 sampled decisions, expert consensus supports 90.0%--92.0% of confirmations and 92.6%--99.0% of label corrections. With model parameters fixed, corrected evaluation lowers F1 for all five tested detectors on both BigVul and DiverseVul; retraining with corrected labels improves F1 for four of five detectors on each dataset. We also release a reusable VulValidate skill, corrected datasets, and reproducible evidence for future vulnerability-detection research.
Figures & tables
Figure 1 . Overview of VulValidate . The workflow reconstructs historical source and environment (S1), plans the experiment (S2), iteratively constructs triggers (S3), validates vulnerable and fixed executions (S4), preserves evidence and Docker-based replay materials (S5), and assigns an evidence-based audit outcome (S6). Two experts review randomly sampled instances blinded to the agent’s decisions.
Check
Required evidence
1
Measured execution
A replay script, actual integer exit codes, and complete logs, including both sides for a two-sided claim
2
Function execution
Trace, coverage, or debugger evidence of labelled-function execution
3
Attack input
A saved input, failure condition, or schedule targeting the mechanism, with construction and purpose documented
4
Oracle relevance
An oracle capable of observing the suspected defect and a mechanism-specific rationale ( Barr et al., 2015 )
5
Search scope
Tested entry points, value ranges, corpus or schedules, duration or iteration bounds, and stopping reason ( Klees et al., 2018 )
6
Source completeness
Comparison of the tested instance with the complete function at the verified upstream parent revision
Table 1 . Seven necessary checks for a mechanism-specific dynamic negative. Each requires inspectable source or execution evidence.
Labelled function
BigVul index
Role in the fault
Outcome and deciding evidence
ZSTD_buildCTable
182843
frame #0
Confirmed. ASan WRITE of size 1 past a 192-byte buffer at the disputed store; exit codes 1 against 0
ZSTD_compressSequences_internal
182844
frame #1, a caller
Label correction. gcov: 12 calls per side, disputed line 3 times per side, no fault
ZSTD_encodeSequences and …_body
182845, 182846
absent from the stack
Label correction. 3 calls per side, disputed lines executed on both sides, no fault
Table 2 . The four functions labelled vulnerable by the single fixing commit 3e5cdf1 , and the evidence that decided each.
Table 3 . Validation of the agent’s decisions. In (a), Risse and Croft denote the manual labellings of Risse et al. ( Risse et al., 2025 ) and Croft et al. ( Croft et al., 2023 ) , n counts functions with a decision on both sides, and chance is the agreement expected from the marginal rates. In (b), S, C, and I denote Supported, Contradicted, and Insufficient evidence; S/n includes all reviewed instances, including those later excluded from the final datasets; Label correction denotes the S6 outcome defined in Section 2.5 . Agreement, chance, and S/n are percentages.
Outcome
BigVul
PrimeVul
DiverseVul
Total
Distinct total
Confirmed
4,412 (40.48)
5,537 (92.22)
10,561 (55.74)
20,510 (57.21)
15,890
Label correction
4,730 (43.39)
87 (1.45)
2,002 (10.57)
6,819 (19.02)
6,774
Attacked but undecided
1,480 (13.58)
374 (6.23)
6,127 (32.34)
7,981 (22.26)
7,777
Not successfully measured
278 (2.55)
6 (0.10)
255 (1.35)
539 (1.51)
539
Overall total
10,900 (100.00)
6,004 (100.00)
18,945 (100.00)
35,849 (100.00)
30,980
Table 4 . Post-review outcome distribution. The first four numeric columns give counts and percentages of each column’s originally vulnerable population. Distinct total counts byte-identical bodies within each outcome, preserving comments and whitespace; the bottom value counts each body once across all outcomes. The Attacked but undecided row includes the 36 reviewed instances withheld from both learning classes after manual review.
Projects: all instances
CWE: all instances
CWE: vulnerable instances
Project
Count
%
CWE
Count
%
CWE
Count
%
linux
123,137
22.91
119
49,574
9.22
119
1,973
12.42
Chrome
74,643
13.88
787
43,230
8.04
125
1,674
10.53
qemu
10,760
2.00
20
43,098
8.02
787
1,673
10.53
php-src
10,602
1.97
125
34,985
6.51
20
1,272
8.01
linux-2.6
10,197
1.90
416
31,061
5.78
200
830
5.22
Table 5 . Ten largest source project groups and CWE categories in the release, ranked independently. Project and full-dataset CWE percentages use 537,591 instances; vulnerable-subset CWE percentages use 15,890. An instance may have multiple CWE annotations.
Model
Train
Test
BigVul
DiverseVul
ACC
Prec.
Rec.
F1
MCC
AUROC
ACC
Prec.
Rec.
F1
MCC
AUROC
CodeT5
O
O
.9878
.9440
.8303
.8835
.8791
.9827
.8962
.2570
.4369
.3236
.2827
.8257
C
C
.9832
.7086
.7007
.7047
.6961
.9800
.9170
.1607
.3658
.2233
.2042
.8162
O
C
.9697
.4817
.8309
.6098
.6193
.9719
.9064
.1623
.4490
.2385
.2299
.8301
O
E
.9891
.8040
.8309
.8172
.8117
.9840
.9083
.1667
.4490
.2432
.2340
.8319
C
E
.9886
.8850
.7007
.7822
.7820
.9846
.9172
.1623
.3658
.2248
.2054
.8166
Table 6 . Vulnerability detection performance on BigVul and DiverseVul. Each model has five settings, O/O , C/C , O/C , O/E , and C/E , where E removes only label-correction instances from the corrected test set and leaves the remaining instances and labels unchanged. Graph availability restricts model-specific test populations. Label combinations and threshold selection follow Section 3.5 . Leading zeros are omitted.
Model
ACC
Prec.
Rec.
F1
MCC
AUROC
CodeT5
.9625
.4471
.3966
.4203
.4018
.9125
LineVul
.9654
.4888
.1688
.2509
.2734
.8423
DeepDFA
.9236
.1563
.2541
.1936
.1609
.7559
ReVeal
.9342
.1840
.2422
.2091
.1773
.7858
Devign
.9462
.2452
.2399
.2425
.2146
.7371
Table 7 . Test performance on the merged benchmark constructed from BigVul, PrimeVul, and DiverseVul. Threshold selection follows Section 3.5 . Leading zeros are omitted.