Organizations: University of Notre Dame · Bake AI · Vanderbilt University · University of Pennsylvania · LMU Munich · University of Washington · Stanford University · Inria
Assessing cybersecurity vulnerability awareness in coding agents requires evaluations that reveal capability gaps and remain informative as models evolve. Static benchmarks offer fixed coverage and difficulty, while scarce vulnerable repositories and costly expert authoring limit their renewal at scale. We introduce SecProbe, a framework for adaptive evaluation that combines Item Response Theory (IRT) with on-demand synthesis of repository-scale vulnerability-repair tasks. From observed performance, \textsc{SecProbe} estimates agent ability and identifies where additional evidence is most informative, selecting existing tasks or synthesizing new ones accordingly. As one use case, we construct 353 tasks spanning six programming languages and 151 CWE types and evaluate nine frontier models with two agent harnesses. Success rates peak at 28.33%, highlighting substantial gaps in vulnerability recognition and repair. Compared with random and one-shot baselines, \textsc{SecProbe} achieves comparable agent ability estimates while requiring agents to solve up to 29.5% fewer tasks. These results support adaptive evaluation as an efficient and discriminative approach to assessing cybersecurity vulnerability awareness.
Figures & tables
Benchmark / Framework
# Tasks
# CWEs
Difficulty control
Cost control
CVE-Bench ( Wang et al., 2025c )
509
16
×
×
SEC-bench ( Lee et al., 2025 )
200
25
×
×
CyberGym ( Wang et al., 2026b )
1,507
88
×
×
CVE-Factory ( Luo et al., 2026a )
190
74
×
×
AutoBaxBuilder ( von Arx et al., 2026 )
40
11
✓
×
SecProbe (Ours)
353
151
✓
✓
Table 1: Comparison of cybersecurity benchmarks and frameworks by task count, CWE coverage, and support for difficulty and cost control.
Figure 2: Coding characteristics of the 353-task pool, showing (a) implementation languages, (b) source lines of code, and (c) implementation-file counts.
Figure 3: Security characteristics of the 353-task pool, showing (a) broad security categories, (b) distinct attack paths per task, and (c) grader scores before repair.
Figure 4: Detailed CWE coverage.
Figure 5: Cross-file security complexity.
Figure 6: Semantic map of the 353-task pool. (a) Task specifications embedded with Qwen3-Embedding-8B and projected with UMAP. Colour gives the security category of each task’s CWE, marker size its number of vulnerabilities, and shaded regions the eight application themes. (b) Cosine similarity of each task and each vulnerability to its nearest neighbour.
Figure 7: Vulnerability-level map of the 1,285 vulnerability descriptions. The first panel colours every vulnerability by security category; each remaining panel highlights one category against the others.
Expert ratings
Dimension
Evaluation focus
Mean ± Std.
≥4
Accept
Agreement
Security
CWE validity and exploit impact
4.60 ± 0.50
92%
94%
0.82
Attacks
Plausible and diverse attack cases
4.20 ± 0.70
85%
87%
0.74
Reference
Complete fix without regressions
4.70 ± 0.40
94%
96%
0.86
Realism
Coherent real-world repository
4.10 ± 0.70
82%
84%
0.72
Complexity
Cross-file diagnosis and repair
4.40 ± 0.60
89%
91%
0.79
Table 2: Human evaluation of generated cybersecurity coding tasks. Experts score each dimension from 1 (poor) to 5 (excellent).
Figure 8: Agent performance on the 353-task pool with Mini-SWE-Agent and Terminus-2 . Bars show strict pass rates; lines show mean normalized grader scores.
Review outcome
Count / Total
Share (%)
Patches assessed by both judges
353 / 353
100.0
Judge ratings within tolerance ( ∣r1−r2∣≤20 )
312 / 353
88.4
Referred for expert review ( ∣r1−r2∣>20 )
41 / 353
11.6
Expert-reviewed patches judged incomplete
29 / 41
70.7
Test-passing patches judged incomplete by an expert
8 / 100
8.0
Table 3: Patch-level review outcomes.
Figure 9: Adaptive evaluation versus random and one-shot baselines. Left shows the normalized maximum information gap as tasks accumulate; right shows the accepted tasks needed to reach a gap of 0.30.
Ability RMSE ↓
Policy
100 tasks
200 tasks
Tasks to reach RMSE ≤0.15↓
Adaptive
0.180
0.110
140
Random
0.240
0.140
181
Fixed order
0.260
0.145
194
Table 4: Ability-estimation efficiency on held-out agents. RMSE is relative to full-pool ability estimates.
Model
Ability rank Spearman ρ
Top-20 selection overlap (%)
Rasch (1PL)
0.96
80
2PL
Reference
Reference
Hierarchical 2PL
0.98
90
Table 5: Robustness of agent rankings and task selection to the IRT specification.
Setting
Pass rate (%) ↓
Mean score ↓
Ordered (%) ↑
OVERALL DIFFICULTY LEVELS
Easy
44.0±3.2
78.6±2.0
–
Medium
27.3±2.8
67.3±2.2
83.3
Hard
13.3±2.1
55.8±2.4
80.0
INDIVIDUAL COMPLEXITY EFFECTS
Repository complexity
−8.0±3.1
−5.4±2.2
73.3
Table 6: Difficulty-control validation.
Figure 10: An example of an agent failure of incomplete repair.
Primary failure mode
Description
Count
Share (%)
Incomplete repair coverage
Repairs some affected functions or paths while leaving others unprotected.
96
37.9
Incorrect diagnosis or localization
Targets the wrong root cause, security mechanism, or code location.
58
22.9
Missed application or state constraints
Fails to enforce authorization rules, permitted values, or lifecycle requirements.
52
20.6
Repair-induced regression
Introduces changes that break required behavior or component interfaces.
24
9.5
Other or unresolved
No effective repair, execution failure, or insufficient evidence for attribution.
23
9.1
Total nonpassing attempts
253
100.0
Table 7: Failure modes of unsuccessful repair attempts.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 11: Path containment leaves a caller-authorization bypass intact.
Figure 12: A version conflict masks missing nonce-consumption checks.
Figure 13: Restricted deserialization still permits out-of-policy values.