cs.CRSep 27, 2026

SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities

Authors: Xiaonan Luo, Yue Huang, Kehan Guo, Ping He, Chuan Zou, Chujie Gao, Lichi Li, Yuchen Ma, +4 more

Organizations: University of Notre Dame · Bake AI · Vanderbilt University · University of Pennsylvania · LMU Munich · University of Washington · Stanford University · Inria

Abstract

Assessing cybersecurity vulnerability awareness in coding agents requires evaluations that reveal capability gaps and remain informative as models evolve. Static benchmarks offer fixed coverage and difficulty, while scarce vulnerable repositories and costly expert authoring limit their renewal at scale. We introduce SecProbe, a framework for adaptive evaluation that combines Item Response Theory (IRT) with on-demand synthesis of repository-scale vulnerability-repair tasks. From observed performance, \textsc{SecProbe} estimates agent ability and identifies where additional evidence is most informative, selecting existing tasks or synthesizing new ones accordingly. As one use case, we construct 353 tasks spanning six programming languages and 151 CWE types and evaluate nine frontier models with two agent harnesses. Success rates peak at 28.33%, highlighting substantial gaps in vulnerability recognition and repair. Compared with random and one-shot baselines, \textsc{SecProbe} achieves comparable agent ability estimates while requiring agents to solve up to 29.5% fewer tasks. These results support adaptive evaluation as an efficient and discriminative approach to assessing cybersecurity vulnerability awareness.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MobileCybench: Evaluating Agent Vulnerability Discovery via Executable Probes

    Sep 21, 2026Andy K. Zhang, Ava Huang, Joey Ji +21Artificial Intelligence AgentsObfuscation

  2. CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities

    Jun 3, 2026Tianneng Shi, Robin Rheem, Dongwei Jiang +13CybersecuritySecurity Evaluation

  3. PatchBench: Evaluating AI Agents for Vulnerability Patching

    Sep 3, 2026Chihao Shen, Jiacheng Li, Aastha Mahajan +3Artificial Intelligence AgentsCrashes