cs.SEOct 6, 2026

CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?

Authors: Hang He, Li Wang, Hao Chen, Yuchen Shao, Yuling Shi, Lisheng Wang, Peiyang Liu, Goose Lin, +4 more

Abstract

Static-analysis checker synthesis requires agents to interpret a defect specification, inspect a repository, implement analyzer-specific logic, and refine the checker through repeated compilation and analysis feedback. Existing coding-agent benchmarks focus on tasks such as patch generation or vulnerability detection and rarely assess whether an agent can develop a working checker in a repository from start to finish. We introduce CheckerBench, an executable benchmark of 300 tasks derived from 297 CVEs across 167 repositories, 85 CWEs, and five language ecosystems. Each task includes vulnerable and fixed revisions, a pinned analysis environment, and a checker scaffold. We further introduce CheckerLab, a common evaluation framework that independently rebuilds submitted checkers and measures vulnerable-fixed diagnostic contrast, patch localization, false positives, and tool use. Across 21 model-harness configurations and three independent repeats per configuration, mean Pass@1 is 32.30%, while the best reaches 45.33%. These results show that reliable, reusable checker development remains challenging for current coding agents.

Explore similar work

CardsList
  1. MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents

    May 5, 2026Jonathan Steinberg, Oren GalLanguage Model Safety EvaluationAI Coding Agents

  2. SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

    Date pendingBingchen Zhao, Dhruv Srikanth, Yuxiang Wu +1Reward HackingAI Coding Agents