Localizing where an LLM-based agent first fails to uphold security during a repair can show which stage of its workflow needs an additional safeguard. This is difficult for silent failures, which are patches that pass syntactic and functional checks but still contain a security vulnerability. Because such patches give no observable failure signal, existing failure attribution methods, which rely on observed task failures and labelled failure steps, are less suited to them. We propose Security Awareness Gap Evaluation (SAGE), a trace-based method that combines an assessment of the security reasoning recorded at each turn with the reconstructed code history to identify the earliest turn at which a repair diverges from the task's security intent. We evaluate SAGE on 95 confirmed silent failures drawn from 3,684 repair traces produced by six agent frameworks and six base models on SecurityEval and CVEfixes. SAGE assigned an origin in 93 cases. Most origins were an unaddressed security requirement or an inadequate defence choice, and only five coincided with the code change itself. When the agent introduced the vulnerable code, the origin preceded the write in 14 of 19 cases. Repeated scoring and a second judge reproduced the origin type more consistently than the exact turn, and agreement was lowest for traces that kept only the final file.
Figures & tables
Figure 1. Examples of the two cases in silent failure origin localization.
Figure 2. Overview of the research design.
Framework
Type
Architecture
Selection rationale
CAMEL ( Li et al., 2023 )
Multi
Two-role dialogue
Planner–engineer roles with explicit handoffs
AG2 ( Wu et al., 2024 )
Multi
Group chat
Round-robin chat with a security reviewer
Deep Agents ( LangChain, 2026 )
Multi
Planner–subagent
Hierarchical decomposition and delegation
Aider ( Gauthier, 2024 )
Single
Edit–feedback loop
Diff edits with lint and test feedback
MiniSWE-Agent ( Yang et al., 2024 )
Single
ReAct loop
Reasoning and tool use, no inter-agent messages
OpenHands ( Wang et al., 2025b )
Single
Edit–feedback loop
Sandboxed file editing and shell execution
Table 1. Selected agent frameworks.
Framework
Excluded runs
Valid traces
Passed L0–L3
Failed L0/L1
SF candidates
AG2
337
563
459
49
55
Aider
156
744
624
58
62
CAMEL
232
668
345
292
31
Deep Agents
378
522
443
29
50
MiniSWE-Agent
383
517
424
27
66
OpenHands
230
670
549
54
67
Table 2. Silent Failure Dataset construction by framework.
Metric
Value
Observed agreement
0.88
Cohen’s κ
0.73
Gwet’s AC 1
0.85
Table 3. Inter-rater reliability on the silent failure candidates subset.
Figure 4. Origin localization outcomes. (a) Code provenance against origin type for all 95 confirmed silent failures; (b) Distance in canonical turns between the origin turn and the introducing write for the 19 cases in which the agent introduced the vulnerable code, one dot per case.
Repeated scoring
Cross judge
All four ratings
Eligible cases ( n )
57
50
48
Type agreement (%)
96.49
92.00
89.58
Exact-turn agreement (%)
87.72
78.00
75.00
Within ±1 turn (%)
91.23
84.00
81.25
Origin type α
0.96
0.86
0.92
Origin turn α
0.82
0.59
0.74
Table 6. Reliability of SAGE localization outputs.
Figure 5. Reliability of the judge-assigned profile scores at the origin turn.
Figure 6. Agreement by confirmed deficiency and by reconstruction condition.
Framework
n
Provenance
Introducing write
CAMEL
4 ∗
4 (100.00%)
4 (100.00%)
AG2
32
28 (87.50%)
28 (87.50%)
Deep Agents
13
9 (69.23%)
9 (69.23%)
Aider
22
18 (81.82%)
18 (81.82%)
MiniSWE-Agent
15
13 (86.67%)
11 (73.33%)
OpenHands
9 ∗
8 (88.89%)
7 (77.78%)
Table 7. Agreement between the localization output and annotator on the 95 confirmed SF cases, by agent framework (a) and by base model (b). ∗ Groups with fewer than ten cases.
Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Recent agentic AI approaches have shown promising results in automated program repair. However, vulnerability repair demands richer program context than general bug repair - context that security engineers routinely assemble in practice but that existing agentic approaches do not engineer. We identify three critical gaps: code-structure context capturing cross-file data flows and memory operation patterns, runtime-execution context revealing crash semantics and memory origins, and commit-history context recovering how fragile code patterns were introduced. We present AgenticRepair, an agentic vulnerability repair framework that addresses the gaps through multi-faceted program context engineering. AgenticRepair orchestrates three specialized LLM subagents to engineer the contexts, which are then embedded into the memory of a dedicated repair subagent for context-conditioned patch synthesis. Evaluated on SEC-Bench comprising 300 real-world instances with sanitizer-based patch verification, AgenticRepair achieves a 73% success rate, substantially outperforming the strongest baseline by 29%. Our ablation study confirms that the three context facets are mutually complementary, and that multi-agent scaffolding and base-model capacity each play an essential role. Collectively, these findings establish multi-faceted program context engineering as a promising design direction for agentic vulnerability repair.
Michael Fu, Qiyue Mei, Patanamon Thongtanunam +1
the School of Computing and Information Systems, The University of Melbourne, Melbourne, Australia. · Kla Tantithamthavorn is with the Faculty of Information Technology, Monash University, Melbourne, Australia.
The advent of agentic vulnerability detection is already becoming a watershed moment for software security. Audits conducted entirely by autonomous LLM agents are uncovering critical vulnerabilities in fundamental software underpinning digital society. Many of these vulnerabilities remained masked for years, surfacing only now with AI agents. Yet the reasoning behind these discoveries remains alarmingly opaque and unvalidated. What assumptions did the agent make about a function's inputs when it deemed that function to be secure? Failures in reasoning and incorrect assumptions can lead to missed vulnerabilities and reduce trust in agentic analysis. We propose a security-specification-first paradigm that (1) exposes the agent's tacit assumptions explicitly as security specifications and (2) continuously refines those specifications via runtime falsification. We realize our approach in Code-Augur, a novel harness for agentic vulnerability detection. Given a codebase, Code-Augur analyzes each component of the system for vulnerable code. When it deems a component to be secure, it commits the local invariants behind that judgment as in-source assertions. In parallel, Code-Augur leverages a guided fuzzer to attempt to falsify those assumptions. When the fuzzer triggers an assertion, this either reveals a genuine vulnerability or a flawed specification to refine. In both cases, this process grounds the agent's understanding, aligning its view of code intent with how the code actually behaves. On real-world subjects, Code-Augur effectively leverages security specifications to detect more vulnerabilities than other state-of-the-art agents. Additionally, Code-Augur found 22 new vulnerabilities in key open-source projects. Compared to curated specialized models like Claude Mythos, Code-Augur offers effective agentic vulnerability detection built on widely available LLMs like Sonnet and DeepSeek.
Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world software. Generally available agents can already aid attackers, who only need to find one exploitable weakness, while defenders must continuously identify and patch all vulnerabilities across fast-growing codebases. Stronger defensive agents would help close this gap, yet the scarcity of security training data with reproducible build and execution environments remains a bottleneck. We present CyberForge, a framework that synthesizes executable, repository-level security training data by injecting vulnerabilities into real C/C++ projects. It validates each instance dynamically: the injected build must pass the project's unit tests, and generated proof-of-vulnerability (PoV) must trigger on the injected build and not on the clean one. CyberForge is not limited by the availability of disclosed vulnerabilities, therefore it can scale in comparison to data augmentation techniques which rely on historic CVE data. The resulting corpus holds 1034 validated vulnerabilities across 80 projects and 63 weakness categories, with edit locality similar to real CVE patches under a real-versus-real noise floor. Fine-tuning on trajectories collected over this corpus improves SEC-bench patch repair by +3.3 to +14.7 points, in all six configurations of three model scales and two teachers, with the 31B student reaching its GPT-5.4-mini teacher, 72.7% against 74.0%. These gains generalize out of distribution to PatchEval, a corpus containing other programming languages, where every configuration also improves and the 31B student passes its teacher.
Amine Lbath, Manan Suri, Aurelien Delaitre +4
National Institute of Standards and Technology · Université Grenoble Alpes, CNRS · University of Maryland, College Park