Coding agents are now proficient enough to generate complex software applications from a single prompt. As their capabilities have grown, human oversight has increasingly shifted from line-by-line code review toward hands-off evaluation of outcomes. However, recent studies have shown that such a transition exposes a critical risk: functional correctness alone does not guarantee a secure implementation. Despite growing attention to code security, training safer coding agents remains challenging because reliable security supervision is difficult to obtain at scale from real-world repositories. We introduce AuraForge to synthesize and validate executable security tests for training secure coding agents. Our approach combines attack-oriented test synthesis, language-extensible task construction, and safeguards against reward hacking. Using AuraForge, we construct AuraGym, a multi-language and multi-CWE executable training gym: 679 executable feature-implementation tasks from 344 real-world repositories across Python, JavaScript, and TypeScript, covering 177 CWE categories. On the subset with human-written security tests, AuraForge produces about 3 times as many test cases on average and reduces the false-positive rate by 83.23%, allowing alternative secure implementations to receive correct supervision. Training Qwen3.5-4B with synthesized security tests gains larger improvements than human-written security tests (average 19.7 FuncPass and 6.2 SecPass vs. 14.9 FuncPass and 4.4 SecPass) on three languages. These results demonstrate that AuraForge provides more diverse and reliable security supervision to train secure coding agents.
Figures & tables
Figure 1: AuraForge . Starting from real repositories and historical security fixes, The Forge derives feature requests without leaking security hints, and constructs language-extensible execution environments. The Aura synthesizes security tests and validates them. The Seal removes implementation-leakage channels. Finally, it produces 679 executable instances from 344 repositories across 177 CWE categories and 3 languages.
Dataset
Task setting
Instances
Repos.
CWEs
Languages
General SWE agent training datasets
SWE-Gym
Issue resolution
2,438
11
—
Python
SWE-smith
Synthetic bug repair
≈ 50,000
128
—
Python
SWE-rebench
Issue resolution
21,336
3,468
—
Python
SWE-rebench V2
Issue resolution
32,079
3,617
—
20 languages
Secure-coding evaluation benchmarks
Table 1: Comparison of general SWE training datasets and secure-coding benchmarks. Instances denote tasks rather than trajectories. A dash indicates that repository counts are not applicable or CWE coverage is not reported.
Figure 2: Division between the shared core and the language adapter in AuraForge .
Figure 3: Retrieval rates of SusVibes runs, per model and channel (runs pooled over scaffolds), with the repository history present and the network reachable.
Figure 4: Comparison of human-written AuraGym h and synthesized AuraGym : (a) task and repository counts; (b) distinct CWE categories by language; and (c) average security tests per task.
Table 6
Model
Size
SusVibes (186)
SusVibes tsjs (82)
FuncPass
SecPass
FuncPass
SecPass
GPT 5.6 Sol
–
84.40
21.50
92.70
18.30
Muse Spark 1.3
–
79.57
15.59
82.90
19.50
GLM 4.7 Flash
30B-A3B
24.73
4.30
24.39
4.88
Nemotron 3.5
30B-A3B
15.05
3.23
25.61
3.66
Qwen 3.5 4B
4B
12.90
2.10
12.19
0.00
Table 3: Functional and security performance on the SusVibes (Python) and newly created TypeScript/JavaScript evaluation sets. SusVibes is the original Python version from Zhao et al. (2026) and SusVibes tsjs is the test split for the new TS and JS. Nemotron 3.5 is the Lightning version.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Selection stage
Candidates
Removed
Repos.
Initial JS/TS export
3,189
—
—
Implementation, test, and patch-size filters
967
2,222
—
Candidate deduplication
909
58
514
Fixing-commit date ≥ 2024-01-01
311
598
—
Identified Node.js major ≥ 18
255
56
127
Repository primary language is JS/TS
249
6
—
Appendix
Table 4: MoreFixes filtering. Removed denotes the reduction from the preceding row; dashes indicate unreported counts.
Selection stage
Candidates
Removed
Repos.
Retrieved unique commits
3,742
—
—
Patch structure, parent, and metadata filters
2,371
1,371
—
Patch includes JS/TS
2,204
167
—
Patch includes test files
1,389
815
489
Fixing-commit date ≥ 2024-01-01
896
493
—
Identified Node.js major ≥ 18
817
79
198
Appendix
Table 5: GHSA filtering and cross-source deduplication. Removed denotes the reduction from the preceding row; dashes indicate unreported counts.
Channel
How the reference is obtained
Countermeasure
git history
git show , git log -p , or git checkout on a pre-mask revision
history stripped; the task ships as a single commit
local copy
reading a copy already on disk: site-packages , a build directory, a vendored tree
image sanitized
internet
fetching the upstream file or archive from a code host
command filtering while solving
package manager
installing the upstream package at its fixed version, then reading its source
command filtering while solving
Appendix
Table 6: Retrieval channels observed in the audit and how each is closed.
Version
Language
# Instances
# Repos
# CWEs
Avg. # target files
Avg. # func. tests
Avg. # security tests
AuraGym
Python
476
266
148
1.65
73.22
7.84
JavaScript
46
36
31
1.70
32.22
11.54
TypeScript
161
46
71
1.99
20.52
10.20
Overall
679
344
177
1.72
53.54
8.62
AuraGymh
Python
202
109
98
1.39
89.51
2.38
JavaScript
53
42
32
1.79
32.22
3.80
Appendix
Table 7: Statistics of the synthesized AuraGyms and the human-written comparison set AuraGymh . The numbers of functional tests, security tests, and target files are reported as per-instance averages. An instance whose repository mixes JavaScript and TypeScript is counted under both languages.
Language
Model
Setting
# Traj.
# Inst.
Coverage
Avg. turns
Avg. tokens (K)
Python
DeepSeek
generic
392
163
34.24%
39.75
52.15
security-hint
77
77
16.18%
46.53
59.84
overall
469
226
47.48%
40.86
53.42
Muse Spark
generic
462
152
31.67%
42.91
35.59
security-hint
161
112
23.33%
47.22
43.27
overall
623
264
55.00%
44.02
37.57
Appendix
Table 8: Statistics of the trajectories collected by DeepSeek-V4.1-Flash and Muse Spark 1.3 that are functionally correct and secure. Coverage denotes the percentage of instances represented by at least one retained trajectory within each model’s instance pool (679 for DeepSeek-V4.1-Flash; 710 for Muse Spark 1.3, which was collected on an earlier version of AuraGym ). Token counts are reported in units of 1,024 tokens.
Kind: the question the test fails
Category: the test…
Example
Human
Synth.
Rejects a secure solution (false positive): the test demands more than that the attack has no effect
Fix-bound : does the test demand a detail of the historical fix that security leaves free?
…wants the exact form of a correct result
A converter must render an over-long heading tag as <h6> , as the fix does; a solution that renders it as plain text stops the attack just as well (Appendix F.6.2 ); likewise a query must have the exact SQL text the fix sends (Appendix F.6.3 ).
34
1
…wants one particular safe reaction where another exists
A model loader must refuse a file in safe mode; a solution instead loads it through a restricted loader that cannot run code (Appendix F.6.1 ).
20
2
…wants the error’s wording or type
The error must say “arbitrary code execution”; a solution refuses the same file with a different message (Appendix F.6.1 ).
7
1
…calls a helper only the fix defines
The test imports a validator function the fix added; a solution that validates inline has no function of that name.
7
0
all fix-bound
68 (32)
4 (3)
Appendix
Table 9: Causes of wrong verdicts. Every wrong verdict behind Table 2 is assigned, from the adjudication record, to the category of test error that produced it; the categories are grouped into kinds, each named by the question the test fails. Counts are wrong verdicts (human / synthesized): secure solutions rejected for the false-positive kinds, vulnerable solutions accepted for the false-negative kinds; kind rows give the total, with tasks in parentheses. Each example is a real case; those in Section F.6 are referenced.
Suites that…
Human
Synthesized
share of tasks
FPR
share of tasks
FPR
assert the attack’s effect is absent
48%
12.0%
87%
3.0%
assert the outcome the fix produces
52%
20.7%
13%
2.0%
Appendix
Table 10: Assertion audit of both suites on the 183 tasks with a secure solution. Each suite’s security tests were classified, blind to the grades, by what they assert: that the attack’s effect is absent, or the outcome the historical fix produces; mixed suites are counted by their majority. FPR is over the secure solutions of the tasks in that class.
Family
Solutions
Tasks
FP (secure rejected)
FN (vulnerable accepted)
Human
Synth.
Human
Synth.
Authentication / authorization
495
66
10/75
4/75
5/420
3/420
Input validation / other
368
45
15/74
8/74
1/294
0/294
Resource exhaustion / ReDoS
334
54
19/99
6/99
10/235
6/235
Injection (SQL / command / code)
255
36
7/51
0/51
0/204
0/204
Path traversal / link following
244
33
4/55
0/55
2/189
8/189
Appendix
Table 11: Wrong verdicts by vulnerability family. A solution counts once under each family its task’s CWE ids fall in; 48 solutions carry no informative CWE and are in no family. Cells are wrong verdicts over the solutions of that kind (secure for FP, vulnerable for FN). Most secure-side denominators are under 100, so these are counts to read, not rates to rank.