The release decision for frontier AI systems increasingly relies on cyber capability benchmarks, yet public vulnerability benchmarks can expose agents to previously published advisories, exploits, and fixes, making it difficult to distinguish prior exposure from capability on unseen vulnerabilities. We evaluate open-weight and proprietary AI models on exploit generation, vulnerability repair and subsequent attacks in five nonpublic software environments, including vulnerabilities we privately disclosed while they remained unpatched. Researcher-developed and reviewed deterministic graders, not LLM judges, determine task scores. Comparisons with 209 disclosed vulnerabilities and cryptographic challenges reveal substantial variation across systems and vulnerability types. Repair scores exceed attack scores in two nonpublic environments and fall below them in three. Passing an initial security test is also insufficient: another exploit succeeds in 92 of 524 non-independent defender test intervals after the initial exploit is stopped. These results motivate vulnerability-specific attack-repair comparisons and subsequent resistance tests.
Figures & tables
Fig. 1: Study workflow. The repeated interactions use continuing agent sessions; they are a separate experiment, not a follow-up to every repair in Q1.
Question
Unit, outcome and weighting
Paired success
Same vulnerability; repair minus attack success; equal vulnerability weights.
Task scores
Task-specific repair minus attack score; equal system weights within each information condition.
Information
Specific minus limited score for tasks with both outcomes.
Review
Same environment/repetition; structured minus generic review score.
Residual tests
Test interval with no recorded software change; another exploit succeeds after the original is stopped; equal interval weights.
Subsequent attack
Defender-first round with source-code changes and a winner; attacker wins; equal recorded rounds.
TABLE I: Observation units and outcome measures. Collections are analysed separately.
Environment
Attack
Repair
D−A
ENV-A
.333
.500
+.167
ENV-B
.324
.379
+.055
ENV-C
.190
.000
-.190
ENV-D0
.333
.000
-.333
ENV-D1
.603
.083
-.520
TABLE II: Nonpublic task scores: 14 paired comparisons per environment except ENV-B (12). Systems have equal weight within each information condition; the two conditions also have equal weight.
Harness / model
n
Attack
Repair
D−A
Claude Code / Opus 5
10
.600
.217
−.383
Codex / GPT-5.6 Sol
10
.133
.233
+.100
Codex / Daybreak Blue alias
10
.133
.200
+.067
OpenCode / Gemini 3.8 Flash
10
.044
.217
+.172
OpenCode / Kimi K3
9
.519
.130
−.389
OpenCode / Qwen3.8 Max 0902
10
.533
.000
−.533
TABLE III: Nonpublic attack and repair scores by agent system. n is the number of complete environment-information cells after the common execution-status screen. Cells have equal weight; the two roles have distinct scoring criteria. Full harness versions and missingness are in the supplement.
System
Attack
Repair
Δ , pp
OpenCode / GPT-5.5
201
178
−11.0
Claude Code / Opus 4.7
167
120
−22.5
OpenCode / GPT-5.3-Codex
161
175
+6.7
TABLE IV: Attack and repair success counts on the same 209 disclosed vulnerabilities. Δ is repair minus attack success in percentage points.
A. Disclosed memory-safety vulnerabilities
GPT-5.5
Opus 4.7
GPT-5.3-Codex
CWE
Weakness
n
A
R
Δ
A
R
Δ
A
R
Δ
125
Out-of-bounds read
113
111
101
−8.8
88
69
−16.8
97
98
+0.9
787
Out-of-bounds write
30
28
26
−6.7
23
18
−16.7
23
29
+20.0
416
Use after free
20
17
15
−10.0
20
11
−45.0
14
14
0.0
457
Uninitialized variable
10
10
9
−10.0
6
3
−30.0
6
7
+10.0
908
Uninitialized resource
5
4
3
−20.0
3
2
−20.0
3
4
+20.0
TABLE V: Attack and repair by weakness type. In panel A, A and R are success counts on the same vulnerabilities and Δ is repair minus attack in percentage points. In panel B, A and R are native-score means and Δ=R−A ; n counts complete system/information comparisons for nonpublic tasks and complete system comparisons for cryptographic tasks. Nonpublic group names are withheld to protect undisclosed flaws. Classifications are model-assisted; small groups describe these cases only.
Repair − attack
Repair: specific − limited
Attack: specific − limited
Structured − generic review
Environment
n
Δ
n
Δ
n
Δ
n
Δ
ENV-A
3
+1.000
2
.000
1
.000
2
.000
ENV-B
3
.000
2
+.500
1
.000
2
+.417
ENV-C
3
−.222
2
.000
1
+.667
2
.000
ENV-D0
4
−.167
2
.000
2
−.333
2
.000
ENV-D1
2
−.333
1
.000
2
−.333
1
.000
TABLE VI: The 50-assignment Codex 0.144.4/GPT-6 Sol experiment on the five nonpublic environments. Each cell reports a matched mean score difference Δ and the number of complete pairs n ; suite-error and unmeasured outcomes are excluded. Repair versus attack combines both information conditions. Information compares source-location hints with limited instructions; review compares structured with generic continuation. Positive values favour the first named condition.
Environment
Rounds
Attacker wins
Defender wins
ENV-A
49
47
2
ENV-B
21
21
0
ENV-C
11
11
0
ENV-D0
14
14
0
ENV-D1
22
22
0
Total
117
115
2
TABLE VII: Subsequent attacks after a defender submits source-code changes: 117 rounds across five environments. Attacker wins include an exploit succeeding or the repair failing build or functionality requirements. Rounds without a new code change are excluded.