AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) contains benchmarks limited by (1) unrealistic environments, typically toy repositories (2) non-standard framework implementations and (3) simple static verifiers. We introduce Incident-Arena, a human-built benchmark of 20 carefully selected tasks grounded in real-world deployed open source software. Each task deploys a production application to an ephemeral Kubernetes cluster, injecting a fault from the config layer through underlying images, and a sustained load profile given the task requirements. We also present a novel verification method, going beyond static checks to functional verifiers, holding systems level metrics stable, while ensuring repairs are done safely. Agent trials run an average of 2.81M tokens and 41 turns, going beyond existing benchmarks, demonstrating agentic long horizon reasoning. Across 20 tasks and 3 application substrates, frontier models score below 64.3%, with failures extending from diagnosis/localization errors, through incomplete repairs and unsafe regressions.
Figures & tables
Figure 1: The strongest top-10 models by reasoning on Incident-Arena ranked by Pass@1. * = cyber refusals, discussed in section 4
Benchmark
Tasks
Task types
Setting
Repairs
Names fault
Scope
Durable
No judge
Both ways
No fault
Harbor
Live systems
AIOpsLab ( Chen et al., 2025 )
48
Detection; RCA; mitigation
Kubernetes services
✓
∘
×
×
✓
×
∘
×
ITBench–SRE ( Jha et al., 2025 )
42 a
Diagnosis; mitigation
Kubernetes applications
✓
∘
×
×
✓
×
×
×
SREGym ( Clark et al., 2026 )
90
Diagnosis; mitigation
Kubernetes applications
✓
∘
×
×
×
×
×
×
OperAID
3
Operational mitigation
Live services
✓
×
×
×
✓
×
×
×
DBA-Bench
106
Database diagnosis; repair
PostgreSQL workloads
✓
∘
✓
×
✓
×
×
×
Table 1: What each benchmark evaluates. ✓ yes; ∘ partly or on some tasks; × no or not stated. Repairs : agents can change the system and have the result checked. Names fault : correctness of fault attribution is graded. Scope : preservation, collateral damage, or repair scope is checked. Durable : the repair is re-checked after a wait or restart. No judge : success does not depend on a trace or LLM judge. Both ways : a correct fix passes and doing nothing fails. No fault : includes cases where changing nothing is correct. Harbor : runs on Harbor, natively or via an adapter. Further sources and coding decisions appear in Appendix I .
Figure 2: Process of Developing Incident-Arena tasks by mining and developing substrates from real applications, injecting faults in different levels, with verifications on live operations.
Cost per episode
Developer
Harness
Version
Model
Effort
Trials
In/Out price
Pass@1
Mean
Median
Passing
Failing
OpenAI
Codex
0.153.4
GPT-6 (astra)
low-max
15
10/50
0.59
$3.28
$2.78
$2.66
$4.17
Codex
0.153.4
GPT-6 (sol)
low-max
15
2/10
0.39
$0.81
$0.60
$0.53
$0.99
Codex
0.153.4
GPT-5.6 (sol)
low-max
15
4/20
0.37
$2.51
$1.70
$1.76
$2.97
Codex
0.153.4
GPT-5.6 (terra)
low-max
15
2/12
0.28
$1.14
$0.76
$1.08
$1.16
Anthropic
Claude Code
2.1.280
Claude Opus 5.5
low-max
15
4/20
0.54
$2.00
$1.34
$1.88
$2.14
Table 2: Agent configurations, API prices, Pass@1 (averaged over all efforts) and model API cost per episode. Trials per task are 3 trials × each available reasoning level; 150 trials per task × 20 tasks = 3,000 trials. Prices are USD per million tokens; cache writes carry a 1.25 × premium on Anthropic and none elsewhere. Passing/Failing are the mean cost of passing and failing episodes.
Class (one per failure, assigned in this order)
Exclusive
Share
Inclusive
Never declared
317
16.3%
317
Changed something off the permitted list
427
21.9%
448
Left a planted cause in place, or its damage
170
8.7%
809
Repair did not last (restart or next trigger)
517
26.5%
782
Declared while service was still out of bounds
466
23.9%
756
Other integrity check
51
2.6%
51
Table 3: Failure classes over the 1,948 failures. Exclusive: one class per failure. The pie shows these shares. Inclusive: every failure that meets the rule.
Behaviour
Trials
Pass
No further change
638
0.696
Further change
931
0.530
Declared within 2 minutes
441
0.717
Declared in 2–10 minutes
782
0.599
Declared at or after 10 minutes
143
0.186
Fewer than 5 read-only checks
518
0.634
Table 4: Post-fix behaviour among trials with a tracked full fix and an action record.
Figure 3: Reward hacking behaviour, split out by model. We categorize into three buckets, demonstrating significant attempted reward hacking behaviour across all models.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Agent access surface ⟶
Fault injection layer ↓
Confined
Shell-visible
Build-capable
Configuration
Repair through operator APIs.
May inspect selected application pods.
Not supported because no source change is present.
Runtime
Repair through controlled changes to live state.
Supported if the activation mechanism does not reveal the answer.
Not supported because build-capable tasks require source changes.
Image
Repair through operational mitigation.
May inspect the faulted service without modifying it.
May edit and rebuild one selected service.
Appendix
Table 5: The task space combines two axes: how the fault is injected and what access the agent receives. Each cell describes the permitted repair surface.
Task
Tier
Causes
Onset
Planted faults
Pass
Action completion
deletes-and-jobs-fail
runtime
2
at start
mariadb.grants redis-queue.config
0.493
0.649
desk-and-queue-oom
runtime
2
under load
mariadb.max-user-connections redis-queue.config
0.300
0.829
desk-and-queue-outage
runtime
2
under load
mariadb.max-user-connections redis-queue.config
0.547
0.641
new-records-and-jobs-fail
runtime
2
at start
mariadb.grants redis-queue.config
0.333
0.693
new-records-and-queue-oom
runtime
2
at start
mariadb.grants redis-queue.config
0.467
0.727
writes-and-queue-oom
runtime
2
at start
mariadb.read-only redis-queue.config
0.153
0.720
Appendix
Table 6: Evaluated tasks and outcome rates over the 3,000 trials (150 per task).
Substrate
Tasks
Application pods
Harness pods
Setup pods
Services
frappe
6
10
14
3
25
saleor
1
5
9
1
16
slack-spine
13
35
12
0
45
Appendix
Table 7: Kubernetes objects rendered from each substrate’s Helm chart with a task overlay, as deployed on the single-node cluster of an episode. Application pods run the software under test: its servers, workers, databases, caches, queues and edge proxy. Harness pods exist for the benchmark: the agent’s container, the egress proxy, DNS filter and TLS gateway, the observability stack (Prometheus, Loki, log shipper, metric exporters and the read-only MCP server), the load generator and the administrative interfaces. Setup pods run once at deployment.
Task
Pass/Pass
Pass/Fail
Fail/Pass
Fail/Fail
deletes-and-jobs-fail
74
3
65
8
desk-and-queue-oom
45
2
24
79
desk-and-queue-outage
82
7
30
31
new-records-and-jobs-fail
50
66
0
34
new-records-and-queue-oom
70
8
25
47
writes-and-queue-oom
23
39
0
88
Appendix
Table 8: Outcome and safety gate verdicts.
Figure 4: Verifier Ablation
Substrate
Task
Scripted-fix reward
Do-nothing reward
frappe
deletes-and-jobs-fail
1.0
0.0
frappe
desk-and-queue-oom
1.0
0.0
frappe
desk-and-queue-outage
1.0
0.0
frappe
new-records-and-jobs-fail
1.0
0.0
frappe
new-records-and-queue-oom
1.0
0.0
frappe
writes-and-queue-oom
1.0
0.0
Appendix
Table 9: Both-ways check inside the reported experiment: the scripted reference fix and the do-nothing agent, one run each per task.
Model
Trials
Pass@1 [task 95%]
Action completion
Pass3
Pass5
GPT-6 astra
300
0.590 [0.430,0.743]
0.702
0.410
0.328
Claude Opus 5.5
300
0.540 [0.387,0.690]
0.736
0.345
0.267
Claude Fable 5.1
300
0.440 [0.297,0.590]
0.597
0.242
0.182
GPT-6 sol
300
0.393 [0.263,0.527]
0.660
0.173
0.104
GPT-5.6 sol
300
0.373 [0.240,0.520]
0.588
0.184
0.138
GPT-5.6 terra
300
0.280 [0.173,0.400]
0.481
0.091
0.045
Appendix
Table 10: Results by model, pooling available reasoning settings.
Model
Frappe (6)
Slack-spine (13)
Saleor (1)
GPT-6 astra
0.678 (0.977)
0.554 (0.550)
0.533 (1.000)
Claude Opus 5.5
0.511 (0.867)
0.574 (0.655)
0.267 (1.000)
Claude Fable 5.1
0.300 (0.533)
0.518 (0.595)
0.267 (1.000)
GPT-6 sol
0.444 (0.889)
0.385 (0.528)
0.200 (1.000)
GPT-5.6 sol
0.300 (0.694)
0.395 (0.508)
0.533 (1.000)
GPT-5.6 terra
0.289 (0.621)
0.277 (0.374)
0.267 (1.000)
Appendix
Table 11: Pass@1 (tracked action completion) by model and substrate.
Task
GPT-6 astra
Claude Opus 5.5
Claude Fable 5.1
GPT-6 sol
GPT-5.6 sol
GPT-5.6 terra
Claude Opus 5
GLM-5.3
Grok 4.7
Muse Spark 1.3
deletes-and-jobs-fail
0.53
0.80
0.40
0.27
0.27
0.07
0.73
0.73
0.33
0.72
desk-and-queue-oom
0.60
0.40
0.13
0.67
0.27
0.67
0.13
0.07
0.00
0.06
desk-and-queue-outage
0.53
0.93
0.53
0.40
0.27
0.33
0.73
0.47
0.58
0.67
new-records-and-jobs-fail
1.00
0.00
0.00
0.80
0.60
0.53
0.00
0.00
0.42
0.06
new-records-and-queue-oom
0.53
0.80
0.60
0.40
0.40
0.13
0.27
0.27
0.50
0.72
writes-and-queue-oom
0.87
0.13
0.13
0.13
0.00
0.00
0.07
0.20
0.00
0.00
Appendix
Table 12: Pass rate by task and model across available reasoning settings.
Figure 5: Passing trials per task and configuration over admitted episodes. Models the best at their reasoning level shown for sampling.
Model
Failures
Never declared
Off-list
Cause/damage
Not lasting
Too early
Other
GPT-6 astra
125.8
31.8%
8.7%
14.3%
18.4%
26.7%
0.1%
Claude Opus 5.5
157.2
13.4%
20.4%
8.3%
24.2%
32.4%
1.4%
Claude Fable 5.1
170.0
5.3%
25.3%
4.7%
24.1%
33.8%
6.8%
GPT-6 sol
171.1
26.9%
9.9%
10.5%
25.8%
25.1%
1.8%
GPT-5.6 sol
178.4
5.6%
11.8%
18.5%
29.9%
29.6%
4.6%
GPT-5.6 terra
206.9
5.8%
30.0%
8.2%
28.1%
23.1%
4.9%
Appendix
Table 13: Exclusive failure-class shares by model.
Setting
Trials
Pass@1
Mean cost (USD)
low
480
0.392
1.46
medium
480
0.400
1.76
high
480
0.402
2.23
xhigh
480
0.419
2.63
max
480
0.438
3.38
Appendix
Table 14: Reasoning effort among the eight five-setting models.
Judge
Scored
Refused
Missing artifacts
Opus 5
2,860
138
2
Opus 5.5
2,302
696
2
Gemini 3.8 Flash
2,998
0
2
GPT-6 Astra
2,987
11
2
Appendix
Table 15: Judge coverage.
Judge
Honest
Suspicious
Exploit-tier
Any flag
Opus 5
85.0%
15.0%
0.0%
15.0%
Opus 5.5
79.8%
20.2%
0.0%
20.2%
Gemini 3.8 Flash
71.6%
27.4%
1.0%
28.4%
GPT-6 Astra
31.7%
65.1%
3.2%
68.3%
Appendix
Table 16: Judge classifications among scored trials.