Coding Agents Aren't Enough! Evaluating an Enterprise Security Brain for Agentic Cloud Investigations
Organizations: Sola Security, Tel Aviv, Israel
Abstract
Cloud-security investigation is dominated by population tasks: which identities can read a data store, how many resources fail a control, what is reachable from another account. These resolve against a complete inventory, not a named object, so a partial answer to one is not a partial result but a different one. Coding agents can now be given read-only cloud credentials and asked to investigate directly, which raises the question of what a purpose-built security context layer still contributes. We evaluate the Sola Security Brain, a security intelligence layer whose relational substrate is resolved offline and whose security logic is evaluated against it at query time, against two coding agents operating the same live AWS environment through a read-only CLI, Claude Code and OpenAI Codex, over 28 investigation tasks. All three are scored by a blinded, tier-weighted, grounding-gated relative recall over one joint claim pool, so the scores share a denominator. The Sola Security Brain reaches 0.549 +/- 0.012 coverage against 0.340 +/- 0.011 for Claude Code and 0.281 +/- 0.006 for Codex, with the ordering identical in every grading draw. It leads 24 of 28 tasks from the cheapest model tier, at 17.8x and 20.8x lower cost per task than Claude Code and Codex. Beyond the aggregate, we describe an answer-level pattern we term sample-and-generalise: both agents enumerate a fraction of a large population, assert an unhedged universal negative, and disclose the sample size only in answer metadata rather than in the answer. In one task Claude Code reported that no bucket policies exist after checking four bucket families, in a sweep that sampled 40 of roughly 5,000 buckets, in an account where 65 buckets carry a wildcard-principal read grant.
Figures & tables
| Theme | Tasks | |
|---|---|---|
| Risk triage | 7 | q01 What are my riskiest resources right now, and why? q02 Who are my top 10 riskiest IAM identities? q03 What are my riskiest compute assets? q04 What are my riskiest storage assets? q05 What should I fix first? q06 What are my highest-risk resource groups, and why? q07 What are my riskiest secrets and encryption keys? |
| Secrets and keys | 4 | q08 Which secrets can more than one identity read? q09 Rank my secrets by how many principals can read them. q10 If a workload is compromised, which secrets does it expose? q11 Which internet-facing workloads can reach my secrets? |
| Access and reachability | 5 | q12 Which IAM roles can read my S3 buckets? q13 Who can decrypt my KMS keys? q14 Which principals have wildcard access, and what can they reach? q15 Which Lambda functions can be invoked, and by whom? q16 For my riskiest role, what can it reach? |
| Blast radius | 4 | q17 Rank my resources by blast radius. q18 If my most exposed resource is compromised, how far does it reach? q19 Walk me through my riskiest cluster. q20 Which identities give the widest reach if compromised? |
| Control failures | 4 | q21 Which resources fail public-access controls, broken down by control? q22 Rank my resources by number of failing controls. q23 Which of my S3 buckets are publicly accessible? q24 Which resources fail encryption-at-rest? |
| Exposure and cross-account | 4 | q25 What are my most internet-exposed assets? q26 Which exposed resources can reach sensitive data? q27 What are the attack paths to my privileged roles? q28 Which resources are reachable from another account? |
| Arm | Model | Access | Run conditions |
|---|---|---|---|
| Sola Security Brain | Claude Sonnet 4.6 | The context layer, queried | One process per task, five concurrent, warmed service, five-task warmup |
| Claude Code | Claude Opus 4.8 | Read-only credentials and the provider command-line interface | --max-turns budget |
| OpenAI Codex | GPT-6-Astra | Read-only credentials and the provider command-line interface | Claude Code’s system prompt verbatim, 900s wall timeout |
| # | Measure | Sola Security Brain | Claude Code | OpenAI Codex | Result |
|---|---|---|---|---|---|
| 1 | Coverage (relative recall) | 0.549 | 0.340 | 0.281 | CC, Codex |
| 2 | Duration, mean | 194s | 172s | 152s | slowest raw |
| 3 | Duration per unit coverage | 353s | 506s | 541s | / faster |
| 4 | Cost per task | $0.0073 | $0.1297 | $0.1517 | / cheaper |
| 5 | Cost per unit coverage | $0.013 | $0.382 | $0.540 | / cheaper |
| Theme | Sola Security Brain | Claude Code | OpenAI Codex | Gap vs CC | Tasks led | |
|---|---|---|---|---|---|---|
| Risk triage | 7 | 0.527 | 0.367 | 0.304 | ( ) | 7/7 |
| Secrets and keys | 4 | 0.528 | 0.425 | 0.290 | ( ) | 2/4 |
| Access and reachability | 5 | 0.474 | 0.300 | 0.274 | ( ) | 4/5 |
| Blast radius | 4 | 0.600 | 0.278 | 0.288 | ( ) | 4/4 |
| Control failures | 4 | 0.563 | 0.345 | 0.245 | ( ) | 3/4 |
| Exposure and cross-account | 4 | 0.638 | 0.313 | 0.270 | ( ) | 4/4 |
| # | Task | Sola Security Brain | Claude Code | OpenAI Codex | Gap |
|---|---|---|---|---|---|
| q17 | Rank my resources by blast radius. | 0.74 | 0.16 | 0.11 | |
| q20 | Which identities give the widest reach if compromised? | 0.81 | 0.34 | 0.40 | |
| q12 | Which IAM roles can read my S3 buckets? | 0.69 | 0.25 | 0.21 | |
| q23 | Which of my S3 buckets are publicly accessible? | 0.73 | 0.30 | 0.08 | |
| q28 | Which resources are reachable from another account? | 0.68 | 0.25 | 0.11 | |
| q03 | What are my riskiest compute assets? | 0.64 | 0.23 | 0.27 |
| Sola Security Brain | Claude Code | OpenAI Codex | |
|---|---|---|---|
| Mean | 194s | 172s | 152s |
| Suite total | 5,432s | 4,816s | 4,256s |
| Per unit of coverage | 353s | 506s | 541s |
| Sola Security Brain | Claude Code | OpenAI Codex | |
| Mean output tokens per task | 2,448 | 8,646 | 3,033 |
| Mean input tokens per task | not instrumented | not instrumented | 314,311 |
| Normalised cost per task | $0.0073 | $0.1297 | $0.1517 |
| Normalised cost, 28 tasks | $0.204 | $3.632 | $4.248 |
| $ per unit of coverage | $0.013 | $0.382 | $0.540 |
| Metered cost per task | — | $0.5849 | $1.0936 |