Scalable Delphi: Large Language Models for Structured Risk Estimation
Organizations: CISPA Helmholtz Center for Information Security, Germany
Abstract
Quantitative risk assessment relies on structured expert elicitation to estimate unobservable properties. The Delphi method produces calibrated, auditable estimates but requires months of coordination and specialist time, placing rigorous risk assessment out of reach for most applications. We propose Scalable Delphi, adapting the classical protocol for LLMs with diverse expert personas, iterative refinement, and rationale sharing. Beyond lowering cost, this makes the assessment analyzable and dynamic. Rationales and revision histories record what each estimate rests on, information can be ablated to test which evidence matters, and the elicitation can be rerun with new evidence, changed assumptions, or adverse scenarios. Because target quantities are unobservable by construction, we design an evaluation framework based on necessary conditions any reliable estimator must satisfy: accuracy and calibration on verifiable proxies, and sensitivity to evidence. Agreement with expert panels and reasoning quality serve as corroboration. Across two domains (AI-augmented cybersecurity risk, ice-sheet contribution to sea-level rise), three benchmarks, and three reproduced expert studies, the estimates pass these tests: they improve systematically as evidence is added, agree with expert panels on most quantities, and correlate strongly with ground truth (Pearson r=0.91-0.98).
Figures & tables
| BountyBench | Cybench | CyberGym | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | MAE | ECE | MAE | ECE | MAE | ECE | |||
| GPT-5.1 | 0.97 | 5.7 | 1.6 | 0.97 | 4.5 | 4.1 | 0.91 | 5.6 | 2.6 |
| Opus-4.1 | 0.95 | 6.1 | 2.8 | 0.98 | 4.9 | 3.0 | 0.93 | 4.9 | 2.2 |
| GLM-5 | 0.96 | 6.3 | 2.2 | 0.97 | 4.2 | 4.0 | 0.94 | 5.0 | 2.1 |
| Kimi K2.5 | 0.96 | 5.9 | 2.5 | 0.98 | 4.0 | 2.3 | 0.92 | 5.8 | 3.4 |
| Llama-4 Maverick | 0.94 | 7.3 | 4.0 | 0.74 | 10.8 | 3.0 | 0.86 | 6.0 | 4.5 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| ID | Role | Background | Analytical Approach |
|---|---|---|---|
| A | Defensive Security Specialist | 10 years SOC experience, APT detection | Defender’s perspective, detection points |
| B | Malware Reverse Engineer | Anti-virus research lab | Bottom-up from technical implementation |
| C | AI/ML Security Researcher | PhD in CS, AI security | Systematic analysis of LLM assistance |
| D | Threat Intelligence Analyst | Former intelligence community | Observed attacker behavior patterns |
| E | Security Compliance Officer | CISSP, CISM certified | Framework-based, control effectiveness |
| ID | Role | Background | Analytical Approach |
|---|---|---|---|
| A | Greenland Outlet Glacier Dynamicist | 15 years tracking marine-terminating outlet glaciers with ICESat/ICESat-2 altimetry | Process-based ice-dynamics intuition anchored to observed retreat histories; emphasizes grounding-line retreat and dynamic discharge over SMB |
| B | Antarctic Ice Shelf Geophysicist | PhD in geophysics; field campaigns on Thwaites, Pine Island, Ross Ice Shelf | Reasons in terms of cavity geometry, basal melt rates, and grounding-line sensitivity; concerned with marine ice-sheet instability |
| C | Paleoclimate & Sea-Level Reconstruction Specialist | Builds Holocene and LGM sea-level databases; collaborates with GIA modellers | Constrains estimates using proxy reconstructions and regional GIA adjustments; anchors variability to past rates |
| D | Ice-Sheet Numerical Modeller | Computational glaciologist running ISMIP6-class ensembles (PISM, CISM, ISSM) | Cross-checks observations against model-plausible ranges; widens intervals for structural uncertainty |
| E | Remote Sensing / Mass Balance Observations Lead | Leads a mass-balance observations group; author on IMBIE 2018/2019 reports | Direct observational anchoring (GRACE/GRACE-FO, CryoSat-2, IMBIE); cautious extrapolation beyond observed record |
| Model | BountyBench | Cybench | CyberGym |
|---|---|---|---|
| GPT-5.1 | 0.1556 | 0.1463 | 0.1276 |
| Opus-4.1 | 0.1576 | 0.1465 | 0.1270 |
| GLM-5 | 0.1571 | 0.1472 | 0.1259 |
| Kimi K2.5 | 0.1564 | 0.1457 | 0.1274 |
| Llama-4 Maverick | 0.1591 | 0.1703 | 0.1310 |
| Reasoning | Human | GPT-5.1 | Opus-4.1 |
|---|---|---|---|
| Invokes supplied pre-LLM baseline | 15.4% | +12.2 | +9.0 |
| References values from previous steps | 4.1% | +4.3 | -3.3 |
| Names a technical ceiling | 25.3% | +55.1 | +37.1 |
| Names a non-technical ceiling | 13.7% | +28.1 | +20.3 |
| Judges task relevance of benchmark | 16.8% | +30.0 | -0.2 |
| Reasons about shape of distribution | 1.9% | +11.9 | -1.5 |
| Panel | Persona | Tier 1 | Tier 2 | Tier 3 | Tier 4 | Tier 5 |
| GPT-5.1, original | A | 560 | 570 | 580 | 600 | 620 |
| B | 180 | 190 | 205 | 210 | 220 | |
| C | 180 | 190 | 205 | 220 | 235 | |
| D | 171 | 176 | 187 | 198 | 209 | |
| E | 180 | 190 | 200 | 210 | 220 | |
| median | 180 | 190 | 205 | 210 | 220 |
| Task | Anchor | GPT-5.1 | Opus-4.1 | Human |
|---|---|---|---|---|
| 01 Number of Actors | 10 | 27 | 175 | 30 |
| 02 Attempts per Actor | 200 | 310 | 900 | 400 |
| Task | Tier | Anchor | GPT-5.1 | Opus-4.1 | Human |
|---|---|---|---|---|---|
| 01 Number of Actors | 1 | 10 | 12 | 17 | 22 |
| 01 Number of Actors | 5 | 10 | 27 | 176 | 52 |
| 02 Attempts per Actor | 1 | 200 | 218 | 220 | 354 |
| 02 Attempts per Actor | 5 | 200 | 312 | 970 | 524 |
| Run | Model | Calls | Input (M) | Output (M) | Reasoning | Cost (USD) | Wall (min) |
|---|---|---|---|---|---|---|---|
| BountyBench | GPT-5.1 | 600 | 0.54 | 5.98 | 5.82 | $118.76 | 34.0 |
| BountyBench | Opus-4.1 | 600 | 0.76 | 0.47 | – | $46.93 | 6.9 |
| Cybench | GPT-5.1 | 400 | 0.36 | 2.76 | 2.66 | $54.66 | 12.2 |
| Cybench | Opus-4.1 | 400 | 0.53 | 0.28 | – | $29.13 | 3.3 |
| CyberGym | GPT-5.1 | 460 | 0.50 | 3.14 | 3.02 | $62.16 | 19.5 |
| CyberGym | Opus-4.1 | 460 | 0.69 | 0.35 | – | $36.61 | 3.6 |
| Run | Model | Calls | Input (M) | Output (M) | Reasoning (M) | Cost (USD) | Wall (min) |
|---|---|---|---|---|---|---|---|
| Murray | GPT-5.1 | 55 | 0.23 | 0.07 | 0.07 | $1.69 | 5.0 |
| Barrett | GPT-5.1 | 100 | 4.39 | 0.32 | 0.23 | $10.94 | 10.7 |
| Barrett | Opus-4.1 | 100 | 4.56 | 0.10 | – | $75.59 | 10.8 |
| Bamber | GPT-5.1 | 270 | 4.19 | 0.62 | 0.56 | $16.99 | 23.5 |
| Bamber | Opus-4.1 | 270 | 4.29 | 0.13 | – | $74.01 | 13.2 |
| Subtotal | 795 | 17.66 | 1.24 | 0.85 | $179.22 | 63.3 |
| Agent | Description | Detect | Exploit | Patch |
|---|---|---|---|---|
| Claude Code | Anthropic’s terminal-based agentic coding tool (Claude 3.7 Sonnet) with built-in code navigation and editing tools; optimized for software engineering and defensive tasks such as patching vulnerabilities. | 5.0 | 57.5 | 87.5 |
| OpenAI Codex CLI: o3-high | OpenAI’s Codex command-line coding agent using the o3-high reasoning model; can read, modify, and execute code with strong tool support, emphasizing reliable code understanding and patch generation. | 12.5 | 47.5 | 90.0 |
| OpenAI Codex CLI: o4-mini | Lightweight OpenAI Codex CLI variant using the o4-mini model; lower-cost, faster reasoning with the same coding-agent tool interface, optimized for efficient patching. | 5.0 | 32.5 | 90.0 |
| C-Agent: o3-high | Custom research agent based on the Cybench framework, running raw bash commands in Kali Linux with iterative memory and reasoning, powered by OpenAI’s o3-high model. | 0.0 | 37.5 | 35.0 |
| C-Agent: GPT-4.1 | Custom Cybench-style agent using OpenAI GPT-4.1, operating purely through terminal commands without specialized coding tools; balanced offensive and defensive behavior. | 0.0 | 55.0 | 50.0 |
| C-Agent: Gemini 2.5 | Custom Cybench-style agent powered by Google Gemini 2.5 Pro Preview, interacting via shell commands in Kali Linux for vulnerability exploitation and patching. | 2.5 | 40.0 | 45.0 |
| Agent | Description | Success |
|---|---|---|
| Claude 3 Opus | Large-scale Anthropic frontier model evaluated within the Cybench structured bash agent. Operates via iterative terminal commands in a Kali Linux environment to solve professional-level CTF challenges, measuring autonomous end-to-end exploitation without subtask guidance. | 10.0 |
| Claude 3.5 Sonnet | Anthropic’s high-performing general-purpose model used as the core reasoning engine in the Cybench agent. Demonstrates strong unguided capability on CTF tasks solvable by expert humans within 11 minutes, indicating practical autonomous offensive skill. | 17.5 |
| Claude 3.7 Sonnet | Updated Anthropic model variant evaluated in unguided Cybench mode. Represents incremental improvements in reasoning and tool use over prior Claude 3.x models, though only unguided success is reported here. | 20.0 |
| Claude 4 Opus | Next-generation Anthropic flagship model evaluated on Cybench-style unguided CTF tasks. Higher success rate suggests substantially improved autonomous cybersecurity reasoning and execution compared to earlier Claude generations. | 38.0 |
| Claude 4 Sonnet | Anthropic mid-tier Claude 4 model evaluated in unguided Cybench tasks. Trades some raw capability for efficiency, but still demonstrates strong autonomous problem-solving on professional CTF challenges. | 35.0 |
| Claude 4.1 Opus | Refined Claude 4 Opus variant with improved reasoning robustness and continued gains in autonomous exploitation capability. | 38.0 |
| Agent | Description | Success |
|---|---|---|
| Anthropic Agent: Claude-Opus-4.6 | Anthropic’s strongest model (Feb 2026) with 1M-token context, agent teams, and a 14.5-hour task horizon, paired with Anthropic’s own agent scaffolding; excels at multi-step planning, debugging, and extended autonomous coding workflows. | 66.6 |
| SageAgent: GPT-5 | OpenSage’s self-programming agent framework—capable of generating agent topology, synthesizing tools, and managing hierarchical memory at runtime—combined with OpenAI’s GPT-5 (Aug 2025), a frontier model with state-of-the-art coding and dynamic reasoning. | 60.2 |
| Anthropic Agent: Claude-Opus-4.5 | Anthropic’s previous flagship (200K context, 64K output) with strong enterprise reasoning, programmatic tool calling, and tool-search capabilities, run through Anthropic’s agent scaffolding for autonomous multi-step tasks. | 50.6 |
| Claude Code: GLM-5 | Zhipu AI’s 744B-parameter MoE model (Feb 2026, 44B active, MIT-licensed, Ascend-trained) with native agent mode and strong coding abilities, driven by Anthropic’s open-source Claude Code CLI for terminal-based file editing and code execution. | 43.2 |
| Kimi Agent: Kimi K2.5 | Moonshot AI’s 1T-parameter MoE model (32B active, Jan 2026) with native multimodal understanding, 256K context, and Agent Swarm technology coordinating up to 100 specialized agents, paired with Kimi Code CLI for agentic task execution. | 41.3 |
| Anthropic Agent: Claude-Sonnet-4.5 | Anthropic’s high-performance mid-tier model (200K context, 64K output) noted for strong coding, agent building, and computer use, with reduced sycophancy; run through Anthropic’s agent scaffolding. | 28.9 |