APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry
Organizations: Zhongguancun Laboratory
Abstract
Large language model (LLM) agents could help security operations centers (SOCs) investigate advanced persistent threats (APTs) by turning weak leads into evidence for intrusion scoping and response. Yet success under one telemetry setting does not establish robustness to changes in log collection, retention, or sampling. We introduce APTInvestBench, a benchmark for evaluating cross-telemetry robustness in autonomous APT investigation. It comprises 370 cases across seven SOC-inspired conditions, derived from 56 report-informed attack reconstructions with 16.4 million log records. Agents investigate unverified leads and submit reports with record-level citations. Fixed action-level support requirements track sufficient evidence across available logs, query returns, and formal citations, separating telemetry limitations from acquisition and reporting gaps. Across eleven LLMs, agents acquire sufficient evidence for 44.3% of recoverable attack actions on average, while formal citations support only 25.0%. More importantly, aggregate coverage can conceal substantial instability: from Full to endpoint-only telemetry, coverage declines by only 1.6 percentage points, yet 35.5% of previously covered actions lose sufficient citation support despite remaining recoverable. Across four frameworks, such losses persist even when registered supporting records remain unchanged. APTInvestBench provides reusable investigation environments and diagnostic evaluation for identifying these gaps and developing more reliable defensive agents.
Figures & tables
| Benchmark | Task | Scenarios | Data | Construction | Telemetry variation | Gap diagnosis | Support-matched comparisons |
| ExCyTIn-Bench ( Wu et al., 2026b ) | Investigation QA | 8 | 923.2K rows | Attack reenactment | Sources / time | ✗ | ✗ |
| DiagChain ( Liu et al., 2026 ) | Attack-chain reconstruction | 23 | 12.8K cards | Public-log integration | Source profiles | ✓ | ✗ |
| SynthChain ( Tan et al., 2026 ) | Supply-chain reconstruction | 7 | 593.7K records | Attack reenactment | Sources / sampling | ✗ | ✗ |
| RAG-SIA ( Cadet et al., 2026 ) | Forensic QA / reconstruction | 94.9K logs | Public data + reenactment | Sources / context | ✓ | ✗ | |
| APTInvestBench | Multi-step APT investigation | 56 | 16.4M records | Report-informed reconstruction | Seven SOC views | ✓ | Common actions + unchanged witnesses |
| Metric | Rate | Interpretation |
| Availability | Fraction of reference actions supportable in the view. | |
| Acquisition | Fraction of recoverable actions supported by query returns. | |
| Citation coverage | Fraction of recoverable actions supported by formal-stage citations. | |
| Citation retention | Fraction of actions with acquired support that retain sufficient formal-stage citations. |
Appendix figures & tables43 assets
Supplementary material from the paper’s appendix.
Appendix
| Inventory item | Complete resource (56 scenarios) | Controlled study (10 scenarios) |
| Scenarios | 56 | 10 |
| Physical records | 16,431,768 | 1,831,162 |
| Malicious records | 10,507 | 1,473 |
| Benign background records | 16,421,261 | 1,829,689 |
| Malicious share (%) | 0.0639 | 0.0804 |
| Records per scenario (k), range | 19.6–527.6 | 109.0–231.1 |
| Scenario ID | Threat label | Logs (k) | Mal. (%) | Acts | Attack (min) | Window (h) |
|---|---|---|---|---|---|---|
| Controlled-study subset: ten scenarios | ||||||
| 2014-06-30-dragonfly-energetic-bear | Dragonfly | 109.0 | 0.0670 | 7 | 90.7 | 16 |
| 2018-04-25-apt-c-01 | APT-C-01 | 231.1 | 0.1653 | 5 | 268.0 | 34 |
| 2019-05-28-apt27 | APT27 | 166.3 | 0.0505 | 13 | 318.5 | 44 |
| 2021-12-23-apt-q-38-donot | Donot | 212.5 | 0.0339 | 11 | 1,422.8 | 34 |
| 2023-04-07-muddywater | MuddyWater | 200.4 | 0.0200 | 5 | 304.4 | 30 |
| Condition | Benchmark definition | Operational motivation | Sources |
| Full telemetry | Retain all generated telemetry as the reference record set. | SOC investigations correlate host, identity, network, DNS, and other records to establish incident scope. | Joint logging guidance ( Australian Signals Directorate’s Australian Cyber Security Centre and others, 2024 ) |
| Endpoint-only | Retain endpoint-origin records, including network facts recorded by endpoint sensors. Remove independent network sources. | Host monitoring and network-flow logging are configured separately, allowing asymmetric source coverage. | Sysmon ( Russinovich and Garnier, 2026 ) VPC Flow Logs setup ( Google Cloud, 2026a ) |
| Network-only | Retain network-origin records and remove endpoint sources. | Devices can lack endpoint agents or usable local logs. Network monitoring supplies visibility into such gaps, also documented in BRICKSTORM intrusions. | Joint logging guidance ( Australian Signals Directorate’s Australian Cyber Security Centre and others, 2024 ) BRICKSTORM report ( Yoder et al., 2025 ) |
| Tiered retention | Use source-specific history windows: endpoint 8 h, network/application 6 h, and perimeter 4 h. | Log types and licenses have different retention periods. Long intrusions can outlast retained evidence of initial access. | Entra retention ( Microsoft, 2026 ) BRICKSTORM report ( Yoder et al., 2025 ) |
| Group sampling | Retain network/perimeter event groups with seeded probability , keeping records within a group together. | Flow selection and production flow-log sampling reduce collection and processing demands. Event groups preserve within-group associations in this controlled instance. | RFC 7014 ( D’Antonio et al., 2013 ) VPC Flow Logs sampling ( Google Cloud, 2026b ) |
| Targeted suppression | Remove records intersecting every registered sufficient witness of an initially supported target action. | Selective log deletion conceals attacker activity. The Volt Typhoon advisory documents this behavior in confirmed intrusions. | Volt Typhoon advisory ( Cybersecurity and Infrastructure Security Agency et al., 2024 ) |
| Reference action | Required observations | Binding and compatibility |
| APT27 process creates error2.aspx | ECAR process creation and file creation. | Process objectID equals file actorID , linking process identity and start to the write. |
| APT-C-01 recurring HTTP communication | At least three distinct HTTP observations spanning at least 60 seconds. | Compatible source and host, the required action binding, and a largest-to-smallest gap ratio of at most two. |
| Record | UTC timestamp | Gap from previous (s) |
| 10:48:09.820287 | – | |
| 10:52:09.473867 | 239.653580 | |
| 10:55:00.606433 | 171.132566 |
| Check | What is established | Scope of the evidence |
| Source-byte and binding checks | A public record maps to archived content and the inspected event identity. | Traceability of inspected bindings. Rule sufficiency also depends on the registered fact requirements. |
| Oracle, empty, and report transformations | Scoring reaches expected extremes and preserves under citation-preserving transformations. | Implementation behavior under the declared rules. |
| Complete-content checks | The inspected witness appears in complete persisted tool output. | Certified records and saved outputs. The main acquisition endpoint uses returned IDs. |
| Bounded assertion review | Selected assertions are inspected against their own citations for activity, linkage, timing, and effects. | A case study with unresolved judgments retained. Whole-report accuracy remains unevaluated. |
| Framework / models | Version | Effort | Response (K) |
| Codex / MiniMax models | 0.145.0 | High | 32.768 |
| Claude Code / MiniMax-M3 | 2.1.228 | High | 32.000 |
| OpenCode / MiniMax-M3 | 1.18.4 | Not set | 32.768 |
| Ref-ReAct / MiniMax-M3 | 0.2.0 | Not set | 32.768 |
| Codex / DeepSeek-V4-Pro | 0.145.0 | High | 16.384 |
| Codex / GPT-5.5, GPT-5.6-Luna, GPT-5.6-Terra, GPT-5.6-Sol, GPT-6-Astra | 0.145.0 | High | 32.768 |
| System | Output (K) | P90 (K) | Over 50K (%) | Time (s) |
| Fixed 50K allowance | ||||
| Codex / MiniMax-M2 | 8.7 | 17.2 | 0.0 | 212.9 |
| Codex / MiniMax-M2.1-HS | 9.0 | 20.5 | 0.0 | 177.8 |
| Codex / MiniMax-M2.5-HS | 8.8 | 23.8 | 0.0 | 179.2 |
| Codex / MiniMax-M2.7-HS | 9.4 | 20.4 | 0.0 | 182.7 |
| Codex / MiniMax-M3 | 26.0 | 37.5 | 0.0 | 562.0 |
| Latency and output | Acquisition progress | ||||||
| System | Median [IQR] (min) | P90 (min) | Output (K tokens) | First (%) | Partial (%) | Last by half (%) | Queries after |
| A. Models with Codex | |||||||
| MiniMax-M2 | 3.6 [3.1–4.0] | 4.9 | 8.7 | 19.4 | 72.3 | 88.2 | 9 |
| MiniMax-M2.1-HS | 3.0 [2.6–3.5] | 4.7 | 9.1 | 17.2 | 72.3 | 85.3 | 12.5 |
| MiniMax-M2.5-HS | 3.0 [2.7–4.0] | 4.5 | 8.8 | 16.7 | 74.5 | 82.9 | 14 |
| MiniMax-M2.7-HS | 3.1 [2.7–3.8] | 4.8 | 8.9 | 17.6 | 72.3 | 70.6 | 11 |
| Eleven models | Four frameworks | |||||
| Condition | Cases | Groups | Cases | Groups | ||
| Full | 8 | 8 | 63/61 | 9 | 8 | 69/67 |
| Endpoint-only | 7 | 6 | 51/44 | 8 | 7 | 64/55 |
| Network-only | 5 | 5 | 34/18 | 5 | 5 | 32/17 |
| Tiered retention | 8 | 7 | 61/37 | 8 | 8 | 52/39 |
| Group sampling | 5 | 5 | 34/31 | 9 | 8 | 70/68 |
| System | Reach (%) | (%) | (%) | (%) | Citations/ covered |
| A. Codex models: , , | |||||
| MiniMax-M2 | 21.3 | 18.3 | 23.5 | 4.3 | 15.6 |
| MiniMax-M2.1-HS | 21.3 | 19.4 | 20.4 | 4.0 | 27.7 |
| MiniMax-M2.5-HS | 21.3 | 22.3 | 30.6 | 6.8 | 17.4 |
| MiniMax-M2.7-HS | 25.5 | 29.1 | 33.3 | 9.7 | 11.7 |
| MiniMax-M3 | 66.0 | 45.7 | 66.9 | 30.6 | 8.0 |
| Reference | Record set | Comparison question |
| Fixed lead-query reference | Independently execute the supplied pivots and pagination without adaptive follow-up. | How much support does autonomous investigation acquire beyond those queries? |
| Shared first-page returns | Intersect first-page IDs actually returned to every model in a case. | Is support returned to all models preserved in their reports? |
| Run-specific lead returns | Exact lead queries and linked continuation pages actually consumed in a run. Their supported actions form . | Which supported actions are added to or omitted from formal-stage citations? |
| System | ||||
| Models with Codex. | ||||
| MiniMax-M2 | 23.3 | 9.3 | 0.0 | 0.0 |
| MiniMax-M2.1-HS | 23.3 | 9.3 | 0.0 | 0.0 |
| MiniMax-M2.5-HS | 23.3 | 14.0 | 2.3 | 0.0 |
| MiniMax-M2.7-HS | 27.9 | 18.6 | 2.3 | 0.0 |
| MiniMax-M3 | 72.1 | 51.2 | 25.6 | 0.0 |
| System | |||||||||
| Execution | Collection | C2 | Execution | Collection | C2 | Execution | Collection | C2 | |
| Codex | 60.2 | 47.1 | 43.5 | 69.8 | 62.5 | 60.0 | 42.0 | 29.4 | 26.1 |
| Claude Code | 43.2 | 52.9 | 39.1 | 68.4 | 100.0 | 33.3 | 29.5 | 52.9 | 13.0 |
| OpenCode | 48.9 | 58.8 | 30.4 | 74.4 | 90.0 | 57.1 | 36.4 | 52.9 | 17.4 |
| Ref-ReAct | 34.1 | 47.1 | 39.1 | 90.0 | 87.5 | 66.7 | 30.7 | 41.2 | 26.1 |
| System | Endpoints | Submitted | Supported | Time order | Source deletion(%) |
| Codex models: 208 registered relation opportunities | |||||
| MiniMax-M2 | 0 | 0 | 0 | 0/0 | 0.0–0.0 |
| MiniMax-M2.1-HS | 0 | 0 | 0 | 0/0 | 0.0–0.0 |
| MiniMax-M2.5-HS | 1 | 0 | 0 | 1/1 | 0.0–0.0 |
| MiniMax-M2.7-HS | 2 | 0 | 0 | 1/2 | 0.0–0.0 |
| MiniMax-M3 | 22 | 17 | 8 | 15/22 | 2.3–4.4 |
| Bootstrap | Leave-one-group-out | |||
| Comparison ( ) | (pp) | 95% interval (pp) | Range (pp) | (%) |
| Models with Codex | ||||
| MiniMax-M3 MiniMax-M2 | +26.3 | |||
| MiniMax-M3 MiniMax-M2.1-HS | +26.6 | |||
| MiniMax-M3 MiniMax-M2.5-HS | +23.7 | |||
| MiniMax-M3 MiniMax-M2.7-HS | +20.9 | |||
| Model | Full | Endpoint | Network | Retention | Sampling | Targeted | Matched | |||||||
| MiniMax-M2 | 21.3 | 4.9 | 15.9 | 2.3 | 16.7 | 5.6 | 16.2 | 2.7 | 16.1 | 3.2 | 20.8 | 8.3 | 17.9 | 2.6 |
| MiniMax-M2.1-HS | 21.3 | 4.9 | 25.0 | 2.3 | 16.7 | 0.0 | 21.6 | 8.1 | 16.1 | 3.2 | 14.6 | 4.2 | 17.9 | 2.6 |
| MiniMax-M2.5-HS | 16.4 | 6.6 | 27.3 | 6.8 | 22.2 | 5.6 | 40.5 | 18.9 | 19.4 | 3.2 | 14.6 | 6.3 | 20.5 | 0.0 |
| MiniMax-M2.7-HS | 42.6 | 19.7 | 25.0 | 9.1 | 27.8 | 0.0 | 27.0 | 13.5 | 22.6 | 6.5 | 20.8 | 4.2 | 30.8 | 5.1 |
| MiniMax-M3 | 59.0 | 45.9 | 34.1 | 22.7 | 50.0 | 33.3 | 43.2 | 21.6 | 38.7 | 25.8 | 43.8 | 25.0 | 46.2 | 33.3 |
| View | Codex | Claude Code | OpenCode | Ref-ReAct | ||||
| Full | 55.2 | 40.3 | 34.3 | 26.9 | 46.3 | 32.8 | 58.2 | 52.2 |
| Endpoint | 25.5 | 16.4 | 23.6 | 14.5 | 25.5 | 14.5 | 18.2 | 14.5 |
| Network | 47.1 | 35.3 | 47.1 | 41.2 | 70.6 | 47.1 | 47.1 | 47.1 |
| Retention | 38.5 | 20.5 | 38.5 | 20.5 | 25.6 | 20.5 | 28.2 | 20.5 |
| Sampling | 44.1 | 32.4 | 33.8 | 20.6 | 45.6 | 35.3 | 33.8 | 32.4 |
| Condition | Sc. | Gr. | Act. | Own | Accepted | Failed | Excluded | |
| Eleven models | ||||||||
| Endpoint | 6 | 6 | 41 | 451 | 104 | 107 | 3 | 4 |
| Network | 4 | 4 | 14 | 154 | 101 | 104 | 6 | 3 |
| Retention | 7 | 7 | 33 | 363 | 105 | 106 | 4 | 2 |
| Sampling | 3 | 3 | 19 | 209 | 100 | 106 | 4 | 6 |
| Suppression | 6 | 6 | 38 | 418 | 104 | 106 | 4 | 2 |
| Condition | UU | UA | UC | AU | AA | AC | CU | CA | CC | |
| Eleven models | ||||||||||
| Endpoint | 451 | 193 | 26 | 19 | 16 | 55 | 18 | 32 | 12 | 80 |
| Network | 154 | 61 | 7 | 7 | 3 | 14 | 9 | 1 | 6 | 46 |
| Retention | 363 | 189 | 19 | 9 | 18 | 36 | 10 | 12 | 15 | 55 |
| Sampling | 209 | 120 | 6 | 0 | 6 | 28 | 6 | 2 | 3 | 38 |
| Suppression | 418 | 186 | 7 | 7 | 16 | 47 | 10 | 15 | 15 | 115 |
| System | View | Both | Lost | Gained | Neither | Lost U | Lost A | ||
|---|---|---|---|---|---|---|---|---|---|
| Eleven models, Codex fixed | |||||||||
| MiniMax-M2 | Endpoint | 41 | 8/7 | 0 | 0 | 1 | 40 | 0 | 0 |
| MiniMax-M2 | Network | 14 | 4/3 | 0 | 1 | 1 | 12 | 0 | 1 |
| MiniMax-M2 | Retention | 33 | 5/6 | 0 | 0 | 1 | 32 | 0 | 0 |
| MiniMax-M2 | Sampling | 19 | 5/5 | 0 | 0 | 1 | 18 | 0 | 0 |
| MiniMax-M2 | Suppression | 38 | 8/8 | 2 | 1 | 2 | 33 | 0 | 1 |
| Condition | Primary | Source | Scenario | Own | Accepted | |||||
| M | F | M | F | M | F | M | F | M | F | |
| Endpoint | ||||||||||
| Network | ||||||||||
| Retention | ||||||||||
| Sampling | ||||||||||
| Suppression | ||||||||||
| Screening outcome | Actions |
| All candidates and minimal witnesses retained | 45 |
| Support survives, evidence sets change | 19 |
| Full supported, no Endpoint witness | 9 |
| No Full-view witness | 2 |
| Total | 75 |
| System | Supp. | Matched | Difference [95% interval] |
| Eleven models | |||
| MiniMax-M2 | 6.2 | 3.1 | +3.1 [-3.1, +9.6] |
| MiniMax-M2.1-HS | 3.1 | 6.2 | -3.1 [-7.3, -0.0] |
| MiniMax-M2.5-HS | 5.4 | 3.6 | +1.8 [-0.0, +7.0] |
| MiniMax-M2.7-HS | 3.4 | 10.3 | -6.9 [-14.6, -0.0] |
| MiniMax-M3 | 20.3 | 32.8 | -12.5 [-26.9, +3.8] |
| Full | Matched | |||
| Panel | ||||
| Eleven models | 39.4 | 24.2 | 36.4 | 21.2 |
| Four frameworks | 29.2 | 12.5 | 29.2 | 16.7 |
| Model | Lead | Final | Added / | Omitted / | Net (pp) |
| MiniMax-M2 | 14.4 | 4.3 | 0.7 | 75.0 | -10.1 |
| MiniMax-M2.1-HS | 13.7 | 4.0 | 1.1 | 78.9 | -9.7 |
| MiniMax-M2.5-HS | 13.7 | 6.8 | 4.3 | 81.6 | -6.8 |
| MiniMax-M2.7-HS | 14.4 | 9.7 | 7.2 | 82.5 | -4.7 |
| MiniMax-M3 | 14.4 | 30.6 | 20.5 | 30.0 | +16.2 |
| DeepSeek-V4-Pro | 14.4 | 30.2 | 21.6 | 40.0 | +15.8 |
| Investigator | Common acquired (%) | Own opportunities | Own (%) |
| Models with Codex | |||
| MiniMax-M2 | 31.6 | 9 | 0.0 |
| MiniMax-M2.1-HS | 23.7 | 10 | 10.0 |
| MiniMax-M2.5-HS | 23.7 | 12 | 8.3 |
| MiniMax-M2.7-HS | 23.7 | 20 | 15.0 |
| MiniMax-M3 | 68.4 | 38 | 55.3 |
| Left / right | Both | Left only | Right only | Neither | Share † |
| MiniMax-M3 / MiniMax-M2.7-HS | 30.6 | 43.1 | 4.2 | 22.2 | 50.0 |
| MiniMax-M3 / DeepSeek-V4-Pro | 52.9 | 15.7 | 14.7 | 16.7 | 50.8 |
| GPT-5.6-Sol / GPT-6-Astra | 56.5 | 7.1 | 12.5 | 23.9 | 83.7 |
| Codex / Ref-ReAct | 73.6 | 2.8 | 16.0 | 7.5 | 29.4 |
| Framework | Actual | Random own count | Random common count | Thin common count | Expanded fields |
| Codex | 29.6 | 1.98 [0.83, 3.31] | 1.16 [0.28, 2.21] | 20.66 [19.06, 22.10] | 31.5 |
| Claude Code | 27.1 | 3.65 [1.93, 5.25] | 2.01 [0.83, 3.31] | 18.80 [17.40, 20.17] | 28.2 |
| OpenCode | 29.6 | 4.05 [2.21, 6.08] | 2.03 [0.83, 3.31] | 18.15 [16.57, 19.89] | 30.7 |
| Ref-ReAct | 32.3 | 3.34 [1.66, 5.25] | 1.00 [0.00, 2.21] | 14.20 [12.15, 16.30] | 33.1 |
| Structure ( ) | Opportunities | Any in (%) | (%) | (%) | ||||
| M | F | M | F | M | F | M | F | |
| Single record (1) | 1,265 | 528 | 60.4 | 62.3 | 60.1 | 62.1 | 42.8 | 54.7 |
| Fixed conjunction (2) | 1,573 | 808 | 37.9 | 34.2 | 28.1 | 22.3 | 11.6 | 14.1 |
| Fixed conjunction (3) | 33 | 20 | 48.5 | 45.0 | 48.5 | 45.0 | 15.2 | 35.0 |
| Periodic (3) | 132 | 72 | 66.7 | 52.8 | 62.1 | 45.8 | 24.2 | 23.6 |
| Fixed conjunction (4) | 55 | 20 | 100.0 | 95.0 | 100.0 | 95.0 | 3.6 | 10.0 |
| Panel / set | No candidate | Count gap | Constraint gap | Sufficient |
| Eleven models: | 33.3 | 4.5 | 0.0 | 62.1 |
| Eleven models: | 49.2 | 20.5 | 6.1 | 24.2 |
| Four frameworks: | 47.2 | 6.9 | 0.0 | 45.8 |
| Four frameworks: | 54.2 | 22.2 | 0.0 | 23.6 |
| Models | Frameworks | |||
| Reference process action | Views | Any acquired / sufficiently cited (%) | Views | Any acquired / sufficiently cited (%) |
| MuddyWater: first script process | 5 | 100.0 / 100.0 | 5 | 100.0 / 100.0 |
| MuddyWater: later DarkBit-associated process | 5 | 0.0 / 0.0 | 5 | 0.0 / 0.0 |
| Spyder: later Remcos-associated process | 2 | 50.0 / 0.0 | 3 | 0.0 / 0.0 |
| PurpleFox: rootkit-associated process | 3 | 100.0 / 0.0 | 3 | 66.7 / 0.0 |
| PurpleFox: elevation-associated process | 3 | 100.0 / 66.7 | 3 | 33.3 / 0.0 |
| System | Original | Original background citations | retained (%) | Background removed (%) |
| Codex models | ||||
| MiniMax-M2 | 12 | 143 | 83.3 | 38.5 |
| MiniMax-M2.1-HS | 11 | 243 | 100.0 | 25.1 |
| MiniMax-M2.5-HS | 19 | 248 | 68.4 | 22.2 |
| MiniMax-M2.7-HS | 27 | 211 | 96.3 | 35.1 |
| MiniMax-M3 | 85 | 203 | 81.2 | 40.9 |
| Inspected example | Established by the records | Stronger interpretation requiring further support |
| Spyder: local command | Local shell invokes curl against a metadata address | Server-side request forgery or membership in the target incident |
| APT27: network activity | Two process-associated connections, followed by process termination | A ten-minute cadence or communication through 14:10:02 |
| Donot: Security 4698 | Named task creation and its stored recurrence/action configuration | Subsequent execution, reboot survival, or the initiating script without a separate binding |
| A. Treatment of the selected target (28 reports per arm) | ||
| Target-report state | Suppression | Matched control |
| Formal narrower-activity claim; registered witness complete | 0 (0.0%) | 4 (14.3%) |
| Formal narrower-activity claim; support outside the registered rule | 4 (14.3%) | 2 (7.1%) |
| Formal target claim; registered witness incomplete | 3 (10.7%) | 4 (14.3%) |
| Target discussed only as hypothesis or uncertainty | 0 (0.0%) | 0 (0.0%) |
| Target not discussed | 21 (75.0%) | 18 (64.3%) |