Autonomous and agentic clients can distribute reconnaissance across many identities so that each request remains valid, low-rate, and benign-looking while the population collectively acquires broad system knowledge. We formalize this threat as Distributed Collective Reconnaissance (DCR) and present SwarmReconGuard, a reproducible black-box benchmark in which the defender observes only service-boundary telemetry. The Docker-isolated study evaluates 11 benign and attack behaviors across 10-10,000 virtual identities, comprising 440 test runs and 3,666,300 requests, with complete telemetry integrity. We compare semantic, Gaussian, conditional, graph, kernel, hybrid, and CUSUM-based detectors. Gaussian likelihood-ratio detection achieves 100% detection with 0% observed false positives on known attacks but only 3% on unseen policies. CUSUM yields 36.1% overall detection at 1.25% false positives, while hybrid CUSUM reaches 85.7% detection with 0% observed false positives at 10,000 identities. Results expose a major policy-generalization gap and motivate exposure-aware, scale-aware defenses.
Figures & tables
Figure 1: SwarmReconGuard architecture. Attacker-side policy and coordination are hidden. The defender receives only API-boundary telemetry, from which window-level semantic, graph, distributional, and sequential evidence is derived.
Scenario
Label
Purpose
benign_flash
Benign
Locality, popular resources, and repeated ordinary access.
benign_diagnostic
Benign
Legitimate diagnostic access; removes an attack-exclusive metadata cue.
benign_explorer
Benign
Independent broad exploration with low diagnostic probability.
benign_bulk
Benign
Broad repetitive batch workflow and high-volume hard negative.
independent_recon
Known attack
Independent uniform reconnaissance; one of two attack policies used for fitting.
coordinated_swarm
Known attack
Randomized stratified partitioning that reduces duplicate work.
Table 1: Benchmark Scenarios
Method
AUROC
AP
FPR
Detection
Precision
F1
Exposure alarm/end
Median delay
Heuristic
0.828
0.895
0.625%
32.1%
98.9%
0.485
6.15%
0.90 s
Gaussian LLR
0.820
0.890
0.000%
30.7%
100%
0.470
6.15%
6.45 s
Conditional LLR
0.789
0.867
0.625%
21.1%
98.3%
0.347
6.04%
0.85 s
CUSUM
0.787
0.873
1.250%
36.1%
98.1%
0.527
4.89%
1.11 s
Hybrid CUSUM
0.752
0.855
0.625%
29.3%
98.8%
0.452
5.69%
35.74 s
Hybrid without graph
0.842
0.901
0.625%
28.9%
98.8%
0.448
6.16%
0.70 s
Table 2: Selected Online Results Across 440 Test Runs
Figure 2: Observed operating-point trade-off between false positives and attack detection. Threshold-independent discrimination and calibrated recall are distinct properties.
Method
Known policies (80 attacks)
Unseen policies (200 attacks)
Detection
FPR
AUROC
Detection
FPR
AUROC
Heuristic
93.75%
0.625%
0.996
7.50%
0.625%
0.761
Gaussian LLR
100%
0%
1.000
3.00%
0%
0.748
CUSUM
75.00%
1.250%
0.996
20.50%
1.250%
0.703
Hybrid CUSUM
48.75%
0.625%
0.813
21.50%
0.625%
0.728
Hybrid without graph
86.25%
0.625%
0.998
6.00%
0.625%
0.780
Table 3: Known-Policy Versus Unseen-Policy Generalization
Figure 3: Detection rate on known and unseen attack policies. Near-perfect known-policy performance does not imply robustness to policy shift.
Figure 4: Detection rate by population scale for selected methods. The response to scale is detector-dependent; hybrid CUSUM becomes especially effective at 10,000 identities.
Variant
AUROC
Detection
FPR
Full hybrid
0.674
8.2%
0.0%
No graph
0.842
28.9%
0.625%
No heuristic
0.603
1.4%
0.0%
No LLR
0.660
7.9%
0.0%
Table 4: Pointwise Hybrid Ablation
Figure 5: AUROC for the pointwise hybrid and component-removal ablations. Removing graph evidence improves performance, identifying a concrete limitation of the current graph descriptor set.
Figure 6: Measured service p95 latency and throughput versus population size. The synthetic service saturates at high scale; security conclusions therefore compare attacks with matched benign load rather than claiming a production SLA.
Language models can find thousands of severe software vulnerabilities, and agents are increasingly being misused for cyberattacks. To avoid detection, attackers frequently distribute their misuse, splitting a harmful task across many user accounts so each individual transcript looks benign. Because safety monitors score only one agent context at a time, they are structurally blind to misuse that is only visible in aggregate, across many accounts. We show this gap is real by building, to our knowledge, the first distributed agent attack, a multi-agent scaffold that completes hard cybersecurity tasks while hiding the harmful objective across subagents with limited contexts, evading a standard monitor that catches it only a fifth as often as prior agent attacks. Towards a defense, we develop an online stateful monitor that uses real-time clustering to collect weak suspiciousness signals across many agent transcripts, and escalates only rarely to a language model that flags misuse across user accounts. In evaluations with large-scale simulated datacenter traffic, our monitor Pareto dominates standard monitors, catching distributed attacks 30% earlier and flagging cyber misuse before it reaches the most harmful stages. Crucially, this comes at negligible additional latency for ~99% of user traffic. This detection advantage persists but narrows as the benign background traffic grows very large. After an extensive red-teaming exercise, we improve the defense and surprisingly also find that it catches standard jailbreaks, since adaptive attackers reuse attack variants across accounts. Our results point toward a new class of safety monitors which reason over groups of users rather than isolated transcripts.
Davis Brown, Samarth Bhargav, Arav Santhanam +7
University of Pennsylvania · 2Carnegie Mellon University
Traditional pentesting uses reconnaissance at each step to uncover unseen weaknesses, build stronger attacks, and advance the objective; we argue that AI agents require the same treatment. We formalize agent reconnaissance by modeling the process and identifying the knowledge assets it seeks to extract: what they are, how they are used, and which agent weaknesses they exploit to give adversaries leverage in indirect prompt injection attacks. We instantiate these insights in Know Your Agent (KYA), a framework that automates black-box, reconnaissance-driven pentesting by probing agents, building target profiles, and using those profiles to craft stronger attacks. We evaluate KYA on agent-security benchmarks and a real-world coding agent, and release KYA, its benchmarks, and baseline implementations for reproducibility.
Or Zion Eliav, Eyal Lenga, Shir Bernstien +1
Faculty of Computer and Information Science Ben-Gurion University of the Negev Beer-Sheva, Israel
Agentic security systems increasingly audit live targets with tool-using LLMs, but prior systems fix a single coordination topology, leaving unclear when additional agents help and when they only add cost. We treat topology choice as an empirical systems question. We introduce a controlled benchmark of 20 interactive targets (10 web/API and 10 binary), each exposing one endpoint-reachable ground-truth vulnerability, evaluated in whitebox and blackbox modes. The core study executes 600 runs over five architecture families, three model families, and both access modes, with a separate 60-run long-context pilot reported only in the appendix. On the completed core benchmark, detection-any reaches 58.0% and validated detection reaches 49.8%. MAS-Indep attains the highest validated detection rate (64.2%), while SAS is the strongest efficiency baseline at $0.058 per validated finding. Whitebox materially outperforms blackbox (67.0% vs. 32.7% validated detection), and web materially outperforms binary (74.3% vs. 25.3%). Bootstrap confidence intervals and paired target-level deltas show that the dominant effects are observability and domain, while some leading whitebox topologies remain statistically close. The main result is a non-monotonic cost-quality frontier: broader coordination can improve coverage, but it does not dominate once latency, token cost, and exploit-validation difficulty are taken into account.