ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing
Authors: Fan Li, Xiangyu Gao, Zixuan Liu, Tong Li, Chuanpu Fu, Ziqiang Wang, Ke Xu
Organizations: Tsinghua University, Beijing, China · Zhongguancun Laboratory, Beijing, China · Renmin University of China, Beijing, China · Nanyang Technological University, Singapore
The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic offers an observable source of evidence, but how much it reveals about agent tasks and operations remains unclear. Existing traffic datasets lack the joint task and stage annotations needed to evaluate this question. We introduce ANT (Agent Network Traffic), a dataset providing agent behavior information at risk, scenario, and behavior primitive granularities alongside network traffic. ANT contains 3,114 execution episodes across 20 tasks and five scenarios, comprising 276,417 bidirectional flows and 40,049 behavior primitive segments organized into 47 macro groups. We establish a benchmark for agent risk identification, scenario recognition, and behavior primitive classification using 13 representative traffic analysis baselines. The results show that existing methods recover useful but uneven behavioral signals. They struggle to identify risk when malicious workflows resemble benign tasks and to distinguish scenarios with similar traffic patterns. Primitive classification is more reliable for frequent macro groups and those with distinctive traffic patterns than for rare or semantically similar groups. ANT provides a common basis for developing more precise auditing and forensic analysis of agent behavior from network traffic. Our data and code are available at https://anonymous.4open.science/r/ant-main-suite-7BC0/.
Figures & tables
Figure 1: Execution events and network flow intervals of a question answering task
Figure 2
Figure 3: Overview of the agent network traffic data collection system.
Figure 4
Baseline
FlowLens
YaTC
NetMamba
TFusion
TrafficFormer
DF
Accuracy ↑
0.923
0.928
0.928
0.933
0.933
0.976
F1 ↑
0.333
0.400
0.400
0.462
0.462
0.872
FPR ↓
0.000
0.000
0.000
0.000
0.000
0.011
Recall ↑
0.200
0.250
0.250
0.300
0.300
0.850
Table 3: Risk identification under full traffic.
Figure 6Figure 7
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Group
Canonical label
Primitive types
Train
Validation
Test
N001
analysis
90
1308
152
67
N002
api_discovery
73
531
34
19
N003
api_inspection
91
1408
130
103
N004
api_probing
91
671
21
37
N005
api_query
95
1685
73
55
N006
application_reconnaissance
86
451
26
8
Appendix
Table 7: Primitive macro groups and split counts, including empty segments and segments without usable TCP packets, excluding synthetic prelude segments. Group IDs refer to the released primitive-to-group mapping.
Figure 11: Primitive clustering diagnostics. Left: cosine silhouette and seed stability across candidate K , with crosses marking partitions containing groups of fewer than 10 primitive types. Right: macro group size, mean cosine silhouette, and the fraction of primitive types with negative silhouette at K=47 . Group IDs match Table 7 .
Representation
Dwithin
Dbetween
Gd [95% CI] (%)
pH
Primitive embedding MMD
0.431680
0.510646
15.46 [14.32, 16.64]
0.0006
Primitive macro group Hellinger
0.698573
0.819826
14.79 [13.98, 15.59]
0.0050
Primitive embedding area
0.147906
0.158600
6.74 [5.85, 7.68]
0.0182
Primitive macro group residual
0.389167
0.412703
5.70 [5.09, 6.36]
0.0182
Appendix
Table 8: Aggregate scenario cohesion across distinct tasks. Intervals summarize execution variability within the observed task inventory. The adjusted pH values account for the six primary tests. Distances use the scale of each representation, and gains report relative cohesion.
Component
Information in bits
Share of total in percent
Scenario
0.7703
34.86
Task within scenario
0.5045
22.84
Episode within task
0.9346
42.30
Total
2.2094
100.00
Appendix
Table 9: Aggregate decomposition of empirical primitive macro group composition. Components are averaged across settings before dividing by the average total.
Representation
Reference order
Observed
Reference mean
ΔG
pH
Directed area
Unrestricted
6.74
0.97
5.77
0.0050
Transition residual
Unrestricted
5.70
0.46
5.24
0.0050
Directed area
Role-preserving
6.74
3.26
3.48
0.0020
Transition residual
Role-preserving
5.70
1.42
4.28
0.0020
Appendix
Table 10: Aggregate additional sequence cohesion. Observed and reference-mean gains are percentages, and ΔG is in percentage points. Unrestricted-order tests use the six-test Holm adjustment. Role-preserving checks use a separate two-test adjustment.
Oct 5, 2026·Christoph Bühler, Matteo Biagiola, Luca Di Grazia +1Security Evaluation
University of St. Gallen, St. Gallen, SG, Switzerland · University of St. Gallen and Università della Svizzera italiana (USI), St. Gallen and Lugano, SG and TI, Switzerland
School of Computing, Wichita State University, Kansas, USA · Department of Computer Science, American International University-Bangladesh, Dhaka, Bangladesh