cs.AIOct 5, 2026

ANT: A Multi-Granularity Network Traffic Dataset and Benchmark for Agents Behavior Auditing

Authors: Fan Li, Xiangyu Gao, Zixuan Liu, Tong Li, Chuanpu Fu, Ziqiang Wang, Ke Xu

Organizations: Tsinghua University, Beijing, China · Zhongguancun Laboratory, Beijing, China · Renmin University of China, Beijing, China · Nanyang Technological University, Singapore

Abstract

The growing adoption of large language model (LLM) agents creates a need for network administrators and security teams to audit agent behavior within organizational networks without inspecting private user content. Network traffic offers an observable source of evidence, but how much it reveals about agent tasks and operations remains unclear. Existing traffic datasets lack the joint task and stage annotations needed to evaluate this question. We introduce ANT (Agent Network Traffic), a dataset providing agent behavior information at risk, scenario, and behavior primitive granularities alongside network traffic. ANT contains 3,114 execution episodes across 20 tasks and five scenarios, comprising 276,417 bidirectional flows and 40,049 behavior primitive segments organized into 47 macro groups. We establish a benchmark for agent risk identification, scenario recognition, and behavior primitive classification using 13 representative traffic analysis baselines. The results show that existing methods recover useful but uneven behavioral signals. They struggle to identify risk when malicious workflows resemble benign tasks and to distinguish scenarios with similar traffic patterns. Primitive classification is more reliable for frequent macro groups and those with distinctive traffic patterns than for rare or semantically similar groups. ANT provides a common basis for developing more precise auditing and forensic analysis of agent behavior from network traffic. Our data and code are available at https://anonymous.4open.science/r/ant-main-suite-7BC0/.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework

    Aug 7, 2026Jing Chen, Yang Sun, Li Zhang +2

  2. AgentSpy: Making AI Agent Behavior Observable

    Oct 5, 2026Christoph Bühler, Matteo Biagiola, Luca Di Grazia +1Security Evaluation

  3. NetInjectBench: Benchmarking Indirect Prompt Injection in Tool-Using Large Language Model Agents for Network Operations

    Jul 11, 2026Ruksat Khan Shayoni, Muhammad Faraz Shoaib, S M Asif Hossain +1Indirect Prompt InjectionLanguage-Model Agents