cs.CROct 4, 2026

AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models

Authors: Bao-Phong Nguyen, Gia-Khanh Pham, Thai-Duong Do, Mai Xuan Trang, Minh-Tuan Le, Xuan-Nam Tran, Huan Vu, Tien-Cuong Nguyen, +2 more

Organizations: Business AI Lab, College of Technology, National Economics University, Hanoi, Vietnam

Abstract

The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems (IDS). However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort and repeated evaluation of alternative configurations. For large, high-dimensional network traffic data, this process can create a significant computational burden. In this work, we propose AutoDP-LLM, an automated pre-processing framework designed to reduce manual pipeline development and computational overhead. Specifically, AutoDP-LLM leverages Large Language Models (LLMs) to autonomously generate and validate executable data pre-processing pipelines. The framework combines deterministic host-side planning with LLM-based specialist agents to formulate data-processing strategies, synthesize executable code, and adaptively determine retained feature sets using semantic reasoning and training-derived statistical evidence, without requiring a predefined feature budget. Focusing on multiclass intrusion detection, we evaluate AutoDP-LLM on the UNSW-NB15 and NSL-KDD benchmark datasets using multiple downstream classifiers. Comparative experiments against conventional feature-selection methods show that AutoDP-LLM achieves competitive detection performance while automating the generation of compact and executable pre-processing pipelines. Component-level ablation experiments further demonstrate the complementary contributions of the semantic and statistical feature-reduction components. The repeated generation, validation, execution, and assessment of candidate pipelines are amenable to parallel execution, highlighting the potential of scalable computing environments, including high-performance computing (HPC) systems, to support automated IDS pipeline development.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GenTI: Benchmarking LLMs for Autonomous IDPS Rule Generation for Unseen Attacks

    Jun 4, 2026Hassan Jalil Hadi, Rehana Yasmin, Ali ShokerIntrusion DetectionCyber Threat Intelligence

  2. Jev-IDS: System One Models for Network Intrusion Detection

    Oct 1, 2026Paulo Severo, Silvio E. Quincozes, Amanda DiasIntrusion DetectionRandom Forest

  3. Evaluation of AutoML Frameworks for IDS under Imbalanced Data Conditions of the NSL-KDD Dataset

    Jun 10, 2026Wiliane Carolina Silva, Evandro César Vilas Boas, Felipe A. P. de FigueiredoImbalanced ClassificationIntrusion Detection