AutoDP-LLM: Automating Data Pre-processing for Intrusion Detection Systems using Large Language Models
Organizations: Business AI Lab, College of Technology, National Economics University, Hanoi, Vietnam
Abstract
The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems (IDS). However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort and repeated evaluation of alternative configurations. For large, high-dimensional network traffic data, this process can create a significant computational burden. In this work, we propose AutoDP-LLM, an automated pre-processing framework designed to reduce manual pipeline development and computational overhead. Specifically, AutoDP-LLM leverages Large Language Models (LLMs) to autonomously generate and validate executable data pre-processing pipelines. The framework combines deterministic host-side planning with LLM-based specialist agents to formulate data-processing strategies, synthesize executable code, and adaptively determine retained feature sets using semantic reasoning and training-derived statistical evidence, without requiring a predefined feature budget. Focusing on multiclass intrusion detection, we evaluate AutoDP-LLM on the UNSW-NB15 and NSL-KDD benchmark datasets using multiple downstream classifiers. Comparative experiments against conventional feature-selection methods show that AutoDP-LLM achieves competitive detection performance while automating the generation of compact and executable pre-processing pipelines. Component-level ablation experiments further demonstrate the complementary contributions of the semantic and statistical feature-reduction components. The repeated generation, validation, execution, and assessment of candidate pipelines are amenable to parallel execution, highlighting the potential of scalable computing environments, including high-performance computing (HPC) systems, to support automated IDS pipeline development.
Figures & tables
| Classifier | Accuracy | Precision | Recall | -score | Execution Time |
| Random Forest | |||||
| KNN | |||||
| MLP | 77.51 0.64 | 77.51 0.64 | |||
| XGBoost | 82.84 0.83 | 78.41 0.22 | 106.38 9.77 | ||
| LightGBM | |||||
| CatBoost |
| Classifier | Accuracy | Precision | Recall | -score | Execution Time |
| Random Forest | |||||
| KNN | 90.97 14.65 | ||||
| MLP | 76.38 0.99 | ||||
| XGBoost | |||||
| LightGBM | |||||
| CatBoost | 79.24 0.26 | 83.82 0.19 | 79.24 0.26 |
| Framework / Method | UNSW-NB15 (XGBoost) | NSL-KDD (MLP) | ||
| Weighted (%) | Time (s) | Weighted (%) | Time (s) | |
| LiteShield-MI-FS † [ 19 ] | 78.17 | 114.07 | 67.34 | 96.47 |
| XGBoost-FS † [ 20 ] | 77.92 | 10.74 | 71.53 | 78.28 |
| IGRF-RFE † [ 21 ] | 78.13 | 857.20 | – | – |
| DSSTE † [ 22 ] | – | – | 73.59 | 77.56 |
| AutoDP-LLM (Ours) | 78.41 0.22 | 106.38 9.77 | 76.38 0.99 | |
| Ablation Variant | Component Configuration | Results (Weighted ) | |||
| Semantic Guard | Semantic Pruning | RF-MI Completion | UNSW-NB15 | NSL-KDD | |
| w/o RF-MI Completion | 77.28 0.17 | 75.03 0.47 | |||
| w/o Semantic Pruning | 77.90 0.10 | 76.26 1.14 | |||
| w/o Semantic Components | 78.33 0.00 | 74.53 0.00 | |||
| Full method | 78.41 0.22 | 76.38 0.99 | |||
| XGBoost and MLP are used as the fixed representative classifiers for UNSW-NB15 and NSL-KDD, respectively. | |||||
| LLM Engine | Classifier | Classification Performance | |||
| Accuracy | Precision | Recall | -score | ||
| Claude-Sonnet-4.6 | Random Forest | 74.96 0.31 | 83.38 0.38 | 74.96 0.31 | 77.74 0.15 |
| XGBoost | 74.84 0.25 | 83.20 0.48 | 74.84 0.25 | 77.44 0.18 | |
| MLP | 76.19 0.95 | 81.47 0.48 | 76.19 0.95 | 77.13 0.48 | |
| DeepSeek-V4-Pro | Random Forest | 76.84 0.17 | 82.31 0.08 | 76.84 0.17 | 78.56 0.08 |
| XGBoost | 77.23 0.18 | 81.83 0.13 | 77.23 0.18 | 78.34 0.11 | |
| LLM Engine | Classifier | Classification Performance | |||
| Accuracy | Precision | Recall | -score | ||
| Claude-Sonnet-4.6 | Random Forest | 76.42 0.05 | 82.33 0.03 | 76.42 0.05 | 72.81 0.10 |
| XGBoost | 78.67 0.42 | 83.30 0.22 | 78.67 0.42 | 76.05 0.41 | |
| MLP | 76.87 0.26 | 79.73 1.01 | 76.87 0.26 | 74.82 0.40 | |
| DeepSeek-V4-Pro | Random Forest | 76.35 0.63 | 82.13 0.44 | 76.35 0.63 | 72.41 0.55 |
| XGBoost | 77.31 0.85 | 82.43 0.33 | 77.31 0.85 | 73.94 1.00 | |
| Design-time Overhead | Run-time Efficiency | ||||
| LLM Engine | Output Tokens | API Latency (s) | Cost (USD) | Number of Features | Training Time (s) |
| Dataset: UNSW-NB15 (Representative classifier: XGBoost) | |||||
| Claude-Sonnet-4.6 | 18,199 11,226 | 283.26 99.65 | \pm$ 0.168 | 14.73 0.46 | |
| GPT-5.4 | 10,418 2,139 | 60.23 8.44 | \pm$ 0.032 | 13.55 0.45 | |
| DeepSeek-V4-Pro | 79,152 33,595 | 858.33 354.50 | \pm$ 0.059 | 14.19 0.42 | |
| Dataset: NSL-KDD (Representative classifier: MLP) | |||||
| Step / Agent | Parameter | Value | Design Rationale & Objective |
| Step 1: Analysis & Host Planning | Temperature | 0.0 | Promotes consistent planning and strict JSON schema compliance. |
| Step 2: Transformation & Resampling Specialist | Temperature | 0.0 | Eliminates syntax hallucinations and guarantees deterministic code execution. |
| Step 3: Feature-Reduction Specialist | Temperature | 0.0 | Ensures consistent feature selection and deterministic reduction logic. |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.