Unmasking Shortcut Learning in IoT Intrusion Detection: A Forensic, Multi-Paradigm Evaluation of Feature Dependence and Data Leakage
Organizations: Computer Science and Engineering Khulna University of Engineering & Technology Khulna, Bangladesh
Abstract
Machine learning-based Network Intrusion Detection Systems often report near-perfect performance on IoT benchmarks. However, whether these models learn generalizable attack behavior or exploit spurious dataset shortcuts- such as static testbed IP/MAC addresses and chronological recording artifacts-remains an important question. We evaluate the CyberFlowIoT-GICAP benchmark, containing 3,617,388 flow records across 126 PCAP sessions with 849,395 benign flows. Four learning paradigms are evaluated across four feature configurations using PCAP-disjoint splits; LightGBM is additionally evaluated using conventional random-flow splitting. When only statistical flow behavior is used (Fbehav), LightGBM (92.58% +/- 8.18%), Random Forest (92.59% +/- 8.18%), and Deep MLP (92.55% +/- 8.18%) achieve nearly identical Macro-F1, indicating that performance is constrained by feature representation rather than model complexity. With raw timestamps (Ftstamp), tree-based models reach 99.28% Macro-F1, while the linear model remains at 90.62%, showing that nonlinear models can exploit dataset-specific temporal structure. Attack detectability is highly asymmetric: high-rate and active attacks maintain >99.8% recall from flow behavior alone in nonlinear models, whereas the DNS Beaconing drops from 27.78% to 0.00% recall when contextual features are removed. Conventional random-flow splitting increases attack recall by up to 14.00%, highlighting the effect of placing flows from the same sessions in both training and test sets. We conclude with a 4-point protocol checklist for realistic IoT NIDS evaluation.
Figures & tables
| Category | PCAPs ∗ | Total Flows | Attacks | Atk % |
| BENIGN (Dedicated) | — | 849,395 | 0 | 0.00% |
| Denial of Service (DoS) | 43 (35/8) | 960,657 | 593,576 | 61.79% |
| ARP Spoofing (MITM) | 8 (8/0) | 70,915 | 61,042 | 86.08% |
| Nmap Reconnaissance | 30 (30/0) | 130,229 | 51,155 | 39.28% |
| SQL Injection (SQLi) | 6 (6/0) | 105,107 | 33,814 | 32.17% |
| DNS Beaconing | 22 (13/9) | 1,072,347 | 1,611 | 0.15% |
| Feature Set | Count | Included Feature Groups |
|---|---|---|
| 75 | 55 Behavioral, 12 App/Proto, 2 Ports, 6 Host IDs | |
| 69 | 55 Behavioral, 12 App/Proto, 2 Ports | |
| 55 | 55 Behavioral Features (Volumes, PIATs, TCP Flags, etc.) | |
| 61 | 55 Behavioral, 6 Absolute Epoch Timestamps |
| Model | Paradigm | Core Hyperparameters |
|---|---|---|
| LightGBM | GBDT | , , , |
| Random Forest | Bagging | , , |
| Deep MLP | Deep NN | , BatchNorm, Dropout 0.2, AdamW |
| Linear SGD | Linear | log_loss , L2, , |
| Model | ||||
|---|---|---|---|---|
| LightGBM | ||||
| RF | ||||
| Deep MLP | ||||
| Linear SGD |
| Attack | LightGBM | RF | MLP | SGD |
|---|---|---|---|---|
| DoS | ||||
| Nmap Recon | ||||
| SQLi | ||||
| FuerzaBruta | ||||
| MQTT | ||||
| ARP Spoofing |
| Model | DoS ( ) | ARP Spoof ( ) | DNS Beacon ( ) |
|---|---|---|---|
| LightGBM | |||
| RF | |||
| Deep MLP | |||
| Linear SGD |
| Macro-F1 (%) | Attack Recall (%) | |||||
|---|---|---|---|---|---|---|
| Feature Set | PCAP | Rand | PCAP | Rand | ||
| 88.36 | 87.99 | 80.60 | 93.49 | |||
| 88.07 | 88.00 | 79.65 | 93.65 | |||
| 92.58 | 88.65 | 89.25 | 94.45 | |||
| 99.28 | 99.57 | 99.86 | 99.42 | |||