ModSec-Learn: Boosting ModSecurity with Machine Learning
Authors: Christian Scano, Giuseppe Floris, Biagio Montaruli, Luca Demetrio, Andrea Valenza, Luca Compagna, Davide Ariu, Luca Piras, +2 more
Organizations: University of Cagliari, Cagliari, Italy · Pluribus One, Cagliari, Italy · SAP Security Research, Mougins, France · EURECOM, Biot, France · University of Genova, Genova, Italy · Prima Assicurazioni, Milano, Italy
ModSecurity is widely recognized as the standard open-source Web Application Firewall (WAF), maintained by the OWASP Foundation. It detects malicious requests by matching them against the Core Rule Set (CRS), identifying well-known attack patterns. Each rule is manually assigned a weight based on the severity of the corresponding attack, and a request is blocked if the sum of the weights of matched rules exceeds a given threshold. However, we argue that this strategy is largely ineffective against web attacks, as detection is only based on heuristics and not customized on the application to protect. In this work, we overcome this issue by proposing a machine-learning model that uses the CRS rules as input features. Through training, ModSec-Learn is able to tune the contribution of each CRS rule to predictions, thus adapting the severity level to the web applications to protect. Our experiments show that ModSec-Learn achieves a significantly better trade-off between detection and false positive rates. Finally, we analyze how sparse regularization can reduce the number of rules that are relevant at inference time, by discarding more than 30% of the CRS rules. We release our open-source code and the dataset at https://github.com/pralab/modsec-learn and https://github.com/pralab/http-traffic-dataset, respectively.
Figures & tables
Figure 1 : ModSec-Learn architecture. A machine-learning model is trained using the CRS rules as input features (52 features) to improve the trade-off between detection rate and false alarms. This amounts to learning a model of the incoming traffic directed towards the protected web services. Sparse regularization can also be used to select a subset of the available rules, instead of using PLs.
Figure 2 : ROC curves of ModSecurity vanilla (ModSec) and ModSec-Learn (SVM, RF, and LR), evaluated on test . Each curve reports the average detection rate of SQLi attacks (i.e., the True Positive Rate) against the fraction of misclassified benign SQL queries (i.e., the False Positive Rate). The zoomed section helps to understand the performance of each model when lines overlap.
PL1
PL2
PL3
PL4
ModSec vanilla
92.50 %
75.45%
68.55%
68.55%
ModSec-Learn SVM ( ℓ1 )
92.50%
99.22 %
99.04%
99.02%
ModSec-Learn SVM ( ℓ2 )
92.50%
99.22 %
99.04%
99.02%
ModSec-Learn LR ( ℓ1 )
92.50%
99.34%
99.35%
99.35 %
ModSec-Learn LR ( ℓ2 )
92.50%
99.34%
99.34%
99.34 %
ModSec-Learn RF
92.50%
99.41%
99.45%
99.45 %
Table 1 : TPR at 1% FPR of ModSec and ModSec-Learn (SVM, RF, and LR) evaluated on the test sets. For each WAF, we higlight the best results in bold.
Figure 3 : Weight values learned at PL 4 by ModSec-Learn LR - ℓ1 (blue) and ModSec-Learn LR - ℓ2 (light red), and the weight used by ModSecurity vanilla (green). The additional color, i.e., red, is given by the overlapping of the green and blue bars with the light red ones. We only report the last three digits of the rule IDs on the x-axis as the first three digits are equal to 942 for all rules.
Higher Institute for Applied Sciences and Technology (HIAST), Damascus, Syria. · Syrian Private University, Damascus, Syria. · Arab International University, Damascus, Syria. +1
University of Texas at Dallas, Richardson, TX · DEVCOM, Army Research Laboratory, Adelphi, MD · Department of Computer Science, Virginia Tech, Blacksburg, VA +1