Organizations: Fortinet, Inc. Sunnyvale, USA · Google LLC Mountain View, USA · Independent Researcher Mountain View, USA · Sony Corporate of America San Jose, USA
Protocol choices can change the conclusions drawn from log anomaly detection benchmarks even when detector settings are fixed. We present a joint empirical study of split construction, representation visibility, and component costs using six fixed count, sequence, and semantic configurations on Hadoop Distributed File System (HDFS) and Blue Gene/L (BGL) logs. Random splits place several configurations near the average-precision ceiling, whereas group-disjoint HDFS and chronological BGL evaluation produce lower scores and different observed orderings. At a fixed BGL cutoff, parser choice spans 0.124 in semantic XGBoost mean average precision while preserving its lead over count XGBoost; the earliest rolling period reverses that ordering. A two-factor cross-system ablation contrasts source-only representations with offline transductive access to unlabeled target templates through the representation corpus and inverse document frequency: HDFS-to-BGL mean average precision moves from 0.191 with source-only access to 0.325 with union-corpus, target-IDF access, and the intermediate conditions reveal direction-dependent interactions in average precision and retrieval at fixed review budgets. Component-level profiling separates parsing and representation costs from classifier training, prediction, and storage. Together, these findings connect detector comparisons to the test population, preprocessing state, visible information, and measured pipeline stages, and identify the protocol fields needed alongside a score to support interpretable comparisons of log anomaly detection accuracy and resource use.
Figures & tables
Characteristic
HDFS
BGL
Unit
block session
100-line window
Sessions/windows
575,061
47,479
Anomalies
16,838 (2.93%)
4,802 (10.11%)
Event vocabulary
29
2,746
Mean length
19.4
100
Distinct count vectors
589
6,280
TABLE I: Dataset characteristics. HDFS and BGL differ in scale, vocabulary, prevalence, repetition, and session structure.
32-d embedding; 64 hidden; Adam 10−3 ; batch 256; at most five epochs
TABLE III: Fixed hyperparameters used across protocols and seeds.
Random split
Group-disjoint
Configuration
AP
F1
AP
F1
Count/LR
.991 ± .004
.989 ± .006
.867 ± .103
.934 ± .046
Count/RF
1.000 ± .000
.998 ± .000
.743 ± .227
.808 ± .259
Count/XGB
.999 ± .001
.998 ± .000
.692 ± .175
.582 ± .337
Sequence/LSTM
.999 ± .001
.993 ± .002
.627 ± .210
.502 ± .324
Semantic/LR
.982 ± .003
.981 ± .003
.725 ± .174
.792 ± .275
TABLE IV: HDFS results across five seeds (mean ± SD); random/group access is transductive.
Random split
Chronological
Configuration
AP
F1
AP
F1
Count/LR
.987 ± .003
.976 ± .004
.277 ± .000
.423 ± .000
Count/RF
.998 ± .001
.984 ± .002
.388 ± .005
.400 ± .012
Count/XGB
.998 ± .001
.989 ± .004
.448 ± .000
.396 ± .000
Sequence/LSTM
.946 ± .012
.916 ± .017
.236 ± .151
.282 ± .091
Semantic/LR
.916 ± .006
.870 ± .007
.205 ± .095
.267 ± .043
TABLE V: BGL results across five seeds (mean ± SD); random access is transductive, chronological access is inductive.
Fig. 1: Mean AP ( ± standard deviation over five seeds) of the six configurations under random splitting and each dataset’s alternative protocol, from Tables IV and V . Group-disjoint HDFS and chronological BGL yield lower mean scores and different leading configurations.
Test anomaly
Count/XGB
LSTM
Semantic/XGB
Panel
Setting
rate
AP
AP
AP
Parser
Compact s=.3
10.1%
.460 ± .000
–
.741 ± .054
Parser
Compact s=.4
10.1%
.448 ± .000
–
.676 ± .086
Parser
Compact s=.5
10.1%
.448 ± .000
–
.631 ± .109
Parser
Drain3 s=.4
10.1%
.432 ± .000
–
.617 ± .177
Rolling
50 → 60%
1.3%
.902 ± .000
.475 ± .147
.838 ± .047
TABLE VI: BGL chronological sensitivity over five seeds. Parser: 70/30 split; rolling: next 10% after each cutoff.
System
Configuration
Fit (s)
Infer. (k/s)
Size (MB)
HDFS
Count/LR
2.6 ± 1.1
6771 ± 2151
.002
HDFS
Count/RF
17.8 ± 2.1
312 ± 91
2.234
HDFS
Count/XGB
3.8 ± 1.8
660 ± 94
.423
HDFS
Sequence/LSTM
309.3 ± 13.2
37 ± 4
.116
HDFS
Semantic/LR
31.4 ± 21.3
1257 ± 113
.004
HDFS
Semantic/XGB
9.4 ± .8
696 ± 189
.416
TABLE VII: Component-level cost over five seeds. Throughput covers classifier-only inference; size covers the serialized classifier.
Word2vec
IDF
AP
P@1%
R@1%
P@5%
R@5%
HDFS → BGL
Source
Source
.191 ± .107
.242 ± .160
.024 ± .016
.220 ± .174
.109 ± .086
Source
Target
.187 ± .084
.259 ± .163
.026 ± .016
.253 ± .163
.125 ± .081
Union
Source
.297 ± .136
.443 ± .266
.044 ± .026
.293 ± .021
.145 ± .010
Union
Target
.325 ± .159
.446 ± .312
.044 ± .031
.318 ± .122
.157 ± .060
BGL → HDFS
TABLE VIII: Cross-system XGBoost target-catalog visibility over five seeds. Union corpus and target IDF are offline transductive; budgets use tie-aware allocation.
System-generated logs underpin security monitoring, yet their rigid template-based format hinders both automated analysis and human comprehension. We present NLLog (Natural-Language Log), a lightweight pipeline that deterministically rewrites parsed templates into WHO-WHAT-SEVERITY sentences, pools them with term-frequency-inverse-document-frequency weighting, classifies sessions with tree ensembles, and back-projects evidence with TreeSHAP for analyst review. On Hadoop Distributed File System (HDFS) and Blue Gene/L (BGL) corpora, NLLog exceeds two reproduced matched-protocol baselines; across HDFS, BGL, and the AIT Alert Data Set, it sustains low false-positive rates with commodity-hardware latency suitable for security operations center triage. Coverage, sparse-versus-dense, faithfulness, and adversarial ablations show that fallback sufficiency is corpus-dependent, that an enrollment-time coverage check can surface refinement requirements before deployment, and that an auditable deterministic rewrite combined with lightweight dense encoding provides a measurable representation layer for log-anomaly detection and triage.
Samuel Ndichu, Tao Ban, Seiichi Ozawa +2
National Institute of Information and Communications Technology, Tokyo, Japan · Kobe University, Kobe, Japan
Detecting anomalies in large-scale system logs is critical for the reliability and security of modern computing infrastructure. We present LogNEO, a log anomaly detector built on EleutherAI's GPT-Neo (1.3B parameters) and fine-tuned with a novel partial-credit, exponentially decaying position-aware reward scheme combined with cross-entropy regularisation via Proximal Policy Optimisation (PPO). The position-aware reward explicitly models prediction difficulty: early positions receive higher rewards for correct predictions, while later positions incur stronger penalties for errors. LogNEO attains F1-scores of 0.927, 0.913, and 0.984 on the HDFS, BGL, and Thunderbird benchmarks, improving recall by up to 6 percentage points over the prior state-of-the-art LogGPT while maintaining comparable precision. A production microservice deployment over Apache Kafka, Redis, and TensorRT-accelerated inference demonstrates 45 ms end-to-end latency at 15,000 events per second.
David Eje, Tanmay Sharma, Khush Patel +2
Department of Computer Science, Innopolis University, Innopolis, Russia
Production systems generate millions of log lines daily, yet most anomaly detectors operate at the session or window-level, flagging groups of lines rather than identifying the specific message responsible. This coarse granularity forces operators to inspect many routine lines per alert. Message-level detection offers finer granularity, but remains challenging. A single event template may correspond to both normal and anomalous messages, failures arise from heterogeneous subsystems, and line-level labeling at scale is impractical. Although large language models (LLMs) can reason over log semantics, applying them to every line is too costly for continuous monitoring. We present FAME (Failure-Aware Mixture-of-Experts), a label-efficient message-level mixture-of-experts framework that uses an LLM only once offline. We annotate at most K labeled lines per template to derive binary normal/anomaly indicators and representative examples. The LLM proposes a partition of templates into failure domains, and a certification step validates the proposal before training. FAME trains a lightweight router and domain experts that run on-premise and output anomaly predictions and failure-domain labels. On BGL, FAME achieves F1 = 98.16 at K = 100 reducing annotation effort by 76x and detects 86.3% of anomalies from unseen EventIDs. On Thunderbird, FAME reaches F1 = 99.95 with perfect recall.
Huanchi Wang, Zihang Huang, Yifang Tian +3
Department of Electrical and Computer Engineering University of Toronto, Toronto, Ontario