cs.DBOct 5, 2026

Protocol-Sensitive Evaluation of Log Anomaly Detection: Component Costs and Target-Access Sensitivity on HDFS and BGL

Authors: Hang Xiao, Janet Sung, Zhaoyi Li, Gangzhen Qian, Chuhong Xu

Organizations: Fortinet, Inc. Sunnyvale, USA · Google LLC Mountain View, USA · Independent Researcher Mountain View, USA · Sony Corporate of America San Jose, USA

Abstract

Protocol choices can change the conclusions drawn from log anomaly detection benchmarks even when detector settings are fixed. We present a joint empirical study of split construction, representation visibility, and component costs using six fixed count, sequence, and semantic configurations on Hadoop Distributed File System (HDFS) and Blue Gene/L (BGL) logs. Random splits place several configurations near the average-precision ceiling, whereas group-disjoint HDFS and chronological BGL evaluation produce lower scores and different observed orderings. At a fixed BGL cutoff, parser choice spans 0.124 in semantic XGBoost mean average precision while preserving its lead over count XGBoost; the earliest rolling period reverses that ordering. A two-factor cross-system ablation contrasts source-only representations with offline transductive access to unlabeled target templates through the representation corpus and inverse document frequency: HDFS-to-BGL mean average precision moves from 0.191 with source-only access to 0.325 with union-corpus, target-IDF access, and the intermediate conditions reveal direction-dependent interactions in average precision and retrieval at fixed review budgets. Component-level profiling separates parsing and representation costs from classifier training, prediction, and storage. Together, these findings connect detector comparisons to the test population, preprocessing state, visible information, and measured pipeline stages, and identify the protocol fields needed alongside a score to support interpretable comparisons of log anomaly detection accuracy and resource use.

Figures & tables

Explore similar work

CardsList
  1. NLLog: Lightweight, Explainable SOC Anomaly Detection via Log-to-Language Rewriting

    Jun 3, 2026Samuel Ndichu, Tao Ban, Seiichi Ozawa +2Interpretable Anomaly DetectionThreat Detection

  2. LogNEO: A GPT-Neo Reinforcement Learning Framework for Accurate Real-Time Log Anomaly Detection

    Jun 6, 2026David Eje, Tanmay Sharma, Khush Patel +2Interpretable Anomaly DetectionOpenai

  3. FAME: Failure-Aware Mixture-of-Experts for Message-Level Log Anomaly Detection

    May 21, 2026Huanchi Wang, Zihang Huang, Yifang Tian +3Interpretable Anomaly DetectionEarly Failure Prediction