HydroJEV: A one-second, training-free screen for cyber-attack and fault attribution in water distribution networks
Authors: Tianwei Mu, Shengyan Jiang, Mingzhe Yuan, Qing Luo, Min Xiao, Wenhong Wang, Jun Li, Manhong Huang
Organizations: School of Municipal Engineering and Environment, Shenyang Jianzhu University, Shenyang 110168, China · Guangzhou Institute of Industrial Intelligence, Guangzhou 510000, China · Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110169, China · Key Laboratory of Ecological Restoration of Regional Contaminated Environment, Ministry of Education, College of Environment, Shenyang University, Shenyang 110044, China · College of Environmental Science and Engineering, State Environmental Protection Engineering Center for Pollution Treatment and Control in Textile Industry, Donghua University, Shanghai 201620, China
When a SCADA alarm is raised in a water distribution network, operators must decide quickly whether it reflects a cyberattack, a physical fault, a normal transient or a faulty sensor. Supervised classifiers need labelled incidents that utilities rarely have, and frontier large language models (LLMs) take tens of seconds per decision. We tested whether Jev, a training-free model that returns class probabilities in about one second, can serve as the first tier of this triage. On a four-class cause-attribution benchmark built on the C-Town network in EPANET, Jev was compared with a hand-written rule tree, a supervised classifier and seven cloud LLMs on identical evidence in four sealed, pre-registered rounds. With only a label-free prior correction, Jev matched the rule tree (macro-F1 0.62-0.64 against 0.56-0.61 in distribution) and exceeded the supervised classifier by 0.36-0.42 on event subtypes absent from its labels, in all four rounds, and it outperformed the classifier whenever fewer than about four labelled events per class were available. Jev also decided 20-40 times faster than frontier LLMs. Accepting only benign Jev verdicts confirmed by the rule tree spared an LLM reviewer 35-38% of windows on fresh sealed sets without loss of macro-F1. Transferred unchanged to two further networks, this gated cascade stayed within the non-inferiority margin of its reviewer on all four sets. A fast, training-free screen can therefore take over about a third of the review load in SCADA anomaly triage while preserving the accuracy of deliberate review.
Figures & tables
Class
Reference (training, in distribution)
Unseen A (round 1-2)
Unseen B (round 3)
Unseen C (rounds 4-6)
Cyberattack
Overpump, masked level; pump off, spoofed running
Pump on, spoofed off; overpump, level replay
Set-point tamper, unmasked; pump stop, 24 h replay
Level spoof, starve; masked duty swap
Physical fault
Burst
Pipe closure
Pump wear; dual leak
Growing leak; pump trip
Normal transient
Demand rise; demand fall; operator switch
Demand pattern shift
Duty pump rotation; temporary demand peak
Gradual demand ramp; erratic demand
Sensor fault
Stuck; drift; offset
Gain error
Dropout to zero; intermittent spikes
Noise burst; lagged channel
Table 1
Figure 1: Study design. (a) The C-Town benchmark network, with its tanks, reservoir and pumps. Four event classes are instantiated by four disjoint subtype families, and 1,279 windows were evaluated in four sealed rounds. (b) The two-tier triage. Every window is first judged by the training-free Jev screen (median 1.2 s per decision, no labels). A benign Jev verdict that the rule tree R1 confirms is accepted; every other window goes to an LLM reviewer (median time per decision, one request at a time). The ungated rule cascade, which accepts every benign Jev verdict and otherwise uses R1, was also tested. (c) The pre-registered rounds and their sealed sets. Blue, in-distribution (reference-family) sets; red, sets from a new event family. Rounds 5 and 6, which transferred the round-4 screen to two further networks, are described in Note S11 .
Round
Sealed sets (windows)
Tested design
Design exposure
Primary tests
1
E1, in distribution (40); E2, unseen A (40)
Jev, evidence format 2
training, development
4
2
E1r2, in distribution (200); E2r2, unseen A (200)
replication
+ round 1
3
2-LLM
E1r2 and E2r2 answered by seven LLMs
same state and question
+ round-2 Jev results
14 (two-sided)
3
E1r3, in distribution (199); E3, unseen B (200)
evidence format 3, physics rules
+ rounds 1-2
5
4
E1c, in distribution (200); E4c, unseen C (200)
screening cascades
cascades designed on round 2, confirmed on round 3
7
5
Net3: N3-V1, N3-N4 (200 each)
round-4 screen, transferred unchanged
C-Town only
4
Table 3
Figure 2: Accuracy of training-free and supervised methods on the eight C-Town sealed sets. (a) Macro-F1 of calibrated Jev, the rule tree R1, the supervised classifier R2 retrained with the evaluated subtype held out, and R2 trained on all 200 labels, for the four in-distribution sets (E1, E1r2, E1r3, E1c) and the four new-family sets (E2, E2r2, E3, E4c). Holding out a subtype is defined only for in-distribution sets. The dotted line marks the macro-F1 of a chance-level four-class classifier. (b) Paired differences in macro-F1 per sealed set; the difference to R2 with the subtype held out was a primary test in all four rounds and significant after Holm correction in each. (c) Recall by subtype over the four in-distribution sets, for Jev and for R2 retrained without that subtype; shading gives the true class. In all panels, bars show the mean over sealed sets, error bars ± 1 s.d. and white dots the individual sealed sets; the mean is printed above each bar.
Figure 3: Label scarcity. (a, b) Macro-F1 of R2 trained on k labelled events per class, against Jev, which uses no labels, on the four in-distribution sets (a) and the four new-family sets (b). For each set, the R2 value at a given k is the mean over 200 random label draws; the dashed line is Jev’s mean. (c) Share of the 200 label draws in which R2 scored below Jev, for each k; the dotted line marks one half. Bars show the mean over the four sealed sets, error bars ± 1 s.d. and white dots the individual sets; the mean is printed above each bar. The range over the 200 draws is shown per set in Note S7 .
Figure 4: Jev against seven general-purpose LLMs, with the same evidence, question and label-free calibration for every model. (a) Macro-F1 on the two round-2 sealed sets (E1r2, E2r2), with the rule tree R1 and R2 trained on 200 labels as references; † marks a set on which the LLM was better than Jev after Holm correction over 14 two-sided tests. With two sets per method, the s.d. only indicates their spread. (b) Time per decision: every round-4 Jev call (n = 400) and the 40 round-4 reviewer decisions per model (20 per set), sent strictly one request at a time. Bars show the mean, error bars ± 1 s.d. and white dots the individual sets (a) or decisions (b); the mean is printed above each bar.
Figure 5: Jev as a first-line screen. (a) Pre-registered cascade comparisons in macro-F1, on the four design and confirmation sets (E1r2, E2r2, E1r3, E3; light, n = 4) and on the two fresh sealed sets (E1c, E4c; dark, n = 2). The dotted line is the non-inferiority margin ( − 0.03) of the LLM cascade; “met” counts the fresh-set primary tests that rejected their null hypothesis after Holm correction over seven tests. (b) Accept rate and accept precision of the rule gate (every benign Jev verdict) and the agreement gate (benign Jev verdicts confirmed by R1) over the eight C-Town sealed sets. (c) Windows accepted per set by true class, for the same two gates. (d) Time per decision of each reviewer alone and behind the Jev screen, on the 40 round-4 windows timed one request at a time; decisions settled by the screen take only the Jev time. Bars show the mean, error bars ± 1 s.d. and white dots the individual sealed sets (a-c) or decisions (d); the mean is printed above each bar. With two fresh sets, the s.d. in (a) only indicates their spread.
Figure 6: Components of the Jev screen. (a) Macro-F1 of Jev with raw probabilities, with the inductive label-free prior used throughout the study, and with a transductive prior re-estimated on each evaluated set, over the six sealed sets of rounds 1-3. (b) Macro-F1 on the two round-3 sealed sets of Jev with evidence formats 2 and 3, of the rule tree R1, of R1 with a physics override (R1b), and of the same override applied to the Jev verdict. (c) Recall by true class of Jev, R1 and R2 trained on 200 labels over the eight C-Town sealed sets. Bars show the mean, error bars ± 1 s.d. and white dots the individual sealed sets; the mean is printed above each bar. With two sets in (b), the s.d. only indicates their spread.
Network
Split
Seed base
Windows
Family
Use
C-Town
Training
1,100,000
200 (50 per class)
Reference
R2 training; label-scarcity draws
C-Town
Development
1,300,000
40 (10 per class)
Reference
Question choice; unlabelled calibration prior
C-Town
Round 1
1,500,000/ 1,700,000
40 / 40
Reference/ Unseen A
Sealed E1 / E2
C-Town
Round 2
5,050,000/ 5,060,000
200 / 200
Reference/ Unseen A
Sealed E1r2 / E2r2
C-Town
Round 3
5,070,000/ 5,080,000
200 / 200
Reference/ Unseen B
Sealed E1r3 / E3
C-Town
Round 4
5,090,000/ 5,095,000
200 / 200
Reference/ Unseen C
Sealed E1c / E4c
Table 9
Rule
Condition
Decision
1
System demand violated and pressure pattern widespread
Normal transient
2
Pump head gain and tank level both violated
Cyberattack
3
Tank mass balance violated and control logic violated
Cyberattack
3 ′
Tank mass balance violated and control logic not violated
Sensor fault
4
Pump status versus flow violated, or exactly one anomalous channel without a system-demand violation
Sensor fault
5
System demand violated, or pressure pattern localized
Physical fault
Table 10
Round
Set
Test
Δ macro-F1 [95% CI]
p
Holm
Met
1
E2
Jev > R2 full
− 0.057 [ − 0.198, 0.078]
0.8055
0.0500
no
1
E2
Jev > R1
+0.144 [ − 0.044, 0.325]
0.0676
0.0167
no
1
E1-LOSO
Jev > R2 LOSO
+0.361 [0.189, 0.523]
<0.0001
0.0125
yes
1
E1-LOSO
Jev > R1
+0.063 [ − 0.144, 0.262]
0.2737
0.0250
no
2
E1r2
Jev > R1
+0.022 [ − 0.063, 0.107]
0.3065
0.0250
no
2
E1r2-LOSO
Jev > R2 LOSO
+0.387 [0.318, 0.455]
<0.0001
0.0167
yes
Table 11
Figure S1: C-Town network. (a) EPANET model used unchanged in rounds 1-4 and in the BATADAL study: 388 junctions, 429 pipes, 7 tanks, 1 reservoir, 11 pumps in five stations and 4 valves. (b) SCADA channels read by the evidence state: tank levels and inflow meters, pump flows and statuses, valve status, pump suction and discharge pressures, eight network pressures and the source outflow meter. Black rings mark the 12 pressure sensors of the BATADAL dataset.
Figure S2: Recall by subtype in rounds 2 and 3. (a) E1r2 and (b) E1r3 (in distribution) for calibrated Jev, raw Jev, R1, R2 full and R2 LOSO. (c) E2r2 and (d) E3 (new families, all subtypes unseen) for Jev, R1, R2 full and the product fusion of Jev and R2. Each cell is the recall of one subtype (window count in brackets); row-label colour gives the true class.
Figure S3: Confusion matrices of calibrated Jev on the eight C-Town sealed sets. Rows are true classes and columns Jev verdicts; cells give window counts and shading the share of the row. Panel titles give the set, round, family and cyberattack recall.
Figure S4: Benign-alarm veto on BATADAL. (a) ROC AUC for separating attack windows from benign-alarm windows (n = 107), with the difference between Jev and each method and its bootstrap 95% CI; the dotted line marks chance. (b) Operating points of R1 alone, Jev alone and the veto cascade in which Jev must confirm each R1 alarm, with their balanced accuracy.
Figure S5: Recall by subtype in round 4. (a) E1c (in distribution) and (b) E4c (look-alike family C) for Jev, R1, R2 full, the rule cascade, both reviewers alone and both LLM cascades. Row-label colour gives the true class and brackets the window count.
Figure S6: Confusion matrices in round 4. E1c (top) and E4c (bottom) for Jev, R1, R2 full, glm-5.2 alone and the LLM cascade with glm-5.2 as reviewer. Cells give window counts and shading the share of the row; panel titles give the macro-F1.
Figure S7: Transfer networks. (a) EPANET Net3: 92 junctions, 117 pipes, 3 tanks, 2 metered reservoirs, 2 pumps and a pressure-zone bypass opened and closed by a tank level; 8 network pressures and 3 pump-end pressures are monitored. (b) EPANET Net1, with its demand patterns resampled to a 1-h step, used in round 6: 9 junctions, 12 pipes, 1 tank, 1 metered reservoir and 1 pump started and stopped by the tank level; 4 network pressures and 1 pump-end pressure are monitored. Every tank carries a level and an inflow meter.
Figure S8: Accuracy and primary tests across networks. (a) Macro-F1 on C-Town (round 4), Net3 (round 5) and Net1 (round 6). Bars show the mean of the reference-family and family-C sets (n = 2 per network), error bars ± 1 s.d. and dots the individual sets (hollow, reference family; filled, family C); the mean is printed above each bar. The dashed line is the mean of R2 fitted on the evaluated network’s own labels. With n = 2 the s.d. only indicates the spread between the two sets. (b) The eight transfer primary tests: LLM cascade minus glm-5.2 alone (non-inferiority, margin − 0.03, dotted) and Jev minus R2 transfer (superiority), with 95% CIs. (c) Windows accepted by the agreement gate by true class, with accept precision and its Wilson 95% interval ( Eq. S2 ).
Figure S9: All metrics on the fresh sealed sets of rounds 4-6. (a) E1c and (b) E4c (C-Town, round 4); (c) N3-Ref and (d) N3-C (Net3, round 5); (e) N1-Ref and (f) N1-C (Net1, round 6). Each row is a method and the columns give macro-F1, accuracy and per-class recall; on C-Town, R2 C-Town is R2 full, and a dash marks a method not defined for the set. R2 own network is a reference that needs labels of the evaluated network.
Figure S10: Realised design, reviewer time and gate statistics. (a) Realised subtype mix of the transfer sets (50 windows per class). (b) Redraws per class before each accepted window (log scale): events that tripped no check (bars) and subtypes the network could not host (diamonds). (c) Mean reviewer time per decision alone and behind the agreement-gated screen on the windows timed one request at a time. (d) Accept rate, (e) accept precision and (f) accepted cyberattacks per set for the rule gate and the agreement gate. In c-f, bars show the mean over the six fresh sets (C-Town round 4, Net3 round 5 and Net1 round 6; n = 6), error bars ± 1 s.d. and white dots the individual sets; the mean is printed above each bar.
Figure S11: Recall by subtype on the transfer sets. (a) N3-Ref and (b) N3-C (Net3, round 5); (c) N1-Ref and (d) N1-C (Net1, round 6). Columns are Jev, R1, R2 fitted on C-Town, R2 fitted on the evaluated network (reference), glm-5.2 alone, the LLM cascade with glm-5.2 and the rule cascade; row-label colour gives the true class and brackets the window count.
Figure S12: Label scarcity on the transfer sets. Mean macro-F1 of R2 refitted on k labelled windows per class of the evaluated network (200 draws per k; shading, 5-95% range of the draws), against Jev and the rule cascade, which use no labels. The dashed vertical line marks the crossover k* ( Eq. 8 ), and the inset gives the share of draws at k = 1 in which R2 fell below Jev.
Figure S13: Secondary comparisons of rounds 4-6. Differences in macro-F1 with paired, class-stratified bootstrap 95% CIs, uncorrected for multiplicity; hollow markers indicate an interval that includes zero. (a) The cascades against their reviewers, their components and each other. (b) Jev and the cascades against R1, the LLMs alone, R2 LOSO, R2 full and R2 fitted on the evaluated network, the last reported as a reference. Colours give the network.
We evaluate Jev on ten dataset-defined application labels in CESNET-QUICEXT-25 using only the first ten packets' sizes, directions, and inter-packet times. To the best of our knowledge, this is the first empirical study of general-purpose decision models, represented here by Jev, for application classification of network flows. Across 52,000 records from 26 collection weeks following the training period, 40 fixed labeled examples raise Jev's accuracy from 9.80% to 28.42%. Random Forest and Extra Trees trained on 8,000 records achieve 69.95% and 66.80% and outperform Jev in every week. Increasing Jev's context to 150 examples yields 34.50% on the first test week. On a paired 100-record subset, Jev with 40 examples achieves 29% accuracy at a median request time of 0.750 s, versus 37% and 6.036 s for the generative language model OpenAI GPT-5.6 Sol with high reasoning effort through Azure; Jev also incurs lower API charges. The paired subset does not establish an accuracy advantage for either service, and the timing reflects different service configurations. Thus, labeled examples substantially improve Jev, but the tested Jev configurations remain less accurate than trained tree ensembles; unequal supervision budgets and fixed configurations prevent attributing the gap to a single cause.
Shenghe Xu, Lifan Mei
Amazon.com, Inc. · Xi’an Jiaotong-Liverpool University
Machine-learning Network Intrusion Detection Systems (IDS) depend on substantial labeled datasets and task-specific training, whereas Large Language Models (LLMs) detection can analyze flow records directly but incurs higher inference cost and latency, with less constrained outputs. This paper presents JEV-IDS, an open experimental general NIDS based on the Jev System One Model (SOM) to detect zero day intrusions Under label scarcity. JEV-IDS serializes one flow per request and asks JEV two questions: a binary attack probability and a finite-choice traffic category. Our results show that, at k=1, JEV was 4.8 times faster and 3.8 times cheaper than GPT-5.6 Luna, with 1.5 times higher novel-attack recall; it also produced 15 times fewer false alarms than a low-data Random Forest. Across 5,400 decisions on a 300-flow NSL-KDD pilot split, JEV achieved F1-Score 0.859, precision 0.941, recall 0.790, and novel-attack recall 0.838. Increasing k to 2 reduced its F1-Score to 0.839.
Paulo Severo, Silvio E. Quincozes, Amanda Dias
Graduate Program in Software Engineering, Federal University of Pampa (UNIPAMPA), Alegrete, Rio Grande do Sul, Brazil
Can a learned model capture how faults propagate through a large-scale network and use this knowledge to causally attribute customer impact to its underlying root cause? Existing root cause analysis techniques often rely on static rules, correlation heuristics, or topology-local reasoning, which struggle to generalize in dynamic environments where faults propagate across complex physical and logical dependencies. We present NetCause, a self-supervised learning-based framework that models network incidents as graph-temporal processes and uses counterfactual simulation to rank candidate root causes. This approach produces an interpretable ranking of root cause hypotheses and integrates naturally with operator-defined mitigation and remediation actions. We train the model on over 1,500 incidents collected over six months from a leading cloud provider's production network and evaluate it on 31 expert-labeled incidents. NetCause consistently improves root cause ranking quality in the regime most relevant to operational decision-making, achieving a 16.1% accuracy improvement over a rule-based heuristic baseline. While training is computationally intensive, inference is lightweight, requiring only seconds of GPU runtime per incident (well below typical telemetry collection latencies).