Physics-Attested Federated Learning: Securing Collaborative Anomaly Detection in Critical Water Infrastructure
Authors: Jeff Nijsse, Shu Su, Benjamin Oholeguy, Sreenivas Sremath Tirumala
Organizations: Department of Software Engineering, RMIT University, Hanoi, Vietnam · Department of Mathematical Sciences, Auckland University of Technology, New Zealand · Independent Researcher, New Zealand · School of Science, Engineering and Technology, RMIT University, Hanoi, Vietnam
Federated learning enables industrial operators to train shared intrusion detection models without disclosing proprietary operational telemetry. However, existing defenses operate strictly in update space, leaving aggregators blind to data poisoning; model updates derived from fabricated telemetry remain indistinguishable from honest contributions. We repurpose cyber-physical process invariants, such as conservation laws and actuator couplings, from runtime detection heuristics into a verifiable admission requirement for federated updates, mined automatically from clean operational data. We evaluate this admission gate across two physical water testbeds (SWaT, WADI) and a distribution benchmark (BATADAL), testing seven aggregation rules against telemetry fabrication, exposure-only replay poisoning, and an invariant-aware adaptive adversary. Across three testbeds the mined invariants reject none of 100 honest shards and all naively fabricated ones, including optimised perturbations that FoolsGold admits in full. On real telemetry, five mined invariants detect 12 of SWaT's 35 attacks, while nine invariants detect 20, with no honest shard rejected. With nine rules, the physics gate recovers 69--100% of the targeted-attack recall lost to replay poisoning, and 54--100% of that lost to fabricated telemetry, across five standard aggregators. To reconcile physical admission control with federated data privacy, we show invariant compliance using zero-knowledge proofs (zk-SNARKs) to allow clients to prove batch adherence without revealing operational telemetry.
Figures & tables
Category
Method
Mechanism
Key Assumption
Coordinate-wise statistics
Coordinate-wise median [ 54 ]
Evaluates updates coordinate-wise, selecting the median value for each parameter dimension.
Adversarial perturbations occupy an empirical minority in every coordinate.
Trimmed mean [ 54 ]
Discards the upper and lower β -fractions along each coordinate before averaging.
Adversarial perturbations occupy an empirical minority in every coordinate.
Geometric proximity
Krum [ 8 ]
Identifies the single update that minimises the sum of squared Euclidean distances to nearest neighbours.
Benign updates form a dense spatial cluster in parameter space.
Multi-Krum [ 8 ]
Averages multiple low-scoring updates selected via Krum’s distance metric.
Honest updates maintain geometric coherence relative to Byzantine updates.
Surrogate-autoencoder error maximisation within ±2.5σ
Autoencoder reconstruction
Update Baselines
Update ( Δi )
Sign-flip, gradient scaling, free-riding, min-max
Krum, Median, Trimmed Mean
Table 2: Classification of evaluated adversarial threat vectors across data and update spaces.
Record
Clients (malicious)
Rules
Attacks
Modes
Seeds
Invariant sets
Cells
SWaT
10 (3)
7
8
5
5
narrow (5), wide (9)
907
WADI
10 (3)
3
2
5
3
default (7)
90
BATADAL
5 (2)
3
2
4
3
mined (10)
72
Simulator
10 (3)
6
1
5
3
exact
51
Table 3: Experimental grid across industrial benchmark records, detailing participant counts, aggregation rules, attack modes, federations, seeds, invariant sets, and total trained models.
Record, setting
Rules
Couplings / balances
Honest rows violating
Attacks covered
Attack rows covered
Share
SWaT, narrow
5
3 / 2
0.21 %
12 / 35
1,137 / 10,931
10 %
SWaT, wide
9
6 / 3
0.36 %
20 / 35
8,991 / 10,931
82 %
WADI, default
7
7 / 0
0.14 %
3 / 14
395 / 1,996
20 %
Table 4: Coverage of labelled attacks across invariant sets on SWaT and WADI. An attack segment is covered when its data violation rate strictly exceeds the 1% threshold. Honest row violations are audited over the first 60,000 observations following the discovery slice.
malicious batches admitted
honest shards
Record (rules)
fabricated
replay, before
replay, after proj.
malicious excluded
rejected
worst violating
BATADAL (10)
0 / 6
0 / 6
6 / 6
0.0 of 2
0 / 9
0.00 %
SWaT narrow (5)
0 / 15
4 / 15
12 / 15
0.6 of 3
0 / 35
0.32 %
SWaT wide (9)
0 / 15
0 / 15
6 / 15
1.8 of 3
0 / 35
0.51 %
WADI (7)
0 / 9
1 / 9
3 / 9
2.0 of 3
0 / 21
0.13 %
Table 5: Batch admission and client exclusion rates across industrial benchmark records before and after adaptive projection under the 1% gate.
PA-FL Lite at k=32 sampled rows
Circuit size (R1CS constraints)
297,736
Proof generation on the plant (peak RAM)
12.0 s (3.8 GB)
Proof verification at the server
290 ms
Proof size πi
806 B
Honest batch accepted, P(v≤1) , 19,992 draws
98.5 %
Replayed attack batch detected, P(v>1)
99.8 %
Table 6: PA-FL Lite at the operating point k=32 on wide SWaT with N=1,024 rows per batch (Groth16 over BN128, Apple M3 Pro): the cost of one attestation and the sampling guarantee it buys. The full benchmark, including k=16 , is Appendix Table 8 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
damage
damage removed (%)
rule admits
Rule
n
(points)
proj.
gate
naive
phys.-aware
gated
Replay, wide set (5 seeds; check admits 0.40 after projection)
FedAvg
2
13.8
24
71
1.00
1.00
0.60
Median
3
6.2
42
103
1.00
1.00
0.60
Norm clip
3
7.4
30
104
1.00
1.00
0.60
Trimmed mean
3
5.1
11
91
0.34
0.34
0.17
Appendix
Table 7: Targeted-recall damage and removal across aggregation rules on SWaT (5 seeds). n is the number of informative seeds ( ≥1.0 point of naive damage against the honest-only reference). Removal above 100% means the gated federation scored above the honest-only reference, within seed variance.
Protocol Metric / Circuit Property
k=16 samples
k=32 samples
R1CS arithmetic constraints (total)
148,888
297,736
Constraints per sample ( 2 paths, 2 leaves, physics)
9,306
9,304
Public / private circuit inputs
19 / 928
35 / 1,856
Circuit compilation time (Circom)
7.47 s
15.0 s
Proving key size ( .zkey )
108.7 MB
217.4 MB
Verification key size ( .vk )
6.1 kB
8.9 kB
Appendix
Table 8: PA-FL Lite full benchmark across sample sizes k . Attestation outcomes are for the seed-0 SWaT batches under a budget of vmax=1 violations and umax=8 inapplicable checks; sampling soundness is estimated from 19,992 random index draws over 56 honest windows.
Federated learning enables privacy-conscious collaboration for network intrusion detection without centralizing sensitive traffic data, yet its deployment in operational environments must simultaneously satisfy three competing requirements: formal differential privacy guaranties, tolerance to Byzantine-adversarial participants, and reliable detection coverage across severely imbalanced attack categories. Existing literature treats these properties as independently composable, an assumption that this paper challenges both theoretically and empirically. In this paper, we study how these requirements interact in class-imbalanced federated NIDS and introduce geometric indistinguishability as a conceptual lens for a regime in which privacy-induced dispersion in client updates can make minority-class signals harder for robust aggregation to preserve. Using UNSW-NB15 as a case study, we evaluate DP-SGD combined with coordinate-wise median under label-flip and model-poisoning attacks, with threat coverage assessed across attack categories. Our results provide initial evidence that the joint use of privacy noise and robust aggregation can disproportionately degrade detection of rare attacks relative to majority classes. We also show that part of the observed collapse under strong privacy can arise from training miscalibration, while a residual performance floor may remain for ultra-rare categories even after epsilon-dependent tuning. These findings motivate studying privacy, robustness, and rare-attack coverage jointly rather than as independently composable properties, and suggest that aggregation-aware modeling and sample-aware evaluation are promising directions for trustworthy federated NIDS.
Adrita Rahman Tory, ABM Shawkat Ali, Md Abu Layek +1
Bangladesh University of Business and Technology (BUBT), Mirpur-2, Dhaka-1216, Bangladesh · Jagannath University, Dhaka, Bangladesh · University of New South Wales (UNSW), ACT 2601, Australia
Federated learning enables multiple parties to train a shared model without centralizing raw data with the help of an aggregator, but introduces integrity risks once participants or infrastructure are not fully trustworthy. Two requirements are particularly important: robustness to poisoned or Byzantine client updates, and verifiability of the aggregator so that clients or third parties can audit the reported aggregation without learning individual updates. Existing work has largely treated these goals separately, and efficient public verifiability for robust, outlier-excluding aggregation remains limited. We present a verifiable federated learning protocol that makes a robust aggregation pipeline publicly auditable. Our design combines cryptographic commitments with non-interactive zero-knowledge proofs to certify both (i) cosine-similarity-based outlier exclusion and (ii) aggregation over the selected set, without revealing individual client updates to verifiers. In experiments under representative poisoning attacks, our method maintains high accuracy, with an average accuracy loss below 4% across the evaluated configurations, while keeping verification overhead practical: proof artifacts can be generated and verified within minutes at the scale studied. In summary, our results show that robust outlier exclusion and public verifiability can be jointly achieved in a federated learning setting.
Federated Learning (FL) systems are susceptible to adversarial attacks, such as model poisoning attacks and backdoor attacks. Existing defense mechanisms face critical limitations in deployments, such as relying on impractical assumptions (e.g., adversaries acknowledging the presence of attacks before attacking) or undermining accuracy in model training, even in benign scenarios. To address these challenges, we propose CustodianFL, a two-staged anomaly detection method specifically designed for FL deployments. In the first stage, it flags suspicious client activities. In the second stage that is activated only when needed, it further examines these candidates using Three-Sigma Rule to identify and exclude truly malicious local models from FL training. To ensure integrity and transparency within the FL system, CustodianFL integrates zero-knowledge proofs, enabling clients to cryptographically verify the server's detection process without relying on the server's goodwill. CustodianFL operates without unrealistic assumptions and avoids interfering with FL training in attack-free scenarios. It bridges the gap between theoretical advances in FL security and the practical demands of real FL systems. Experimental results demonstrate that CustodianFL consistently delivers performance comparable to benign cases, highlighting its effectiveness in identifying and eliminating malicious models with high accuracy.
Shanshan Han, Wenxuan Wu, Baturalp Buyukates +4
University of California, Irvine · Texas A&M University · University of Birmingham +3