Physics-Attested Federated Learning: Securing Collaborative Anomaly Detection in Critical Water Infrastructure
Authors: Jeff Nijsse, Shu Su, Benjamin Oholeguy, Sreenivas Sremath Tirumala
Organizations: Department of Software Engineering, RMIT University, Hanoi, Vietnam · Department of Mathematical Sciences, Auckland University of Technology, New Zealand · Independent Researcher, New Zealand · School of Science, Engineering and Technology, RMIT University, Hanoi, Vietnam
Federated learning enables industrial operators to train shared intrusion detection models without disclosing proprietary operational telemetry. However, existing defenses operate strictly in update space, leaving aggregators blind to data poisoning; model updates derived from fabricated telemetry remain indistinguishable from honest contributions. We repurpose cyber-physical process invariants, such as conservation laws and actuator couplings, from runtime detection heuristics into a verifiable admission requirement for federated updates, mined automatically from clean operational data. We evaluate this admission gate across two physical water testbeds (SWaT, WADI) and a distribution benchmark (BATADAL), testing seven aggregation rules against telemetry fabrication, exposure-only replay poisoning, and an invariant-aware adaptive adversary. Across three testbeds the mined invariants reject none of 100 honest shards and all naively fabricated ones, including optimised perturbations that FoolsGold admits in full. On real telemetry, five mined invariants detect 12 of SWaT's 35 attacks, while nine invariants detect 20, with no honest shard rejected. With nine rules, the physics gate recovers 69--100% of the targeted-attack recall lost to replay poisoning, and 54--100% of that lost to fabricated telemetry, across five standard aggregators. To reconcile physical admission control with federated data privacy, we show invariant compliance using zero-knowledge proofs (zk-SNARKs) to allow clients to prove batch adherence without revealing operational telemetry.
Figures & tables
Category
Method
Mechanism
Key Assumption
Coordinate-wise statistics
Coordinate-wise median [ 54 ]
Evaluates updates coordinate-wise, selecting the median value for each parameter dimension.
Adversarial perturbations occupy an empirical minority in every coordinate.
Trimmed mean [ 54 ]
Discards the upper and lower β -fractions along each coordinate before averaging.
Adversarial perturbations occupy an empirical minority in every coordinate.
Geometric proximity
Krum [ 8 ]
Identifies the single update that minimises the sum of squared Euclidean distances to nearest neighbours.
Benign updates form a dense spatial cluster in parameter space.
Multi-Krum [ 8 ]
Averages multiple low-scoring updates selected via Krum’s distance metric.
Honest updates maintain geometric coherence relative to Byzantine updates.
Surrogate-autoencoder error maximisation within ±2.5σ
Autoencoder reconstruction
Update Baselines
Update ( Δi )
Sign-flip, gradient scaling, free-riding, min-max
Krum, Median, Trimmed Mean
Table 2: Classification of evaluated adversarial threat vectors across data and update spaces.
Record
Clients (malicious)
Rules
Attacks
Modes
Seeds
Invariant sets
Cells
SWaT
10 (3)
7
8
5
5
narrow (5), wide (9)
907
WADI
10 (3)
3
2
5
3
default (7)
90
BATADAL
5 (2)
3
2
4
3
mined (10)
72
Simulator
10 (3)
6
1
5
3
exact
51
Table 3: Experimental grid across industrial benchmark records, detailing participant counts, aggregation rules, attack modes, federations, seeds, invariant sets, and total trained models.
Record, setting
Rules
Couplings / balances
Honest rows violating
Attacks covered
Attack rows covered
Share
SWaT, narrow
5
3 / 2
0.21 %
12 / 35
1,137 / 10,931
10 %
SWaT, wide
9
6 / 3
0.36 %
20 / 35
8,991 / 10,931
82 %
WADI, default
7
7 / 0
0.14 %
3 / 14
395 / 1,996
20 %
Table 4: Coverage of labelled attacks across invariant sets on SWaT and WADI. An attack segment is covered when its data violation rate strictly exceeds the 1% threshold. Honest row violations are audited over the first 60,000 observations following the discovery slice.
malicious batches admitted
honest shards
Record (rules)
fabricated
replay, before
replay, after proj.
malicious excluded
rejected
worst violating
BATADAL (10)
0 / 6
0 / 6
6 / 6
0.0 of 2
0 / 9
0.00 %
SWaT narrow (5)
0 / 15
4 / 15
12 / 15
0.6 of 3
0 / 35
0.32 %
SWaT wide (9)
0 / 15
0 / 15
6 / 15
1.8 of 3
0 / 35
0.51 %
WADI (7)
0 / 9
1 / 9
3 / 9
2.0 of 3
0 / 21
0.13 %
Table 5: Batch admission and client exclusion rates across industrial benchmark records before and after adaptive projection under the 1% gate.
PA-FL Lite at k=32 sampled rows
Circuit size (R1CS constraints)
297,736
Proof generation on the plant (peak RAM)
12.0 s (3.8 GB)
Proof verification at the server
290 ms
Proof size πi
806 B
Honest batch accepted, P(v≤1) , 19,992 draws
98.5 %
Replayed attack batch detected, P(v>1)
99.8 %
Table 6: PA-FL Lite at the operating point k=32 on wide SWaT with N=1,024 rows per batch (Groth16 over BN128, Apple M3 Pro): the cost of one attestation and the sampling guarantee it buys. The full benchmark, including k=16 , is Appendix Table 8 .
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
damage
damage removed (%)
rule admits
Rule
n
(points)
proj.
gate
naive
phys.-aware
gated
Replay, wide set (5 seeds; check admits 0.40 after projection)
FedAvg
2
13.8
24
71
1.00
1.00
0.60
Median
3
6.2
42
103
1.00
1.00
0.60
Norm clip
3
7.4
30
104
1.00
1.00
0.60
Trimmed mean
3
5.1
11
91
0.34
0.34
0.17
Appendix
Table 7: Targeted-recall damage and removal across aggregation rules on SWaT (5 seeds). n is the number of informative seeds ( ≥1.0 point of naive damage against the honest-only reference). Removal above 100% means the gated federation scored above the honest-only reference, within seed variance.
Protocol Metric / Circuit Property
k=16 samples
k=32 samples
R1CS arithmetic constraints (total)
148,888
297,736
Constraints per sample ( 2 paths, 2 leaves, physics)
9,306
9,304
Public / private circuit inputs
19 / 928
35 / 1,856
Circuit compilation time (Circom)
7.47 s
15.0 s
Proving key size ( .zkey )
108.7 MB
217.4 MB
Verification key size ( .vk )
6.1 kB
8.9 kB
Appendix
Table 8: PA-FL Lite full benchmark across sample sizes k . Attestation outcomes are for the seed-0 SWaT batches under a budget of vmax=1 violations and umax=8 inapplicable checks; sampling soundness is estimated from 19,992 random index draws over 56 honest windows.
Bangladesh University of Business and Technology (BUBT), Mirpur-2, Dhaka-1216, Bangladesh · Jagannath University, Dhaka, Bangladesh · University of New South Wales (UNSW), ACT 2601, Australia