XPhysICS: Cross-Physical-Domain Threat Grounding for Industrial Control Systems Security
Organizations: University of Utah, Salt Lake City, Utah, USA · Kahlert School of Computing, University of Utah, Salt Lake City, Utah, USA
Abstract
Industrial control system attacks are usually documented in terms of the plant where they occurred: its sensors, actuators, process stages, and control logic. Yet many attacks express a more general physical pattern--such as suppressing flow, corrupting chemical dosing, or driving a vessel toward overflow--that may also matter in a different plant. The challenge is deciding when such a threat remains meaningful on a new system rather than relying on similar component names or broad semantic labels. We present XPhysICS, a methodology for grounding documented cyber-physical threats onto a specific target system. XPhysICS converts source evidence into a provenance-linked description of what is manipulated, what physical consequence is expected, and what observations the evidence calls for. Once this analyst-guided abstraction, its vocabulary and schema version, and a target contract are fixed, XPhysICS applies deterministic grounding checks. An accepted result can be represented as a validation slice that records the mapped roles, signals, dependencies, and context intended to support later evaluation. We study 83 threat abstractions across continuous-process and manufacturing sources using separate evaluation denominators. The continuous-process study evaluates 78 abstractions against target contracts spanning water treatment, water distribution, hydropower, and chemical processes. Selected cases are exercised through controlled perturbations of simulator-role signals. We also test compatibility with several analysis styles, including the released upstream GeCo implementation, and conduct a three-objective, one-target realizability study using a paper-derived search reproduction. Across these evaluated settings, the results support treating explicit target checks and traceable evidence as separate from semantic similarity alone.
Figures & tables
| Construct | Operational definition | Origin / assignment |
|---|---|---|
| Effect family | Normalized physical-consequence category drawn from a closed, versioned extraction vocabulary. Target grounding operates over the narrower grounding vocabulary declared in scope by the target contract (Section 3.3 ). | OTThreat-informed vocabulary; analyst- or parser-assigned. |
| Manipulated roles | Source-side component roles manipulated by the threat, represented by an identifier and functional role with optional process-stage context. | XPhysICS -specific; analyst- or parser-assigned. |
| Consequence roles | Components or consequence records associated with the realized physical effect following manipulation of . | XPhysICS -specific; assigned from supporting source evidence. |
| Observability obligations | Signals or observations indicated by the source evidence as relevant to assessing the stated effect; these obligations seed target-side slice construction (Section 3.4 ). | XPhysICS -specific; may be empty when the source evidence does not identify corresponding observations. |
| Provenance | Evidence-lineage metadata linking populated abstraction fields to the source document, supporting excerpt, extraction method, and confidence assignment. | OTThreat-informed; provenance metadata are recorded by the evidence workflow, while confidence may involve analyst or reviewer judgment. |
| Contract component | Role in XPhysICS |
|---|---|
| Effect-family scope | Declares the source-threat effect families within the target’s current grounding scope. |
| Target signals and roles | Defines the target signal-and-role surface over which candidate mappings are evaluated. |
| Dependency relationships | Provides target relationships used during validation-slice construction and dependency/connectivity evaluation. |
| Threat-surface scope | Declares the boundary directions and threat classes represented by the target study. |
| Target rules and context | Defines the target rules and operating context used to determine rule-surface applicability and support subsequent slice evaluation. |
| Evaluation object | Result | Interpretation |
|---|---|---|
| Structured threat corpus | 83 threats from 12 documents; 55 high-confidence, 18 medium-confidence, 10 low-confidence. Continuous-process matrix sources: SWaT 51, WADI 17, OilTreatment 10 ( ). Manufacturing extension: Fischertechnik 5. | Provides typed source semantics across water, oil/chemical, and manufacturing evidence for deterministic target grounding. |
| WADI source deepening | Naive PDF parsing: 2 rows; XPhysICS : 15 A1 threats + 14 A2 attack windows | Source-specific extraction recovers concrete attack points, effect families, and timing/observability context missed by table scraping. |
| XPhysICS evidence ledger | 333 curated threats; 1,947 field-level evidence spans; 100% core-field coverage | Our OTThreat-aligned audit layer links paper-facing threat objects to source evidence; it is provenance support, not the grounding algorithm. |
| Manual extraction sample | 20 sampled items: 14 full pass, 6 partial, 0 fail | A bounded spot-check found no outright abstraction failures among sampled paper-facing records. |
| Grounding and target case | Target-side result and interpretation |
|---|---|
| SWaT WT: Tank-level high spoof. Positive level offset on the WT tank slice; a target-side probe, not a direction-matched replay of SWaT attack 36. | Nominal-subtracted analysis yields three invariant violations: two rate-of-change and one mass-balance. This is the cleanest WT invariant case. |
| SWaT WT: Pump-flow low spoof. Pump/flow suppression instantiated on the WT slice. | Raw analysis yields nine correlation violations, but nominal-subtracted analysis yields zero. We treat this as nominal-confounded and do not count it as a clean invariant success. |
| WADI WD: Supply-flow low spoof. Low-flow effect instantiated on the WD supply slice. | No baseline invariant violations occur; the first firing appears only at the perturbation boundary with shallow range/correlation evidence. We report this as a near-threshold boundary case. |
| WADI WD: Supply-flow high idle. High-flow/pressure effect instantiated during an idle supply condition. | The slice produces 11 range violations, giving a clean WD flow/pressure case. |
| WADI WD: NaOCl level step. Dosing/storage-level perturbation instantiated on the WD chemical slice. | The slice produces three mass-balance/rate-of-change violations. This is a clean WD dosing case, although it is outside the GeCo predictive surface. |
| Consumer family | WT slice results | WD slice results | Role in evaluation |
|---|---|---|---|
| Invariant checks | Raw 2/2; 1 clean case remains after nominal subtraction. | 2/3; high-flow and NaOCl are strong, low-flow is partial. | Native slice consumer; preserves WT nominal caveat. |
| GeCo-style predictive ( Wolsing et al., 2025 ) | 2/2 hits; nominal FPR 0.0; mean F1 0.4167. | 2/3 hits; nominal FPR 0.0; mean F1 0.1025; NaOCl outside modeled surface. | Bounded predictive compatibility lane. |
| SAIN-style state-aware ( Abbas et al., 2024 ) | 2/2 hits; nominal FPR improves from 0.0167 to 0.0; mean F1 0.66. | 3/3 hits; nominal FPR 0.0; mean F1 1.0. | Bounded state-aware compatibility lane. |
| SCAPHY-style phase-aware ( Ike et al., 2022 ) | 2/2 hits; nominal FPR 0.0; mean F1 0.5661. | 3/3 hits; recovers low-flow case; nominal FPR 0.0; mean F1 0.57. | Bounded phase/context-aware compatibility lane. |
| Case | Effect / perturbation | Observed response | Interpretation |
|---|---|---|---|
| SWaT Hydro: reservoir level | overflow, HY_Res_Level | range, rate-of-change fire (2/2) | Both declared rule types respond |
| SWaT Hydro: turbine speed | speed excursion, HY_Speed_Pct | range, rate-of-change fire (2/2) | Both declared rule types respond |
| SWaT Hydro: flow | pressure/flow, HY_Flow | range fires only (1/3) | Limited observed dependency-rule response |
| WADI GRFICS: tank pressure | pressure/flow, TE_Tank_Pressure | rate-of-change fires only (1/2) | Limited observed rule response |
| WADI GRFICS: tank level | overflow, TE_Tank_Level | rate-of-change and causality responses | Mixed observed rule response |
| WADI GRFICS: feed-1 flow | pressure/flow, TE_Feed1_Flow | causality, range fire (2/3) | Broadest observed GRFICS rule response |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Evidence object | Appendix detail |
|---|---|
| Structured extraction | Result: 83 threats from 12 documents; 55 high-confidence, 18 medium-confidence, and 10 low-confidence. The source set includes SWaT ( Mathur and Tippenhauer, 2016 ) (51), WADI ( Ahmed et al., 2017 ) (17), OilTreatment (10), and Fischertechnik (5). Claim boundary: these are coverage and confidence counts for the paper-facing corpus, not a full extraction precision/recall evaluation. |
| WADI ( Ahmed et al., 2017 ) deepening | Result: naive PDF parsing yields 2 rows; XPhysICS recovers 15 A1 threats and 14 A2 attack windows, a 7.5 gain in attack-level threat records. Claim boundary: A1 seeds extraction, A2 contributes timing/observability windows, and A3 provides clean context; this is source-specific extraction, not a general PDF-table benchmark. |
| CISS ( iTrust, Centre for Research in Cyber Security, 2019 ) semantics | Result: 187 parsed rows, 73 process-relevant rows, and 20 high-confidence derived threats from CISS ( iTrust, Centre for Research in Cyber Security, 2019 ) . Claim boundary: CISS contributes source semantics and observability support; we do not treat it as perturbation replay. |
| Manual sample | Result: 20 sampled items: 14 full pass, 6 partial, and 0 fail. Claim boundary: this bounded spot-check found no outright abstraction failures, but it is not a full precision/recall measurement. |
| XPhysICS evidence ledger | Result: 333 curated threats, 1,947 field-level evidence spans, and 100% core-field coverage. Claim boundary: the ledger is an artifact of XPhysICS , aligned with OTThreat ( Paul et al., 2024 ) ; it supports provenance and auditability, not the grounding algorithm itself. |
| Grounding case | Detailed outcome |
|---|---|
| SWaT WT: tank-level high spoof | Adequacy: target signals declared in the WT contract; 100% source/target/dependency coverage; timing evaluable. Outcome: nominal-subtracted delta yields 3 violations: 2 rate-of-change and 1 mass-balance. Interpretation: clean WT invariant case. |
| SWaT WT: pump-flow low spoof | Adequacy: target signals declared in the WT contract; 100% source/target/dependency coverage; timing evaluable. Outcome: raw analysis yields 9 correlation violations, but nominal-subtracted delta yields 0. Interpretation: nominal-confounded; not counted as a clean invariant success. |
| WADI WD: supply-flow low spoof | Adequacy: target signals declared in the WD contract; 100% source/target/dependency coverage; partial verdict. Outcome: no baseline invariant violations; first firing occurs at with shallow range/correlation evidence. Interpretation: near-threshold boundary case. |
| WADI WD: supply-flow high idle | Adequacy: target signals declared in the WD contract; 100% source/target/dependency coverage; strong verdict. Outcome: 11 range violations. Interpretation: clean WD flow/pressure case. |
| WADI WD: NaOCl level step | Adequacy: target signals declared in the WD contract; 100% source/target/dependency coverage; strong verdict. Outcome: 3 mass-balance/rate-of-change violations. Interpretation: clean WD dosing case; outside the GeCo-style predictive surface. |
| Case | Expected rule types | Viol. | Earl. | Observed rule types (expected hit / total) |
|---|---|---|---|---|
| swat_hydro_uc1 / hydro_reservoir_level_high_spoof | range; rate_of_change | 23 | 20 | range; rate_of_change (2/2) |
| swat_hydro_uc1 / hydro_speed_pct_high_spoof | range; rate_of_change | 23 | 20 | range; rate_of_change (2/2) |
| swat_hydro_uc1 / hydro_flow_low_spoof | range; correlation; causality | 23 | 20 | range; rate_of_change (1/3 expected; additional rate_of_change) |
| wadi_grfics_uc1 / grfics_tank_pressure_high_spoof | range; rate_of_change | 2 | 20 | rate_of_change (1/2) |
| wadi_grfics_uc1 / grfics_tank_level_high_spoof | range; rate_of_change | 3 | 20 | causality; rate_of_change (1/2 expected; additional causality) |
| wadi_grfics_uc1 / grfics_feed1_flow_low_spoof | range; correlation; causality | 25 | 20 | causality; range; rate_of_change (2/3 expected; additional rate_of_change) |
| Consumer family | Bounded implementation details |
|---|---|
| Invariant checks | Slice inputs: signal ranges, correlations, mass-balance relationships, and rate-of-change constraints. Calibration: thresholds from nominal traces. Scoring: violation count and nominal-subtracted delta. Boundary: native XPhysICS consumer and primary validation lane. |
| GeCo-style predictive ( Wolsing et al., 2025 ) | Slice inputs: time-series windows and signal correlations. Calibration: nominal-trace training and threshold calibration. Scoring: hit/miss, nominal FPR, and F1. Boundary: bounded predictive lane inspired by GeCo; not upstream code reuse or faithful reproduction. |
| SAIN-style state-aware ( Abbas et al., 2024 ) | Slice inputs: target signals and coarse discrete process state. Calibration: state-conditioned thresholds from nominal traces. Scoring: hit/miss, nominal FPR, and F1. Boundary: bounded state-aware lane; not full SAIN reproduction. |
| SCAPHY-style phase-aware ( Ike et al., 2022 ) | Slice inputs: target signals, phase labels, and operating context. Calibration: phase-segmented nominal calibration. Scoring: hit/miss, nominal FPR, and F1. Boundary: bounded phase/context-aware lane; not full SCAPHY reproduction. |
| Prototype adapter | Verified prototype evidence and claim boundary |
|---|---|
| Post-analysis classification/typing | Result: at 1 FA/hr, baseline P/R/F1 was 0.375/1.000/0.545; adding XPhysICS semantics yielded 1.000/0.667/0.800, corresponding to a precision gain of +0.625 and F1 gain of +0.255. What it exercises: post-analysis typing over detector windows using grounded temporal predicates and provenance. Boundary: these measurements characterize an earlier XPhysICS prototype adapter only. SCADMAN is cited as motivating literature and was not independently reproduced here; the evaluated prototype does not implement SCADMAN’s control-flow-attestation mechanism and therefore does not measure SCADMAN performance or mechanism fidelity. |
| Provenance-workflow adapter | Result: 17/17 rule templates instantiated, 2,074 typed physics-consistent mappings, and 0.95 reproducibility across reruns. What it exercises: rule expansion, provenance-pack generation, eligibility checks, and typed path export. Boundary: these measurements characterize the prototype provenance workflow only. ICSTracker is cited as motivating literature; because implementation fidelity has not been independently established, these results do not constitute an ICSTracker reproduction or performance evaluation. |
| Extension | Key Result | What It Exercises | Why It Is Bounded | Paper Loc. |
|---|---|---|---|---|
| CISS observability | 20 high-confidence threats; 13/20 full-WT grounded; avg score 2.45 | Source semantics breadth; slice-minimality | Source semantics only, not perturbation replay | § 5.1 |
| TXT S7 Fischertechnik | 5/5 threats grounded; 102 tags; 26/28 components | Manufacturing/program-artifact grounding | Separate manufacturing denominator; not continuous-process matrix | § 5.8 , App. F |
| CrossPLC control study | MiniSWaT/SPHERE-WT: 6/6 role families; direct matrix delta 0 | Controller-code enrichment; provenance | Optional enrichment; does not improve grounding percentages | § 5.8 , App. F |
| STL/RTAMT guardrail | WADI/WD: accepted, rejected; accepted, rejected | Formal slice-level temporal guardrail | One bounded WADI WD branch; not full verification | § 5.8 , App. F |
| EPIC/Hydro scenarios | 6 direct + 2 bridge scenarios; 0.8 semantic coverage; 3/3 strong support | Power/hydro domain breadth | Scenario support only; not EPIC attack-trace replay | § 5.8 , App. F |
| PowerDuck/ GOOSE | 16 IPAL files; 1,133,589 packets; flooding/insertion/replay/suppression | Protocol-level evidence complement | Protocol complement to EPIC; not target replay | § 5.8 , App. F |