Artifact removal routinely precedes the classification of electrodermal activity (EDA), on the assumption that a cleaner signal supports a better decision. We tested this assumption in a virtual-reality (VR) balance-disturbance task. A residual gating network was trained on a benchmark with expert-corrected EDA, frozen, and applied to VR recordings, where raw and gated signals were classified by five published time-series methods under identical leave-one-participant-out evaluation. On the benchmark the gate detected artifacts well (median record AUROC 0.94) and reduced error inside artifact regions by 17.8%. In the VR task it did not improve classification. Changes in balanced accuracy ranged from -1.35 to +0.93 percentage points, no classifier improved and two lost accuracy, and all five were equivalent to raw input within +/- 3.32 points. The benefit was lost between waveform and decision. The correction that lowered waveform error also reduced skin conductance response detection in all 43 benchmark records. Processing left 92.8% of predictions unchanged, and the predictions it did change were corrected and corrupted at similar rates. The VR recordings also carried little contamination (an estimated 4.6% of samples), and even perfect localization of deliberately injected artifacts recovered only 3.3 points in the most sensitive classifier. A pooled association between artifact level and accuracy (11.3 points) disappeared within participants (0.1 points), showing how differences between people can make cleaning look useful. Preprocessing should be judged by the decision it supports, against an unprocessed arm.
Figures & tables
Property
EDABE v2 (benchmark)
VR balance disturbance
Unit
43 records
11 participants
Sampling rate
128 Hz
58.8 Hz
Reference
Expert mask and clean waveform
None
Evaluation
Fixed split, 33 training / 10 test records
11-fold leave-one-participant-out
Task windows
None
984 windows of 720 samples
Artifact burden
10.75% annotated
4.64% estimated
Table 1: Roles of the two corpora. The benchmark supports expert-referenced evaluation of the artifact handler; the VR corpus supplies the task labels.
Figure 1: Study design. (a) The benchmark (EDABE v2) provides expert-clean waveforms and point-wise artifact masks (thumbnail: record E01, raw trace with the expert reference inside the annotated span). Two checkpoints, CRG-A and CRG-B, were trained on the fixed split of 33 training and 10 test records and then frozen, and their detection, reconstruction and event preservation were scored against the reference. (b) The VR corpus provides labels for windows before and after each balance disturbance but no clean reference. The frozen CRG-B weights are applied without fitting on VR data, and raw and gated inputs pass through the same leave-one-participant-out folds (one held-out participant per fold, three seeds). The endpoint is the per-participant difference in balanced accuracy, with intervals from a participant-cluster bootstrap; family colours match figure 3 .
Figure 2: The continuous residual gate. (a) The raw trace x(t) is divided by a per-record robust scale s . A stem and nine dilated residual blocks (48 channels, dilation d from 1 to 256) feed a detection head, whose output p(t)=σ(ℓ(t)) lies between 0 and 1, and a residual head, whose output r(t)=4tanh(d(t)) lies between −4 and 4. Their product is added to the unmodified normalized trace, clamped at zero and multiplied by s (equations 1 and 2 ), so the output equals the input wherever p(t) is close to zero. Shaded modules are learned on the benchmark and then frozen; the other operations have no parameters. (b) The operator on a 14-s window of benchmark record E01 (CRG-A output, 128 Hz, unfiltered). The correction xgate−x=sp(t)r(t) is concentrated in the expert-annotated artifact span, where the output moves towards the expert reference x∗(t) . The bottom track shows the training targets that the mask defines: an L1 loss to x∗ inside the span and to the raw signal outside it. The detection head is trained against the mask throughout, and CRG-B adds distillation from a fold-local teacher. (c) Receptive field after each block, from 9 samples after the first block to 4089 samples (31.9 s at 128 Hz) after the last. The dashed line marks the length of one VR window (720 samples).
Figure 3: Effect of gating on VR classification. (a) Gated-minus-raw balanced accuracy for each model family. Open circles are the 11 held-out participants, each averaged over seeds; filled circles and bars are the mean and its 95% participant-cluster bootstrap interval. The shaded band is the ±3.32 pp margin tested by TOST, and dotted lines mark the tightest common margin ( ±2.37 pp). BF01 uses a JZS prior ( r=0.707 ) against a two-sided alternative. (b) TOST p as a function of the equivalence half-width δ , computed from the same participant differences (paired t , 10 degrees of freedom). A family is equivalent within ±δ wherever its curve lies below α=0.05 .
Stage
Measurement (checkpoint)
Result
Meaning
Contamination
Artifact share
10.75% annotated in the benchmark; 4.64% estimated in VR
Less to remove in VR
Detection
Record AUROC (CRG-A)
Median 0.9409; 36/43 records above 0.90; median probability 3.13 times median artifact fraction
Good ranking, too much intervention
Reconstruction
Artifact-region MAE (CRG-B)
0.1010 to 0.0830 \SIUnitSymbolMicroS ( −17.8 %)
Waveforms improved
Events
SCR peak F1 (CRG-B)
0.8980 to 0.8674; lower in 43/43 records
Events degraded
Features
MiniRocket displacement (CRG-A)
No change needed in 57.4% of windows, all changed; median cosine 0.363
Change where none was needed, partial alignment where it was
Table 2: From contamination to decision: where the expected benefit was lost. Each row is a separate measurement on the checkpoint shown.
Figure 4: Two measured benchmark excerpts at 128 Hz, unfiltered, selected to illustrate failure modes rather than typical behaviour: gate failure at high artifact burden in E01 (test record) and over-intervention at low burden in E14 (training record). Upper panels show the raw trace (blue) and the CRG-A output (vermillion), which is drawn beneath raw so that it is visible only where the gate moved the signal. The expert-clean reference (green, dotted) is drawn within annotated spans; outside them it equals raw in E14 and differs from raw by less than 0.0025\SIUnitSymbolMicroS in E01. Lower strips show reference minus raw (green) and gated minus raw (vermillion). Circles are reference SCR peaks, and crosses are detections on the raw or gated trace left unmatched by one-to-one matching within ±1 s. Dashed boxes mark the windows enlarged at right.
Figure 5: Detector behaviour on the benchmark (CRG-A; out-of-fold predictions for training records, test-ensemble predictions for test records). (a) Point-wise AUROC for each record; bars mark split medians, and 36 of 43 records exceed 0.90. (b) Observed artifact fraction in 100 fixed probability bins of width 0.01, with the number of samples per bin below on a log scale (34,312,745 samples in total). (c) Mean predicted probability against annotated artifact fraction for each record. All 43 records lie above the identity line.
MAE ( \SIUnitSymbolMicroS )
Red.
MAE
F1
Subset
n
raw → gated
(%)
lower
lower
Expert 1
21
0.134 → 0.097
27.6
21/21
21/21
Expert 2
22
0.069 → 0.070
−0.2
10/22
22/22
Training
33
0.102 → 0.085
16.3
25/33
33/33
Test
10
0.098 → 0.075
23.1
6/10
10/10
Table 3: CRG-B reconstruction and event outcomes by expert subset and split. MAE is averaged across records; reduction =100×(1−MAEgated/MAEraw) . The last two columns count records. Expert and split subsets overlap, so rows are not additive. In the test split alone, F1 fell by 2.50 pp ( p=0.002 ) and MAE did not change significantly ( p=0.084 ; paired Wilcoxon).
Figure 6: Waveform repair and event preservation on the 43 benchmark records, CRG-B continuous-strength student versus raw input. (a) Artifact-region mean absolute error to the expert-clean reference, on a log scale. (b) SCR peak F1 . Thin lines join each record and the thick line joins the means. (c) The same records as paired changes. Waveform error fell in 31 records and event F1 fell in all 43. Filled circles are records annotated by Expert 1 and open circles those annotated by Expert 2.
Figure 7: Decision changes after gating for each model family, with 2,952 prediction instances per family (984 windows × 3 seeds). Bars to the right count instances that became correct, and bars to the left those that became incorrect. Diamonds mark the net change, which equals the pooled accuracy difference in equation 5 . The right-hand columns give that difference and the share of instances whose class did not change.
Figure 8: Controlled artifact injection. (a, b) Change in balanced accuracy relative to contaminated input at each effective perturbed fraction for catch22 + RF and MiniRocket; contaminated input defines the zero line. Points are means over 11 participants with seeds averaged, and arms are offset horizontally for legibility. Vertical bars are 95% participant-cluster bootstrap intervals for the oracle-localization contrast. The clean-input arm reuses the native windows at every dose, so its curve shows the damage that perfect repair would undo; MiniRocket’s values vary across doses only because its transform is refitted at each dose. (c) The prespecified validity check (catch22 + RF, highest dose versus native input) and pooled nonzero-dose contrasts against contaminated input.
Figure 9: Participant composition behind the artifact–accuracy association. (a) Difference in raw-input accuracy between low- and high-proxy windows after a global median split and after a median split within each participant, with 95% participant-cluster bootstrap intervals. (b) Share of each participant’s windows above the global proxy median; without differences between participants every share would lie near 50%. Three participants (P06, P10 and P03) supply 50.4% of all high-proxy windows.
Oct 7, 2026·Lu Wang-Nöth, Hai Huang, Philipp Heiler +3
Institute for Applied Computer Science, University of the Bundeswehr Munich, Munich, Germany · brainboost GmbH, Munich, Germany · Graduate School of Engineering Science, The University of Osaka, Osaka, Japan +1