Artifact removal routinely precedes the classification of electrodermal activity (EDA), on the assumption that a cleaner signal supports a better decision. We tested this assumption in a virtual-reality (VR) balance-disturbance task. A residual gating network was trained on a benchmark with expert-corrected EDA, frozen, and applied to VR recordings, where raw and gated signals were classified by five published time-series methods under identical leave-one-participant-out evaluation. On the benchmark the gate detected artifacts well (median record AUROC 0.94) and reduced error inside artifact regions by 17.8%. In the VR task it did not improve classification. Changes in balanced accuracy ranged from -1.35 to +0.93 percentage points, no classifier improved and two lost accuracy, and all five were equivalent to raw input within +/- 3.32 points. The benefit was lost between waveform and decision. The correction that lowered waveform error also reduced skin conductance response detection in all 43 benchmark records. Processing left 92.8% of predictions unchanged, and the predictions it did change were corrected and corrupted at similar rates. The VR recordings also carried little contamination (an estimated 4.6% of samples), and even perfect localization of deliberately injected artifacts recovered only 3.3 points in the most sensitive classifier. A pooled association between artifact level and accuracy (11.3 points) disappeared within participants (0.1 points), showing how differences between people can make cleaning look useful. Preprocessing should be judged by the decision it supports, against an unprocessed arm.
Figures & tables
Property
EDABE v2 (benchmark)
VR balance disturbance
Unit
43 records
11 participants
Sampling rate
128 Hz
58.8 Hz
Reference
Expert mask and clean waveform
None
Evaluation
Fixed split, 33 training / 10 test records
11-fold leave-one-participant-out
Task windows
None
984 windows of 720 samples
Artifact burden
10.75% annotated
4.64% estimated
Table 1: Roles of the two corpora. The benchmark supports expert-referenced evaluation of the artifact handler; the VR corpus supplies the task labels.
Figure 1: Study design. (a) The benchmark (EDABE v2) provides expert-clean waveforms and point-wise artifact masks (thumbnail: record E01, raw trace with the expert reference inside the annotated span). Two checkpoints, CRG-A and CRG-B, were trained on the fixed split of 33 training and 10 test records and then frozen, and their detection, reconstruction and event preservation were scored against the reference. (b) The VR corpus provides labels for windows before and after each balance disturbance but no clean reference. The frozen CRG-B weights are applied without fitting on VR data, and raw and gated inputs pass through the same leave-one-participant-out folds (one held-out participant per fold, three seeds). The endpoint is the per-participant difference in balanced accuracy, with intervals from a participant-cluster bootstrap; family colours match figure 3 .
Figure 2: The continuous residual gate. (a) The raw trace x(t) is divided by a per-record robust scale s . A stem and nine dilated residual blocks (48 channels, dilation d from 1 to 256) feed a detection head, whose output p(t)=σ(ℓ(t)) lies between 0 and 1, and a residual head, whose output r(t)=4tanh(d(t)) lies between −4 and 4. Their product is added to the unmodified normalized trace, clamped at zero and multiplied by s (equations 1 and 2 ), so the output equals the input wherever p(t) is close to zero. Shaded modules are learned on the benchmark and then frozen; the other operations have no parameters. (b) The operator on a 14-s window of benchmark record E01 (CRG-A output, 128 Hz, unfiltered). The correction xgate−x=sp(t)r(t) is concentrated in the expert-annotated artifact span, where the output moves towards the expert reference x∗(t) . The bottom track shows the training targets that the mask defines: an L1 loss to x∗ inside the span and to the raw signal outside it. The detection head is trained against the mask throughout, and CRG-B adds distillation from a fold-local teacher. (c) Receptive field after each block, from 9 samples after the first block to 4089 samples (31.9 s at 128 Hz) after the last. The dashed line marks the length of one VR window (720 samples).
Figure 3: Effect of gating on VR classification. (a) Gated-minus-raw balanced accuracy for each model family. Open circles are the 11 held-out participants, each averaged over seeds; filled circles and bars are the mean and its 95% participant-cluster bootstrap interval. The shaded band is the ±3.32 pp margin tested by TOST, and dotted lines mark the tightest common margin ( ±2.37 pp). BF01 uses a JZS prior ( r=0.707 ) against a two-sided alternative. (b) TOST p as a function of the equivalence half-width δ , computed from the same participant differences (paired t , 10 degrees of freedom). A family is equivalent within ±δ wherever its curve lies below α=0.05 .
Stage
Measurement (checkpoint)
Result
Meaning
Contamination
Artifact share
10.75% annotated in the benchmark; 4.64% estimated in VR
Less to remove in VR
Detection
Record AUROC (CRG-A)
Median 0.9409; 36/43 records above 0.90; median probability 3.13 times median artifact fraction
Good ranking, too much intervention
Reconstruction
Artifact-region MAE (CRG-B)
0.1010 to 0.0830 \SIUnitSymbolMicroS ( −17.8 %)
Waveforms improved
Events
SCR peak F1 (CRG-B)
0.8980 to 0.8674; lower in 43/43 records
Events degraded
Features
MiniRocket displacement (CRG-A)
No change needed in 57.4% of windows, all changed; median cosine 0.363
Change where none was needed, partial alignment where it was
Table 2: From contamination to decision: where the expected benefit was lost. Each row is a separate measurement on the checkpoint shown.
Figure 4: Two measured benchmark excerpts at 128 Hz, unfiltered, selected to illustrate failure modes rather than typical behaviour: gate failure at high artifact burden in E01 (test record) and over-intervention at low burden in E14 (training record). Upper panels show the raw trace (blue) and the CRG-A output (vermillion), which is drawn beneath raw so that it is visible only where the gate moved the signal. The expert-clean reference (green, dotted) is drawn within annotated spans; outside them it equals raw in E14 and differs from raw by less than 0.0025\SIUnitSymbolMicroS in E01. Lower strips show reference minus raw (green) and gated minus raw (vermillion). Circles are reference SCR peaks, and crosses are detections on the raw or gated trace left unmatched by one-to-one matching within ±1 s. Dashed boxes mark the windows enlarged at right.
Figure 5: Detector behaviour on the benchmark (CRG-A; out-of-fold predictions for training records, test-ensemble predictions for test records). (a) Point-wise AUROC for each record; bars mark split medians, and 36 of 43 records exceed 0.90. (b) Observed artifact fraction in 100 fixed probability bins of width 0.01, with the number of samples per bin below on a log scale (34,312,745 samples in total). (c) Mean predicted probability against annotated artifact fraction for each record. All 43 records lie above the identity line.
MAE ( \SIUnitSymbolMicroS )
Red.
MAE
F1
Subset
n
raw → gated
(%)
lower
lower
Expert 1
21
0.134 → 0.097
27.6
21/21
21/21
Expert 2
22
0.069 → 0.070
−0.2
10/22
22/22
Training
33
0.102 → 0.085
16.3
25/33
33/33
Test
10
0.098 → 0.075
23.1
6/10
10/10
Table 3: CRG-B reconstruction and event outcomes by expert subset and split. MAE is averaged across records; reduction =100×(1−MAEgated/MAEraw) . The last two columns count records. Expert and split subsets overlap, so rows are not additive. In the test split alone, F1 fell by 2.50 pp ( p=0.002 ) and MAE did not change significantly ( p=0.084 ; paired Wilcoxon).
Figure 6: Waveform repair and event preservation on the 43 benchmark records, CRG-B continuous-strength student versus raw input. (a) Artifact-region mean absolute error to the expert-clean reference, on a log scale. (b) SCR peak F1 . Thin lines join each record and the thick line joins the means. (c) The same records as paired changes. Waveform error fell in 31 records and event F1 fell in all 43. Filled circles are records annotated by Expert 1 and open circles those annotated by Expert 2.
Figure 7: Decision changes after gating for each model family, with 2,952 prediction instances per family (984 windows × 3 seeds). Bars to the right count instances that became correct, and bars to the left those that became incorrect. Diamonds mark the net change, which equals the pooled accuracy difference in equation 5 . The right-hand columns give that difference and the share of instances whose class did not change.
Figure 8: Controlled artifact injection. (a, b) Change in balanced accuracy relative to contaminated input at each effective perturbed fraction for catch22 + RF and MiniRocket; contaminated input defines the zero line. Points are means over 11 participants with seeds averaged, and arms are offset horizontally for legibility. Vertical bars are 95% participant-cluster bootstrap intervals for the oracle-localization contrast. The clean-input arm reuses the native windows at every dose, so its curve shows the damage that perfect repair would undo; MiniRocket’s values vary across doses only because its transform is refitted at each dose. (c) The prespecified validity check (catch22 + RF, highest dose versus native input) and pooled nonzero-dose contrasts against contaminated input.
Figure 9: Participant composition behind the artifact–accuracy association. (a) Difference in raw-input accuracy between low- and high-proxy windows after a global median split and after a median split within each participant, with 95% participant-cluster bootstrap intervals. (b) Share of each participant’s windows above the global proxy median; without differences between participants every share would lie near 50%. Three participants (P06, P10 and P03) supply 50.4% of all high-proxy windows.
Wrist electrodermal activity (EDA) differs in amplitude from one person to the next, so affect-recognition models normalize their input before classification. Studies that test such models on held-out subjects seldom report where the normalization statistics come from, yet statistics computed from the held-out subject's own recording give the model information that a device does not have when it is first worn. We asked how this choice alters the measured benefit of pretraining. A compact convolutional network, SAFE-EDA, was pretrained on expert artifact annotations from 43 subjects and compared with the same network trained from scratch on the Wearable Stress and Affect Detection (WESAD) dataset (15 subjects, leave-one-subject-out), with two normalization sources crossed with four window hops. When the statistics came only from training subjects, pretraining raised macro-F1 by 0.078 to 0.227; when they came from the held-out user's full recording, the gain fell to between 0.020 and 0.050 and was no longer significant. Artifact supervision was far more useful than self-supervised pretraining on the same recordings (0.078 versus 0.008). Across 13 configurations in two datasets, the pretrained network was better in 12, but on the second dataset (26 subjects) per-user normalization increased the gain instead of reducing it, so the interaction depends on the data. Only five of 50 published WESAD studies state which data were used for normalization. Reporting this choice is necessary to separate first-use performance from performance after calibration.
Haochen Chai, Xinbi Luo, Zining Liu +1
College of Medicine and Biological Information Engineering, Northeastern University, Shenyang, China
Electroencephalography (EEG) is a cornerstone of brain-computer interfaces and clinical neuroscience, yet deep learning models are typically trained and evaluated under a single, unreported preprocessing pipeline. We formalize preprocessing choices as a counterfactual intervention space and show that EEG predictions are surprisingly unstable under this space: across six datasets spanning four paradigms, up to 42% of trial-level predictions flip when only the preprocessing changes, a variability that standard uncertainty methods do not explicitly quantify because they condition on a fixed preprocessing pipeline. We provide three tools to make this instability measurable, decomposable, and reducible. First, a Walsh-Hadamard decomposition of the 2^7 pipeline space reveals that sensitivity is near-additive in practice under the binary intervention design, enabling efficient step-by-step optimization. Second, we introduce Preprocessing Uncertainty (PU), a per-trial diagnostic that captures a dimension of instability complementary to model-based confidence. Third, we study Normalized Adaptive PGI (NA-PGI), a graph-structured regularizer that exploits the compositional structure of preprocessing interventions as one mitigation strategy with clear scope conditions.
Dengzhe Hou, Zihao Wu, Lingyu Jiang +3
Tohoku University · University of Georgia · Texas A&M University +1
Electroencephalography (EEG) is highly susceptible to electromyographic (EMG) artifacts, whose temporal heterogeneity and spatial-spectral overlap with neural activity can leave mixed sources after blind source separation. Existing artifact-removal methods are further limited by scarce reliable component-level ground truth: expert annotations are costly and subjective, while no established method provides realistic simulation-based ground truth for EMG contamination in multichannel scalp EEG. To address these limitations, we propose a framework combining a frequency-aware high-dimensional representation with Multi-Instance Learning. The representation unfolds separated components into frequency-resolved intra-components, creating a space in which mixed neural and muscular activity becomes more separable, while the weakly supervised learning formulation enables artifact-likelihood scores for individual intra-components to be learned from epoch-level labels without finer-grained ground truth. The resulting intra-component classifier supports fine-grained EMG artifact detection and score-guided attenuation. Experiments on held-out subjects show that the framework learns informative intra-component scores and reduces artifact-related spectral deviations most clearly for jaw tension, with moderate effects for raising eyebrows and limited effects for frowning.
Lu Wang-Nöth, Hai Huang, Philipp Heiler +3
Institute for Applied Computer Science, University of the Bundeswehr Munich, Munich, Germany · brainboost GmbH, Munich, Germany · Graduate School of Engineering Science, The University of Osaka, Osaka, Japan +1