A Proof-of-Concept Study of Weakly Supervised Labeling of Fine-Grained EEG Components for Artifact Attenuation
Authors: Lu Wang-Nöth, Hai Huang, Philipp Heiler, Shuqiong Wu, Liyun Zhang, Helmut Mayer
Organizations: Institute for Applied Computer Science, University of the Bundeswehr Munich, Munich, Germany · brainboost GmbH, Munich, Germany · Graduate School of Engineering Science, The University of Osaka, Osaka, Japan · Graduate School of Frontier Sciences, The University of Tokyo, Tokyo, Japan
Electroencephalography (EEG) is highly susceptible to electromyographic (EMG) artifacts, whose temporal heterogeneity and spatial-spectral overlap with neural activity can leave mixed sources after blind source separation. Existing artifact-removal methods are further limited by scarce reliable component-level ground truth: expert annotations are costly and subjective, while no established method provides realistic simulation-based ground truth for EMG contamination in multichannel scalp EEG. To address these limitations, we propose a framework combining a frequency-aware high-dimensional representation with Multi-Instance Learning. The representation unfolds separated components into frequency-resolved intra-components, creating a space in which mixed neural and muscular activity becomes more separable, while the weakly supervised learning formulation enables artifact-likelihood scores for individual intra-components to be learned from epoch-level labels without finer-grained ground truth. The resulting intra-component classifier supports fine-grained EMG artifact detection and score-guided attenuation. Experiments on held-out subjects show that the framework learns informative intra-component scores and reduces artifact-related spectral deviations most clearly for jaw tension, with moderate effects for raising eyebrows and limited effects for frowning.
Figures & tables
Figure 1 : Unfolding components to intra-components with components 0, 1, 12, 13, 22 and 23 as examples. For visual clarity, frequency bins are aggregated into standard EEG frequency bands: theta (6.875–8 Hz; corresponding to the available portion of the conventional 4–8 Hz theta band), low-alpha (8–10 Hz), high-alpha (10–12 Hz), low-beta (12–20 Hz), high-beta (20–30 Hz), low-gamma1 (30–50 Hz), and low-gamma2 (50–80 Hz). Following [ 20 ] , the ILRMA decomposition starts at 6.875 Hz to avoid linear dependencies across frequency; each topomap uses an individual color scale to show its spatial distribution; the numbers beside each topomap are the limits of its color range for unitless, rescaled projected power; hence, colors are not quantitatively comparable across maps. The intra-component representation enables frequency-wise separation of mixed neural and non-neural activity within components.
Figure 2 : Weakly supervised MIL-based intra-component classifier. During training, each EEG epoch is treated as a bag of intra-components. A shared encoder maps each intra-component to an embedding, and a sigmoid head assigns an instance-level artifact-likelihood score. The highest-scoring 2% of intra-components are averaged by top- k MIL pooling to obtain the epoch-level contamination score, which is optimized against the weak epoch label. After training, the MIL pooling layer is discarded, and the encoder and score head are reused for intra-component classification. The framework permits threshold selection on validation data, while the proof-of-concept cleaning evaluation used a fixed threshold of τ=0.5 to define a binary attenuation mask.
Metrics
Test @ τ=0.5
BCE
0.38
Balanced accuracy
0.91
Recall
0.89
Specificity
0.93
Precision
0.86
F1 score
0.88
Table 1 : Epoch-level test classification results for the MIL model at threshold 0.5
Metrics
Test @ τ∈[0.02,0.95]
ROC AUC
0.98
Average Precision
0.96
Table 2 : Epoch-level test classification results for the MIL model across different thresholds
Artifact
Fp1
Fp2
F9
F7
F3
Fz
F4
F8
F10
T9
T7
C3
JT
5.09
7.38
84.77
111.53
86.69
25.21
120.05
144.94
141.23
43.53
18.54
95.97
RE
-0.47
17.97
0.30
2.27
9.86
3.03
10.03
3.33
0.66
0.31
0.52
1.89
Fr
-0.71
0.24
0.02
0.02
0.06
0.02
0.02
0.04
0.03
0.00
0.00
0.02
Table 3 : Channel-wise median relative area-difference improvement after removing artifact-like intra-components, grouped by artifact type: jaw tension (JT), raising eyebrows (RE), and frowning (Fr), with N=24 epochs per type. Values are reported in the units of the relative-deviation metric Δdc and should not be interpreted as percentage improvement. Positive values indicate reduced spectral deviation from the same-subject EO reference, whereas negative values indicate increased deviation after cleaning. Green highlights mark pronounced positive improvements, and red marks undesirable changes. Jaw tension showed the strongest improvements over frontal, fronto-temporal, and central channels; raising eyebrows showed smaller mainly frontal improvements; frowning produced near-zero changes across channels.
Figure 3 : Representative jaw-tension artifact example showing normalized time-averaged power spectra before and after removal of artifact-like intra-components. Each panel corresponds to one EEG channel (refer to the topomap upper right) and shows the artifact-contaminated epoch before cleaning, the reconstructed epoch after cleaning, and the EO reference spectrum. Importantly, complete overlap between the after-cleaning curve and the clean reference curve is not expected. The y-axis is scaled independently for each channel to emphasize channel-specific spectral changes. Red rectangles highlight frontal, fronto-temporal, and central channels in which jaw tension produced a pronounced broadband increase in power, particularly at higher frequencies. These channels also show the largest y-axis ranges. After intra-component removal, the power spectra were strongly attenuated across these channels and shifted substantially closer to the EO reference, with smaller reductions also visible over other channels.
Figure 4 : Representative raising-eyebrows artifact example showing normalized time-averaged power spectra before and after removal of artifact-like intra-components. Each panel corresponds to one EEG channel (refer to the topomap upper right) and shows the artifact-contaminated epoch before cleaning, the reconstructed epoch after cleaning, and the EO reference spectrum. Importantly, complete overlap between the after-cleaning curve and the clean reference curve is not necessarily expected. The y-axis is scaled independently for each channel to emphasize channel-specific spectral changes. The red rectangle highlights frontal channels in which the raising-eyebrows artifact produced elevated broadband power. These channels also show the largest y-axis ranges. After intra-component removal, the spectra were attenuated across most channels, with the strongest visible reduction in the highlighted frontal channels and smaller reductions in non-frontal channels.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Feature group
Formula
Short explanation
Dim.
Mean magnitude activity
T1∑t=1T∣x(t)∣
Average absolute normalized log-magnitude.
1
Burst fraction
T1∑t=1T1[x(t)>τ]
Fraction of samples above the burst threshold.
1
Longest burst duration
T1maxj(zj−uj+1)
Longest contiguous above-threshold segment, normalized by signal length.
1
High-energy concentration
∑t=1Texp(x(t))+ϵ∑t∈Top10%exp(x(t))
Fraction of magnitude energy contained in the largest 10% of samples.
1
Temporal variation
T−11∑t=1T−1∣x~(t+1)−x~(t)∣
Mean absolute first difference of the smoothed magnitude signal.
Autocorrelation of x(t) at four relative lags. xˉ is the mean of x(t) and var(x) the variance.
4
Appendix
Table 0.A.1 : Handcrafted magnitude-based temporal features used for each intra-component. The column “Dim.” gives the number of scalar features contributed by each feature group.
Feature group
Formula
Short explanation
Dim.
Circular concentration
T1∑t=1Texp(jθ(t))
Circular resultant length of the phase sequence.
1
Phase-increment concentration
T−11∑t=1T−1exp(jΔθ(t))
Circular resultant length of phase increments.
1
Phase increment variability
stdt(Δθ(t))
Standard deviation of wrapped phase increments.
1
Mean absolute phase increment
T−11∑t=1T−1∣Δθ(t)∣
Average absolute phase change between adjacent samples.
Table 0.A.2 : Handcrafted phase-based temporal features used for each intra-component. The column “Dim.” gives the number of scalar features contributed by each feature group.
Module
Layer
Kernel/Stride /Padding
Activation /Dropout
Output
Input
Spatial topomaps T
–
–
N×3×32×32
Spatial
Conv2D 3→16
3×3/1/1
GELU
N×16×32×32
Spatial
Conv2D 16→32
3×3/2/1
GELU
N×32×16×16
Spatial
Conv2D 32→64
3×3/2/1
GELU
N×64×8×8
Spatial
Adaptive avg. pool + flatten
1×1
Dropout 0.1
N×64
Spatial
Linear 64→64
–
GELU
N×64
Appendix
Table 0.B.1 : Detailed architecture of the intra-component classifier. All convolutional and linear layers use learned biases; no batch normalization is used, with LayerNorm applied only to the 66-D temporal feature input. Here, F denotes the number of frequency bins, S denotes the number of components (ILRMA sources), and N=F×S is the number of intra-components in one MIL bag, corresponding to one EEG epoch.
Figure 0.B.1 : Detailed architecture of the intra-component classifier. Each intra-component has three inputs: a 3-channel spatial topomap Ti∈R3×32×32 , a 66-dimensional handcrafted temporal feature vector vi∈R66 , and a normalized scalar frequency coordinate fi∈R . Kernel size, stride, and padding are denoted by k , s , and p , respectively. The spatial CNN, temporal MLP, and fusion modules are shared per-instance encoders applied to all N=F×S intra-components, where F is the number of frequency bins and S is the number of ILRMA sources, as indicated by the green block. The frequency-context encoder then gathers all embeddings in the bag (one EEG epoch), reshapes them to [S,64,F] , and applies 1-D convolutions along the frequency axis for each component (ILRMA source). Each final score pi is therefore an intra-component artifact-likelihood score; one EEG epoch produces N such scores.
Artifact class
Channel
p5
p25
Median
p75
p95
JT
Fp1
-6.13
0.04
5.09
13.39
68.34
JT
Fp2
0.31
1.70
7.38
16.72
67.29
JT
F9
0.91
8.81
84.77
329.51
717.42
JT
F7
1.54
8.93
111.53
291.32
781.37
JT
F3
0.98
13.96
86.69
222.52
648.78
JT
Fz
1.79
9.72
25.21
91.12
156.08
Appendix
Table 0.C.1 : Channel-wise relative area-difference improvement after score-guided removal of artifact-like intra-components for jaw tension (JT).
Artifact class
Channel
p5
p25
Median
p75
p95
RE
Fp1
-35.51
-6.57
-0.47
104.93
1326.41
RE
Fp2
0.42
3.35
17.97
145.74
656.11
RE
F9
-0.07
0.08
0.30
3.01
17.68
RE
F7
0.03
0.24
2.27
14.30
216.86
RE
F3
0.01
2.43
9.86
98.71
497.83
RE
Fz
0.05
1.02
3.03
43.88
91.38
Appendix
Table 0.C.2 : Channel-wise relative area-difference improvement after score-guided removal of artifact-like intra-components for raising eyebrows (RE).
Artifact class
Channel
p5
p25
Median
p75
p95
Fr
Fp1
-8.68
-1.75
-0.71
-0.13
11.92
Fr
Fp2
0.00
0.03
0.24
0.74
4.29
Fr
F9
-0.03
-0.00
0.02
0.06
0.32
Fr
F7
-0.00
0.00
0.02
0.06
0.40
Fr
F3
-0.01
0.00
0.06
0.14
0.72
Fr
Fz
-0.01
0.00
0.02
0.04
2.06
Appendix
Table 0.C.3 : Channel-wise relative area-difference improvement after score-guided removal of artifact-like intra-components for frowning (Fr).
Department of Electronic and Information Engineering, Tokyo University of Agriculture and Technology, Koganei-Shi 184–8588, Japan · Department of Electronics and Communication Engineering, National University of Mongolia, Ulaanbaatar, 14200 Mongolia