Learning Domain- and Class-Disentangled Prototypes for Domain-Generalized EEG Emotion Recognition
Authors: Guangli Li, Canbiao Wu, Zhehao Zhou, Na Tian, Li Zhang, Zhen Liang
Organizations: School of Biological Science and Medical Engineering, Hunan University of Technology, Zhuzhou 412008, China · School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen 518060, China · Guangdong Provincial Key Laboratory of Biomedical Measurements and Ultrasound Imaging, Shenzhen 518060, China · Shenzhen Pengrui Brain Science Technology, Shenzhen 518060, China
Electroencephalography (EEG)-based emotion recognition plays a critical role in affective Brain-Computer Interfaces (aBCIs), yet its practical deployment remains limited by inter-subject variability, reliance on target-domain data, and unavoidable label noise. To address these challenges, we propose a Multi-domain Aggregation Transfer learning framework with domain-class prototypes (MAT) for emotion recognition under completely unseen target domains. MAT introduces a feature decoupling module to disentangle class-invariant domain features from domain-invariant class features, enabling more robust and interpretable EEG representations. A Hierarchical-Domain Aggregation (HDA) mechanism based on Maximum Mean Discrepancy (MMD) constructs superdomains to model shared distributional structures across subjects, while adaptive prototype updating refines domain and class prototypes to capture stable intrinsic representations. Moreover, a pairwise learning strategy reformulates classification as similarity estimation between sample pairs, effectively mitigating the effect of label noise. Extensive experiments on three public EEG emotion datasets (SEED, SEED-IV, and SEED-V) show that the accuracy of MAT is improved by 2.87%, 3.84%, and 2.05% compared with the state-of-the-art (SOTA) model for unseen target domains. Our results provide a promising direction for emotion recognition under real-world unseen-subject scenarios.The source code is available at https://github.com/WuCB-BCI/MAT.
Figures & tables
Fig. 1: Illustration of the transfer learning paradigm. Knowledge extracted from one or multiple source domains is transferred to a target domain to enhance feature alignment and model generalization across subjects.
Fig. 2: The T-SNE visualization of the previous prototype representation method and ours. (a) Single class-prototype representation methods. (b) Our proposed dual-prototype representation method, which utilizes both domain-specific and class-specific prototype representations, effectively improves the generalization performance of the model.
Fig. 3: The training phase of the MAT framework. Here, EEG features are first decoupled into domain- and class-specific components through discriminators Dd and Dc with a Gradient Reversal Layer (GRL) for adversarial alignment. Second, domains S1∼Sn are clustered into superdomains S1∼SK by HDA mechanism to capture shared yet distinct subject representations. Within each superdomain, domain prototypes μd and class prototypes μc are adaptively updated to ensure stable and discriminative feature learning. Finally, the traditional classification task is transformed into the relative consistency learning between samples, improving robustness against label noise and unseen target domains.
Fig. 4: The inferencing phase of the MAT framework. The optimal domain prototype μd and class prototype μc obtained during the model training phase are transferred to the unseen target domain. Specifically, we first decouple the sample features, and then the model determines the superdomain space μd by Domain Prototype Inference, and perform Class Prototype Inference within this superdomain space to determine the emotion category μc . Here, h(⋅) represent the bilinear transformation to capture the most relevant domain space (Eq. 15 ). dcos(⋅) represents the similarity evaluation (Eq. 16 ).
Notation
Description
S/T
Source/Target Domain
xd/xc
Domain-/Class-specific Representation
yd/yc
Domain/Class Label
fg(⋅)
Shallow Feature Extractor
fd(⋅)/fc(⋅)
Domain/Class Feature Decoupler
Dd(⋅)/Dc(⋅)
Domain / Class Discriminator
TABLE I: Frequently used notations and descriptions.
Methods
Pacc(%)
Methods
Pacc(%)
Traditional machine learning methods ( ↑ 13.22%)
KNN* [ 29 ]
55.26 ± 12.43
KPCA* [ 30 ]
48.07 ± 09.97
SVM* [ 31 ]
70.62 ± 09.02
SA* [ 32 ]
59.73 ± 05.40
TCA* [ 33 ]
58.12 ± 09.52
CORAL* [ 34 ]
71.48 ± 11.57
GFK* [ 35 ]
56.71 ± 12.29
RF* [ 36 ]
62.78 ± 06.60
Deep transfer learning methods ( ↑ 2.78%)
TABLE II: Cross-subject single-session LOSO cross-validation results on SEED dataset, expressed as (Acc% ± Std%). Here, ’*’ indicates the results are obtained by our own implementation. ’#’ denoted as the model of the unseen target domain and ↑ represents the gap between MAT and them.
Methods
Pacc(%)
Methods
Pacc(%)
Traditional machine learning methods ( ↑ 15.44%)
KNN* [ 29 ]
41.77 ± 09.53
KPCA* [ 30 ]
29.25 ± 09.73
SVM* [ 31 ]
50.50 ± 12.03
SA* [ 32 ]
34.74 ± 05.29
TCA* [ 33 ]
44.11 ± 10.76
CORAL* [ 34 ]
48.14 ± 10.38
GFK* [ 35 ]
43.10 ± 09.77
RF* [ 36 ]
52.67 ± 13.85
Deep transfer learning methods ( ↑ 3.84%)
TABLE III: Cross-subject single-session LOSO cross-validation results on SEED-IV dataset, expressed as (Acc% ± Std%). Here, ’*’ indicates the results are obtained by our own implementation. ’#’ denoted as the model of the unseen target domain and ↑ represents the gap between MAT and them.
Methods
Pacc(%)
Methods
Pacc(%)
Traditional machine learning methods ( ↑ 8.30%)
KNN* [ 29 ]
35.73 ± 07.98
KPCA* [ 30 ]
35.47 ± 09.39
SVM* [ 31 ]
53.14 ± 10.10
SA* [ 32 ]
36.06 ± 11.55
TCA* [ 33 ]
37.57 ± 13.47
CORAL* [ 34 ]
55.18 ± 07.42
GFK* [ 35 ]
38.32 ± 10.11
RF* [ 36 ]
42.29 ± 16.02
Deep transfer learning methods ( ↑ 4.45%)
TABLE IV: Cross-subject single-session LOSO cross-validation results on SEED-V dataset, expressed as (Acc% ± Std%). Here, ’*’ indicates the results are obtained by our own implementation. ’#’ denoted as the model of the unseen target domain and ↑ represents the gap between MAT and them.
Methods
Pacc(%)
Methods
Pacc(%)
Traditional machine learning methods ( ↑ 9.56%)
KNN* [ 29 ]
21.72 ± 06.87
KPCA* [ 30 ]
35.93 ± 09.81
SVM* [ 31 ]
31.21 ± 09.37
SA* [ 32 ]
30.45 ± 11.31
TCA* [ 33 ]
37.34 ± 08.82
CORAL* [ 34 ]
27.96 ± 07.22
GFK* [ 35 ]
31.34 ± 09.66
RF* [ 36 ]
34.82 ± 08.17
Deep transfer learning methods ( ↑ 2.21%)
TABLE V: Cross-subject single-session LOSO cross-validation results on SEED-VII dataset, expressed as (Acc% ± Std%). Here, ’*’ indicates the results are obtained by our own implementation. ’#’ denoted as the model of the unseen target domain and ↑ represents the gap between MAT and them.
Methods
Pacc(%)
Methods
Pacc(%)
Traditional machine learning methods ( ↑ 7.43%)
KNN* [ 29 ]
60.18 ± 08.10
KPCA* [ 30 ]
72.56 ± 06.41
SVM* [ 31 ]
68.01 ± 07.88
SA* [ 32 ]
57.47 ± 10.01
TCA* [ 33 ]
63.63 ± 06.40
CORAL* [ 34 ]
55.18 ± 07.42
GFK* [ 35 ]
60.75 ± 08.32
RF* [ 36 ]
72.78 ± 06.60
Deep transfer learning methods ( ↑ 2.48%)
TABLE VI: Cross-subject cross-session LOSO cross-validation results on SEED dataset, expressed as (Acc% ± Std%). Here, ’*’ indicates the results are obtained by our own implementation. ’#’ denoted as the model of the unseen target domain and ↑ represents the gap between MAT and them.
Methods
Pacc(%)
Methods
Pacc(%)
Traditional machine learning methods ( ↑ 14.5%)
KNN* [ 29 ]
40.06 ± 04.98
KPCA* [ 30 ]
47.79 ± 07.85
SVM* [ 31 ]
48.36 ± 07.51
SA* [ 32 ]
40.34 ± 05.85
TCA* [ 33 ]
43.01 ± 07.13
CORAL* [ 34 ]
50.01 ± 07.93
GFK* [ 35 ]
43.48 ± 06.27
RF* [ 36 ]
48.16 ± 09.43
Deep transfer learning methods ( ↑ 3.23%)
TABLE VII: Cross-subject cross-session LOSO cross-validation results on SEED-IV dataset, expressed as (Acc% ± Std%). Here, ’*’ indicates the results are obtained by our own implementation. ’#’ denoted as the model of the unseen target domain and ↑ represents the gap between MAT and them.
Methods
Pacc(%)
Methods
Pacc(%)
Traditional machine learning methods ( ↑ 4.31%)
KNN* [ 29 ]
35.28 ± 07.57
KPCA* [ 30 ]
39.68 ± 11.28
SVM* [ 31 ]
41.20 ± 10.76
SA* [ 32 ]
31.87 ± 09.87
TCA* [ 33 ]
37.68 ± 08.40
CORAL* [ 34 ]
54.08 ± 07.44
GFK* [ 35 ]
37.89 ± 09.84
RF* [ 36 ]
43.63 ± 11.38
Deep transfer learning methods ( ↑ 3.23%)
TABLE VIII: Cross-subject cross-session LOSO cross-validation results on SEED-V dataset, expressed as (Acc% ± Std%). Here, ’*’ indicates the results are obtained by our own implementation. ’#’ denoted as the model of the unseen target domain and ↑ represents the gap between MAT and them.
Methods
Pacc(%)
Methods
Pacc(%)
Traditional machine learning methods ( ↑ 5.50%)
KNN* [ 29 ]
19.03 ± 03.09
KPCA* [ 30 ]
28.99 ± 05.84
SVM* [ 31 ]
22.50 ± 04.70
SA* [ 32 ]
17.98 ± 03.89
TCA* [ 33 ]
28.58 ± 07.74
CORAL* [ 34 ]
20.44 ± 04.64
GFK* [ 35 ]
27.12 ± 06.25
RF* [ 36 ]
27.13 ± 04.44
Deep transfer learning methods ( ↑ 2.82%)
TABLE IX: Cross-subject cross-session LOSO cross-validation results on SEED-VII dataset, expressed as (Acc% ± Std%). Here, ’*’ indicates the results are obtained by our own implementation. ’#’ denoted as the model of the unseen target domain and ↑ represents the gap between MAT and them.
Ablation Strategy
Pacc(%)
w/o Domain Prototype
78.95 ± 08.92
↓ 5.75%
w/o Class Disc. Loss in Eq. 3
81.62 ± 07.16
↓ 3.08%
w/o Domain Disc. Loss in Eq. 4
79.34 ± 08.74
↓ 5.36%
w/o Class and Domain Disc. Loss
78.11 ± 09.24
↓ 6.59%
w/o HDA mechanism in Sec. III-B
80.23 ± 05.12
↓ 4.47%
w/o Ada. Upd. Coe. α in Eq. 12
81.54 ± 05.86
↓ 3.16%
TABLE X: Results of ablation experiments of the MAT model, expressed as (Acc% ± Std%)
Fig. 5: T-SNE visualization of domain representation and domain alignment without (a) to (c) and with (d) to (f) Hierarchical-Domain Aggregation mechanism. And showing the distribution of domain representations at the beginning of training, after 50 training epochs, and end of training, respectively.
Fig. 6: T-SNE visualizations of class representation and class alignment of the MAT model, showing the distribution of class features at the beginning of training, after 50 training epoch, and end of training, respectively.
Fig. 7: Results for different parameter settings of the proposed MAT model.
Noisy
Pointwise
Pairwise
Pair. − Point.
Ratio (η)
Learning (%)
Learning (%)
(%)
0%
76.73 ± 06.62
84.70 ± 04.63
↑ 07.97
5%
75.89 ± 06.75
83.61 ± 05.28
↑ 07.72
10%
73.45 ± 05.94
83.02 ± 05.06
↑ 09.57
20%
71.84 ± 07.96
82.31 ± 04.94
↑ 10.47
30%
70.04 ± 07.63
81.38 ± 05.68
↑ 11.34
TABLE XI: Results of the MAT model adding different proportions ( η% ) of random label noise to the source domain, expressed as (Mean-Accuracy% ± Standard-Deviation%). ↑ denotes the performance differences.
Fig. 8: Performance of the MAT model in different number of aggregates K .
Methods
Pacc(%)
Methods
Pacc(%)
Traditional machine learning methods ( ↑ 9.60%)
KNN* [ 29 ]
70.22 ± 5.95
KPCA* [ 30 ]
70.56 ± 3.27
SVM* [ 31 ]
68.31 ± 7.98
SA* [ 32 ]
69.31 ± 6.36
TCA* [ 33 ]
67.25 ± 5.38
CORAL* [ 34 ]
67.47 ± 5.70
GFK* [ 35 ]
66.31 ± 5.39
RF* [ 36 ]
67.88 ± 4.67
Deep transfer learning methods ( ↑ 4.99%)
TABLE XII: Cross-subject LOSO cross-validation results on MDD dataset, expressed as (Acc% ± Std%). Here, ’*’ indicates the results are obtained by our own implementation. ’#’ denoted as the model of the unseen target domain and ↑ represents the gap between MAT and them.
Fig. 9: Figure S1: The T-SNE visualization of the without HDA and with HDA.
Fig. 10: Figure S2: Schematic illustration of the interaction and iterative optimization among domain-class disentanglement, prototype learning, and hierarchical domain aggregation. Each sample is decomposed into domain-specific and class-specific representations, which are in one-to-one correspondence. The illustrated samples originate from two different subjects and therefore exhibit distinct subject-specific characteristics (light green and dark green), while samples belonging to the same emotion share similar emotion-specific patterns. Accordingly, the three emotion categories from different subjects are represented by red, yellow, and purple, respectively.
Fig. 11: Figure S3: The loss curve of proposed MAT: (a) Domain feature disentanglement loss Eq. 4 . (b) Class feature disentanglement loss Eq. 3 . (c) Total disentanglement loss Eq. 5 . (d) Pairwise computing loss Eq. 17 . (e) The MAT total loss Eq. 19 .
Fig. 12: Figure S4: Confusion matrices of different baseline models under cross-subject single-session LOSO cross-validation. The SEED database (a) ∼ (d) contains three emotion categories: negative, neutral and positive. The Seed-IV database (e) ∼ (h) contains four emotion categories: neutral, sad, fear and happy. The horizontal axis represents the predicted labels, while the vertical axis represents the true labels.
Fig. 13: Figure S5: Confusion matrices of different baseline models under cross-subject single-session LOSO cross-validation. The Seed-V database contains five emotion categories: happy, fear, neutral, sad, and disgust. The horizontal axis represents the predicted labels, while the vertical axis represents the true labels.
Fig. 14: Figure S6: Obtained p-values of the paired t-test between the proposed MAT and the baseline method in terms of accuracy. The error bars represent the standard deviation of the obtained experimental results. Here, ML and DL represent the baseline comparison results of the machine learning model and the deep transfer learning model, respectively. N/A indicates that the p-value could not be obtained because the results of the baseline model were not implemented by us.
Electroencephalogram (EEG) captures endogenous brain activity with high temporal fidelity and holds substantial promise for precise emotion decoding. However, channel redundancy and pronounced inter-subject variability remain key obstacles to scalable generalization. To address these limitations, we propose a novel framework termed PRioritized channel Importance with Semi-supervised doMain adaptation (PRISM), enabling label-efficient cross-subject emotion decoding. On the channel side, PRISM assigns differentiable, data-dependent channel weights via a lightweight expert ensemble, amplifying reliable electrodes while suppressing distractors. On the domain side, PRISM leverages unlabeled data through confidence-filtered pseudo-labels to drive consistency regularization and domain alignment, mitigating subject-specific heterogeneity. Extensive experiments show that PRISM surpasses state-of-the-art methods on DEAP, DREAMER, and SEED datasets, achieving robust cross-subject generalization given limited annotations.
Xin Zhou, Xiang Zhang, Hao Deng +1
School of Computing, T. J Watson College of Engineering and Applied Science, Binghamton University - State University of New York, Binghamton, NY 13902 USA · Massachusetts General Hospital, Harvard University, Boston, MA 02114 USA
Cross-subject electroencephalogram (EEG)-based emotion recognition remains challenging due to substantial inter-individual variability and discrete formulation that overlooks affective continuity. Existing methods operate in Euclidean space and focus on marginal distribution alignment, failing to preserve the semantic structure of emotions across subjects. This article proposes MGMCL, reconceptualizing emotion recognition as learning continuous representations on symmetric positive definite (SPD) Riemannian manifolds. The frame?work introduces multi-granularity manifold contrastive learning at instance, emotion, and trajectory levels while preserving semantic ordering. Neural ordinary differential equations on manifolds model continuous emotion dynamics. Cross-subject generalization employs Gromov-Wasserstein manifold alignment. Weakly-supervised learning enables continuous valence-arousal-dominance prediction from discrete labels. Extensive experiments on three public datasets demonstrate state-of-the-art performance: 91.23% accuracy on SEED, 73.82% on SEED-IV, and 76.38% on DEAP, achieving consistent improvements of 1.89%, 1.66%, and 1.28% over previous best methods, respectively.
In cross-subject EEG emotion decoding, neighboring windows can inform the representation of a target window, but they contain both background variation and sustained task signals. This creates a challenge for contextual representation learning because subtracting shared activity can also remove useful information. We introduce the Morlet Spectral Transformer (MST), which conditions target-window tokens on a structured spectral summary before attention. MST averages the Morlet log-amplitude spectra of unlabeled same-trial neighbors and uses frequency-specific electrode projections to combine the target spectrum with its reference-relative residual. This retains access to absolute activity while incorporating context into a fixed attended target-token grid. Without external pretraining, MST achieves 66.5%, 40.7%, 36.1%, and 27.9% accuracy on SEED, SEED-IV, SEED-V, and SEED-VII under leave-one-subject-out evaluation, averaged over three random seeds, outperforming the evaluated pretrained and from-scratch baselines. On FACED, MST achieves 17.52% accuracy in nine-class cross-subject evaluation, exceeding the strongest evaluated baseline by 1.15 percentage points. We also performed the ablation studies across all 15 SEED subjects to evaluate the contribution of different components of MST.