A multi-scenario EEG dataset for auditory attention decoding in naturalistic multi-talker environments
Authors: Shu Peng, Rui Liu, Yufei Zhang, Wenlong You, Zhige Chen, Jiachen Xi, Qiyuan Sun, Yan Liu, +2 more
Organizations: Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hong Kong SAR, China · Department of Computing, The Hong Kong Polytechnic University, Hong Kong SAR, China
Understanding how the brain selectively follows relevant speech amid competing voices is a central challenge in auditory neuroscience and a key step toward neuro-steered hearing technologies. However, most open-source Electroencephalography (EEG) datasets for Auditory Attention Decoding (AAD) use idealized single-competing-talker paradigms that oversimplify the acoustic, spatial, and semantic structure of everyday communication. To capture this ecological complexity, we introduce the SoundBubble-EEG dataset: a high-density 128-channel EEG resource comprising more than 25 hours of recordings from 30 participants. The paradigm requires listeners to selectively attend to a dynamic target speaker group, a designated "sound bubble", amid competing multi-speaker distractor bubbles across three realistic scenarios: a restaurant, a home TV viewing, and a meeting discussion. By bridging the gap between constrained laboratory protocols and real-world auditory scenes, this dataset enables investigations of multi-talker speech comprehension, neural speech tracking, and cross-scenario generalization. It also provides a benchmark for AAD algorithms under realistic acoustic and semantic variability and may support auditory neuroscience and the development of neuro-steered hearing technologies.
Figures & tables
Figure 1 : Overview of the SoundBubble-EEG dataset generation, validation, and public release workflow. Participants are engaged in three ecologically valid listening scenarios: meeting, restaurant, and living room (orange). The acquired EEG signals were subsequently processed using a standardized preprocessing pipeline (blue) and underwent rigorous validation, including behavioral assessments, signal quality evaluations, and neural/application-level analyses (green). Finally, the dataset was publicly released in a BIDS-compliant format, accompanied by comprehensive analysis code to facilitate reproducibility and further research (purple).
Figure 2 : Illustration of the "Sound Bubble" paradigm and the experimental procedure. (a) The traditional competing-talker paradigm, which typically features two isolated single speakers (e.g., one on the left versus one on the right) without any environmental background sound. The icon indicates the target sound source that the participant is instructed to focus on. (b) The proposed "Sound Bubble" paradigm replaces individual talker scenarios with clusters of multiple speakers, referred to as "Sound Bubbles", each engaged in discussions centered on a coherent theme and positioned within a realistic acoustic environment. This design requires the listener to direct attention to a target bubble while suppressing a multi-speaker distractor bubble and ambient noise, significantly enhancing ecological validity. (c)–(e) Three ecologically valid listening scenarios were designed for the dataset, in which the green and yellow bubbles indicate the target (attended) and non-target (unattended) sound bubbles, respectively: (c) Restaurant scene characterized by high ambient noise and two competing conversational sound bubbles; (d) Home TV viewing scene, where a fixed media source (TV) competes with side family conversations in a low-noise background; and (e) Meeting scene, where multiple participants form two discussion groups and each speak from a different direction. (f) Experimental procedure. The formal session followed the actual scenario order of Meeting (5 trials), Restaurant (10 trials), and TV (10 trials). In each trial, a black screen with a green directional cue indicated the target sound source and remained visible throughout listening; the first 3 seconds served as baseline and the following 120 seconds as the active-listening task period. Participants then answered a general-comprehension question (Q1) and a specific-details question (Q2), followed by a self-paced rest before the next trial.
Subject ID
Age
Sex
Handedness
Q1 Mean Acc. (%)
Q2 Mean Acc. (%)
sub-01
26
M
R
100.0
100.0
sub-02
25
M
R
96.0
76.0
sub-03
26
M
R
100.0
84.0
sub-04
19
F
R
100.0
80.0
sub-05
26
M
L
96.0
60.0
sub-06
24
M
R
96.0
80.0
Table 1 : Participant demographics and behavioral accuracy. This table details the age, sex, handedness, and formal-trial behavioral accuracy for each participant (‘M’ = Male, ‘F’ = Female; ‘R’ = Right, ‘L’ = Left). Q1 assesses general comprehension of the attended stream, while Q2 focuses on specific details.
Stimulus ID
Scenario
Listener Facing
Sound Bubble 1 (Pos)
Sound Bubble 2 (Pos)
Noise Source (Pos)
S06015
Meeting
0∘
−85∘
+27∘
+57∘
S06746
Meeting
0∘
−53∘
+79∘
−57∘
S06943
Meeting
0∘
+77∘
−36∘
−61∘
S08364
Meeting
0∘
+72∘
−67∘
−50∘
S08440
Meeting
0∘
+61∘
−56∘
−17∘
S00358
Restaurant
0∘
−36∘
+77∘
+44∘
Table 2 : Spatial configurations of the 25 auditory stimulus templates. Azimuth angles are expressed relative to the listener-facing direction ( 0∘ = front; negative values = left; positive values = right). Elevation is 0∘ (eye level) for all sources.
Figure 4 : Behavioral performance and data quality assessment. (a) Behavioral accuracy for general understanding (Q1) and specific details (Q2) across the three acoustic conditions (Restaurant, TV, Meeting). The dashed line indicates chance level (0.25). (b) Sensor layout of the 128-channel high-density EEG net. (c) Distribution of interpolated bad channels per run across three acoustic scenarios. (d) Number of artifact independent components excluded via ICA per run across three acoustic scenarios. (e) Spatial probability heatmap of bad channels across all recordings.
Figure 5 : Neurophysiological sanity checks and application-level AAD validation. (a) Grand average global Power Spectral Density (PSD) for Baseline (dashed grey) and Task (solid red) states, with the alpha band (8–13 Hz) highlighted in yellow. (b) Topography of normalized alpha power change ( 10log10(Task/Baseline) ). Blue regions indicate relative alpha suppression, while red regions denote relative alpha enhancement. (c) Auditory attention decoding performance using a backward mTRF model. Paired raincloud plots show Pearson correlations ( r ) between EEG-reconstructed and actual speech envelopes, demonstrating significantly higher cortical tracking for attended speech (one-sided Wilcoxon signed-rank test, *** p=9.31×10−10 ). (d,e) Trial-label swap null controls for the group-mean decoding advantage Δr=rattend−runattend and single-trial decoding accuracy. Shaded distributions show label-swapped null distributions; vertical colored lines show the observed values computed from the true attended/unattended labels. (f,g) Scenario-wise AAD robustness for Δr and decoding accuracy across Restaurant, TV, and Meeting scenarios. Scenario-wise panels show subject-level scenario means; Restaurant and TV contained 10 trials per subject, whereas Meeting contained 5 trials per subject.
Auditory attention decoding (AAD) aims to infer the attended speaker from neural responses in multi-speaker acoustic environments and is a key problem for neuro-steered hearing systems. Although recent studies have achieved encouraging progress, existing AAD models still do not fully exploit frequency domain electroencephalography (EEG) information. In particular, most approaches introduce multi-band information through handcrafted feature extraction or direct cross-band feature concatenation, which mainly exploit frequency information at a shallow level and may overlook band-specific patterns and cross-band interactions. To address these limitations, this paper proposes FAConformer, a frequency-aware CNN-Transformer framework for AAD that explicitly integrates band-specific encoding and adaptive cross-band interaction. Specifically, FAConformer first decomposes EEG signals into multiple frequency bands and assigns each band to an independent CNN-Transformer encoder for band-specific modeling. The resulting band-wise features are then adaptively fused by a carefully designed frequency-aware attention (FAA) module that models cross-band dependencies by treating band-wise features as tokens. Further, band-wise auxiliary supervision (BAS) is introduced to prevent weakly contributing branches from being under-optimized during joint training. In this way, FAConformer performs frequency-aware modeling that more effectively exploits frequency domain information. Extensive experiments on two public AAD datasets with three decision-window lengths demonstrated that FAConformer consistently outperformed 12 competitive baselines, surpassing the current state-of-the-art model by 4.9%. Further analyses of band importance, ablation, and parameter sensitivity verify the effectiveness, robustness, and interpretability of the proposed framework. Code is available at https://github.com/wzwvv/FAConformer.
Ziwei Wang, Xingyi He, Tianwang Jia +2
Hubei Key Laboratory of Brain-inspired Intelligent Systems, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China
Identifying which speaker a listener is attending to in a noisy room -- the cocktail-party problem -- is the missing ingredient for next-generation hearing aids and brain-computer interfaces: it tells the device whose voice to amplify. Auditory attention decoding (AAD) reads this answer from EEG, but the literature splits into disconnected pieces: directional-AAD classifies side but does not map side to stream; regression-based source-AAD ranks candidate streams by a single Pearson correlation that is intrinsically noisy at the 1-5 s windows real devices need; and envelope reconstruction has no native AAD rule. We argue the right object is not any single statistic but the conditional likelihood of the attended envelope given EEG, and we make this practical with NEUROTOKEN: a single network whose three heads share one EEG front-end, with a conditional flow-matching head (ATTUNEFLOW) that scores candidates by an integrated velocity-residual likelihood ratio. Two inference-time ensembles -- QUADTRACK (four complementary statistics) and ENV-FLOW (z-normalised QUADTRACK+ATTUNEFLOW) -- absorb per-statistic failure modes for free. On KU Leuven, DTU, and NJU at 5 s, ATTUNEFLOW lifts per-segment source-AAD by 9%-16% over the strongest non-generative baseline and shrinks across-subject variance by ~3x; trial-level fusion exceeds 93% on two of three datasets. In parallel reproductions we show that canonical 95-97% direction-AAD numbers collapse by 17%-45% under a strict trial-disjoint protocol, clarifying both the true ceiling and why a likelihood-based formulation is needed.
Ali Alavi, Donald S. Williamson
Department of Computer Science and Engineering Ohio State University Columbus, OH, 43210
In the past decade, numerous studies have applied deep neural networks (DNNs) to decode auditory attention (AAD) from Electroencephalogram (EEG) signals via stimulus reconstruction. However, the influence of dataset balance on the decoding performance of stimulus reconstruction-based AAD remains unexplored. In this study, three publicly available EEG-AAD datasets - KUL, DTU, and NJU cEEGrid - are used to construct both balanced and unbalanced experimental conditions. We hypothesize and demonstrate that stimulus reconstruction-based DNN decoders tend to produce overestimated decoding performance on unbalanced datasets. To address this issue, we propose a leave-one-paired-envelope-out (LOPEO) cross-validation protocol. Experimental results confirm that LOPEO effectively prevents inflated decoding accuracy on unbalanced datasets. While balanced datasets are generally preferred in experimental design, LOPEO provides a principled evaluation framework for unbalanced datasets that have already been published, filling an important gap in the field.
Yuanming Zhang, Yayun Liang, Zhibin Lin +1
Key Lab of Modern Acoustics, Nanjing University, Nanjing 210093, China · NJU-Horizon Intelligent Audio Lab, Horizon Robotics, Beijing 100094, China