A multi-scenario EEG dataset for auditory attention decoding in naturalistic multi-talker environments
Authors: Shu Peng, Rui Liu, Yufei Zhang, Wenlong You, Zhige Chen, Jiachen Xi, Qiyuan Sun, Yan Liu, +2 more
Organizations: Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Hong Kong SAR, China · Department of Computing, The Hong Kong Polytechnic University, Hong Kong SAR, China
Understanding how the brain selectively follows relevant speech amid competing voices is a central challenge in auditory neuroscience and a key step toward neuro-steered hearing technologies. However, most open-source Electroencephalography (EEG) datasets for Auditory Attention Decoding (AAD) use idealized single-competing-talker paradigms that oversimplify the acoustic, spatial, and semantic structure of everyday communication. To capture this ecological complexity, we introduce the SoundBubble-EEG dataset: a high-density 128-channel EEG resource comprising more than 25 hours of recordings from 30 participants. The paradigm requires listeners to selectively attend to a dynamic target speaker group, a designated "sound bubble", amid competing multi-speaker distractor bubbles across three realistic scenarios: a restaurant, a home TV viewing, and a meeting discussion. By bridging the gap between constrained laboratory protocols and real-world auditory scenes, this dataset enables investigations of multi-talker speech comprehension, neural speech tracking, and cross-scenario generalization. It also provides a benchmark for AAD algorithms under realistic acoustic and semantic variability and may support auditory neuroscience and the development of neuro-steered hearing technologies.
Figures & tables
Figure 1 : Overview of the SoundBubble-EEG dataset generation, validation, and public release workflow. Participants are engaged in three ecologically valid listening scenarios: meeting, restaurant, and living room (orange). The acquired EEG signals were subsequently processed using a standardized preprocessing pipeline (blue) and underwent rigorous validation, including behavioral assessments, signal quality evaluations, and neural/application-level analyses (green). Finally, the dataset was publicly released in a BIDS-compliant format, accompanied by comprehensive analysis code to facilitate reproducibility and further research (purple).
Figure 2 : Illustration of the "Sound Bubble" paradigm and the experimental procedure. (a) The traditional competing-talker paradigm, which typically features two isolated single speakers (e.g., one on the left versus one on the right) without any environmental background sound. The icon indicates the target sound source that the participant is instructed to focus on. (b) The proposed "Sound Bubble" paradigm replaces individual talker scenarios with clusters of multiple speakers, referred to as "Sound Bubbles", each engaged in discussions centered on a coherent theme and positioned within a realistic acoustic environment. This design requires the listener to direct attention to a target bubble while suppressing a multi-speaker distractor bubble and ambient noise, significantly enhancing ecological validity. (c)–(e) Three ecologically valid listening scenarios were designed for the dataset, in which the green and yellow bubbles indicate the target (attended) and non-target (unattended) sound bubbles, respectively: (c) Restaurant scene characterized by high ambient noise and two competing conversational sound bubbles; (d) Home TV viewing scene, where a fixed media source (TV) competes with side family conversations in a low-noise background; and (e) Meeting scene, where multiple participants form two discussion groups and each speak from a different direction. (f) Experimental procedure. The formal session followed the actual scenario order of Meeting (5 trials), Restaurant (10 trials), and TV (10 trials). In each trial, a black screen with a green directional cue indicated the target sound source and remained visible throughout listening; the first 3 seconds served as baseline and the following 120 seconds as the active-listening task period. Participants then answered a general-comprehension question (Q1) and a specific-details question (Q2), followed by a self-paced rest before the next trial.
Subject ID
Age
Sex
Handedness
Q1 Mean Acc. (%)
Q2 Mean Acc. (%)
sub-01
26
M
R
100.0
100.0
sub-02
25
M
R
96.0
76.0
sub-03
26
M
R
100.0
84.0
sub-04
19
F
R
100.0
80.0
sub-05
26
M
L
96.0
60.0
sub-06
24
M
R
96.0
80.0
Table 1 : Participant demographics and behavioral accuracy. This table details the age, sex, handedness, and formal-trial behavioral accuracy for each participant (‘M’ = Male, ‘F’ = Female; ‘R’ = Right, ‘L’ = Left). Q1 assesses general comprehension of the attended stream, while Q2 focuses on specific details.
Stimulus ID
Scenario
Listener Facing
Sound Bubble 1 (Pos)
Sound Bubble 2 (Pos)
Noise Source (Pos)
S06015
Meeting
0∘
−85∘
+27∘
+57∘
S06746
Meeting
0∘
−53∘
+79∘
−57∘
S06943
Meeting
0∘
+77∘
−36∘
−61∘
S08364
Meeting
0∘
+72∘
−67∘
−50∘
S08440
Meeting
0∘
+61∘
−56∘
−17∘
S00358
Restaurant
0∘
−36∘
+77∘
+44∘
Table 2 : Spatial configurations of the 25 auditory stimulus templates. Azimuth angles are expressed relative to the listener-facing direction ( 0∘ = front; negative values = left; positive values = right). Elevation is 0∘ (eye level) for all sources.
Figure 4 : Behavioral performance and data quality assessment. (a) Behavioral accuracy for general understanding (Q1) and specific details (Q2) across the three acoustic conditions (Restaurant, TV, Meeting). The dashed line indicates chance level (0.25). (b) Sensor layout of the 128-channel high-density EEG net. (c) Distribution of interpolated bad channels per run across three acoustic scenarios. (d) Number of artifact independent components excluded via ICA per run across three acoustic scenarios. (e) Spatial probability heatmap of bad channels across all recordings.
Figure 5 : Neurophysiological sanity checks and application-level AAD validation. (a) Grand average global Power Spectral Density (PSD) for Baseline (dashed grey) and Task (solid red) states, with the alpha band (8–13 Hz) highlighted in yellow. (b) Topography of normalized alpha power change ( 10log10(Task/Baseline) ). Blue regions indicate relative alpha suppression, while red regions denote relative alpha enhancement. (c) Auditory attention decoding performance using a backward mTRF model. Paired raincloud plots show Pearson correlations ( r ) between EEG-reconstructed and actual speech envelopes, demonstrating significantly higher cortical tracking for attended speech (one-sided Wilcoxon signed-rank test, *** p=9.31×10−10 ). (d,e) Trial-label swap null controls for the group-mean decoding advantage Δr=rattend−runattend and single-trial decoding accuracy. Shaded distributions show label-swapped null distributions; vertical colored lines show the observed values computed from the true attended/unattended labels. (f,g) Scenario-wise AAD robustness for Δr and decoding accuracy across Restaurant, TV, and Meeting scenarios. Scenario-wise panels show subject-level scenario means; Restaurant and TV contained 10 trials per subject, whereas Meeting contained 5 trials per subject.
Hubei Key Laboratory of Brain-inspired Intelligent Systems, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China