NEUROTOKEN: Joint Source and Directional AAD with Envelope Decoding via Conditional Flow Matching
Organizations: Department of Computer Science and Engineering Ohio State University Columbus, OH, 43210
Abstract
Identifying which speaker a listener is attending to in a noisy room -- the cocktail-party problem -- is the missing ingredient for next-generation hearing aids and brain-computer interfaces: it tells the device whose voice to amplify. Auditory attention decoding (AAD) reads this answer from EEG, but the literature splits into disconnected pieces: directional-AAD classifies side but does not map side to stream; regression-based source-AAD ranks candidate streams by a single Pearson correlation that is intrinsically noisy at the 1-5 s windows real devices need; and envelope reconstruction has no native AAD rule. We argue the right object is not any single statistic but the conditional likelihood of the attended envelope given EEG, and we make this practical with NEUROTOKEN: a single network whose three heads share one EEG front-end, with a conditional flow-matching head (ATTUNEFLOW) that scores candidates by an integrated velocity-residual likelihood ratio. Two inference-time ensembles -- QUADTRACK (four complementary statistics) and ENV-FLOW (z-normalised QUADTRACK+ATTUNEFLOW) -- absorb per-statistic failure modes for free. On KU Leuven, DTU, and NJU at 5 s, ATTUNEFLOW lifts per-segment source-AAD by 9%-16% over the strongest non-generative baseline and shrinks across-subject variance by ~3x; trial-level fusion exceeds 93% on two of three datasets. In parallel reproductions we show that canonical 95-97% direction-AAD numbers collapse by 17%-45% under a strict trial-disjoint protocol, clarifying both the true ceiling and why a likelihood-based formulation is needed.
Figures & tables
| Symbol | Loss | Role | S1 | S2 |
| Spatial CE | Direction logits , label-smooth 0.05 | – | ||
| Envelope warm-up Pearson ( 3.4 ) | Maximise | – | ||
| Envelope -margin ( 3.4 ) | Force | – | ||
| Cross-batch envelope InfoNCE ( 5 ) | Contrast attendeds across batch + same-trial unatt. negative | – | ||
| Multi-resolution STFT magnitude | L1 magnitude of vs at frames | – | ||
| Spatial–envelope consistency | Encourage side side of stronger (App. B.6 ) | – |
| Dataset | Win | env- | env-multi | attune-flow | env-multi-flow | env-trial | env-trial-fused | spatial | spatial-trial |
|---|---|---|---|---|---|---|---|---|---|
| KU Leuven (15) | 1s | ||||||||
| KU Leuven (16) | 5s | ||||||||
| DTU (18) | 1s | ||||||||
| DTU (18) | 5s | ||||||||
| NJU (21) | 1s | ||||||||
| NJU (21) | 5s |
| Score | 1 s (%) | 5 s (%) | 1 5 s (pp) |
|---|---|---|---|
| env- | |||
| QuadTrack -Multi | |||
| AttuneFlow -Flow | |||
| Env-Multi-Flow |
| Dataset | per-seg Env-Multi-Flow | per-seg AttuneFlow -Flow | trial AttuneFlow -EM-OR | |
|---|---|---|---|---|
| KU Leuven | 6 | |||
| DTU | 6 | |||
| NJU | 7 |
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Type | Backbone | Cons. audio | Gen. env. | Gen. audio |
|---|---|---|---|---|---|
| mTRF [ 9 ] | Reconstructive | Linear ridge | yes | yes | no |
| VLAAI [ 1 ] | Reconstructive | CNN | yes | yes | no |
| DECAF [ 36 ] | Reconstructive | CNN + margin | yes | yes | no |
| SSM2Mel [ 15 ] | Reconstructive | State-space | yes | yes | no |
| DMF2Mel [ 14 ] | Reconstructive | State-space | yes | yes | no |
| wav2vec align [ 11 ] | Alignment | Transformer (frozen) | yes | no | yes |
| Rule | Inputs | Train cost | Best regime |
|---|---|---|---|
| Env-r | none | easy subjects, long windows | |
| Env-Multi | none | noisy single-statistic failure modes | |
| AttuneFlow | flow head | long windows, strong – | |
| Env-Multi-Flow | all of the above | flow head | default per-segment decision |
| Spatial | raw EEG | none (closed-form CSP) | wide-azimuth ( ) datasets |
| Spatial-Trial | per-trial Spatial votes | none | wide-azimuth, intra-subject |
| Dataset | Env-Multi | ||||
|---|---|---|---|---|---|
| KU Leuven | 60.4 | 60.7 | 64.5 | 62.1 | 71.7 |
| DTU | 58.9 | 59.4 | 67.0 | 60.2 | 75.5 |
| NJU | 61.8 | 62.6 | 65.3 | 60.9 | 70.8 |
| Dataset | QuadTrack -Multi | AttuneFlow -Flow | Env-Multi-Flow |
|---|---|---|---|
| KU Leuven (16) | |||
| DTU (18) | |||
| NJU (21) | |||
| mean |
| Probe | Recipe | Env-multi | Spatial | Env-concat | Sp-trial |
|---|---|---|---|---|---|
| D1 | CSP shrink , env-infonce | ||||
| D2 | Baseline (shrink 0.3, margin 8) | ||||
| E1 | D2 with GroupNorm (BN GN) in SpatialConv | ||||
| E1b | E1 + rel. margin + hard-NCE | ||||
| E2 | E1b + Env-SubjectAdapter (neuro stats) | ||||
| F | D2 + Match-Mismatch head |
| Held-out subject | Env-multi | Spatial | Env-concat | Sp-trial |
|---|---|---|---|---|
| S1 (D2) | ||||
| S16 (J) |
| Epoch | Total loss | |||
|---|---|---|---|---|
| 0 | ||||
| 1 | ||||
| 2 | ||||
| 3 | ||||
| 4 | ||||
| 5 |
| Dataset | env- | env-multi | attune-flow | env-multi-flow | spatial | spatial-trial | em-trial-OR | |
|---|---|---|---|---|---|---|---|---|
| KU Leuven | 6 | |||||||
| DTU | 6 | |||||||
| NJU | 7 |
| Held-out | env- | env-multi | attune-flow | env-multi-flow | spatial | spatial-trial | em-trial-OR |
|---|---|---|---|---|---|---|---|
| S1,S2,S3 | 51.0 | 60.1 | 61.3 | 64.8 | 76.7 | 70.0 | 96.7 |
| S4,S5,S6 | 52.5 | 60.7 | 64.2 | 66.0 | 59.5 | 61.7 | 100.0 |
| S7,S8,S9 | 53.2 | 56.3 | 56.2 | 57.2 | 62.6 | 66.7 | 95.0 |
| S10,S11,S12 | 52.5 | 57.1 | 62.4 | 63.1 | 59.5 | 61.7 | 98.3 |
| S13,S14,S15 | 52.4 | 60.1 | 62.4 | 64.1 | 75.2 | 73.3 | 98.3 |
| S14,S15,S16 | 51.4 | 60.1 | 63.8 | 65.1 | 84.6 | 85.0 | 100.0 |
| Held-out | env- | env-multi | attune-flow | env-multi-flow | spatial | spatial-trial | em-trial-OR |
|---|---|---|---|---|---|---|---|
| S1,S2,S3 | 52.6 | 59.0 | 54.7 | 58.4 | 47.7 | 45.0 | 83.9 |
| S4,S5,S6 | 50.6 | 57.9 | 53.8 | 57.1 | 49.1 | 48.3 | 76.1 |
| S7,S8,S9 | 53.0 | 63.0 | 54.4 | 61.4 | 49.6 | 49.4 | 87.8 |
| S10,S11,S12 | 53.2 | 57.3 | 51.6 | 57.6 | 53.1 | 51.1 | 81.7 |
| S13,S14,S15 | 53.3 | 62.1 | 54.9 | 61.4 | 49.8 | 48.3 | 86.1 |
| S16,S17,S18 | 51.5 | 58.4 | 54.8 | 59.2 | 50.8 | 48.9 | 83.3 |
| Held-out | env- | env-multi | attune-flow | env-multi-flow | spatial | spatial-trial | em-trial-OR |
|---|---|---|---|---|---|---|---|
| S02,S03,S04 | 53.5 | 49.3 | 49.8 | 49.7 | 50.5 | 51.0 | 64.6 |
| S06,S07,S08 | 52.5 | 51.0 | 52.2 | 52.0 | 52.5 | 55.2 | 72.9 |
| S09,S12,S13 | 52.2 | 51.1 | 51.4 | 53.3 | 52.2 | 52.1 | 81.2 |
| S14,S15,S16 | 50.4 | 49.6 | 52.7 | 51.8 | 54.3 | 56.2 | 77.1 |
| S17,S18,S19 | 51.3 | 50.5 | 50.9 | 50.0 | 49.7 | 51.0 | 75.0 |
| S21,S22,S23 | 50.0 | 49.7 | 49.7 | 48.2 | 51.8 | 54.4 | 74.4 |
| Layer | Op | Output shape | #Params |
| SpatialConv | |||
| temporal_conv | Conv2d( , , , no bias) | 400 | |
| spatial_conv | Conv2d( , , groups , no bias) | 1 024 | |
| norm | BatchNorm2d(16) | 32 | |
| act | GELU; squeeze axis 2 | — | |
| Multi-scale temporal | |||
| Layer | Op | Output shape | #Params |
| fixed FIR bandpass | 257 taps, frozen, 0.5–9 Hz | 0 | |
| SpatialConv (8 filters) | Conv2d ; Conv2d groups ; BN; GELU | 728 | |
| temporal_proj | Conv1d( , ) | 432 | |
| Residual dilated stack: 4 layers, kernel 15, dilations 1/2/4/8 | |||
| layer 1 (d=1) | Conv1d(48,48,15,p=7); GELU; Drop(0.2); residual | 34 608 | |
| layer 2 (d=2) | Conv1d(48,48,15,p=14,d=2); GELU; Drop(0.2); residual | 34 608 | |
| Layer | Op | Output shape | #Params |
|---|---|---|---|
| stats | cat([mean, std]) over time | — | |
| encoder | Linear(128 32); GELU; Linear(32 128) (zero-init) | 8 352 | |
| chunk | split into , | — | |
| FiLM apply | — | ||
| lag_predictor | Linear(128 32); GELU; Linear(32 1) (zero-init) | 4 161 | |
| lag bound | frames | — |
| Layer | Op | Output shape | #Params |
|---|---|---|---|
| csp_bandpass | rfft mask 8–30 Hz; irfft | 0 | |
| csp_filters (frozen) | einsum | 0 (buf, 512) | |
| log-var | — | ||
| bp_bandpass (6 bands) | RFFT, mag 2 /T, | 0 | |
| alpha key-pair asym | at 6 pairs | — | |
| concat (CSP BP) | — | — |
| Layer | Op | Output shape | #Params |
| Branch A: residual dilated CNN over output | |||
| proj_in | Conv1d(48 48, ) | — (Identity) | |
| block d=1 | Conv1d(48,48,9,p=4); GELU; Drop(0.2); residual | 20 784 | |
| block d=2 | Conv1d(48,48,9,p=8,d=2); GELU; Drop(0.2); residual | 20 784 | |
| block d=4 | Conv1d(48,48,9,p=16,d=4); GELU; Drop(0.2); residual | 20 784 | |
| block d=8 | Conv1d(48,48,9,p=32,d=8); GELU; Drop(0.2); residual | 20 784 | |
| Layer | Op | Output shape | #Params |
| time embed | sin/cos to , broadcast over | — | |
| time_mlp | Linear(32 128); GELU; Linear(128 128) | 20 736 | |
| concat | — | — | |
| input_proj | Conv1d(177 128, ) | 22 784 | |
| 8 stacked velocity blocks (each, identical) | |||
| conv_a | Conv1d(128,128,9,p=4,d=1) | 147 584 | |
| Layer | Op | Output shape | #Params |
| EEG branch | |||
| adaptive_avg_pool1d; flatten | to 8 bins | — | |
| eeg_mlp | Linear(384 256); GELU; Drop(0.1); Linear(256 128) | 131 456 | |
| L2 normalise | — | ||
| Envelope branch | |||
| adaptive_avg_pool1d; flatten | to 8 bins | — | |
| Layer | Op | Output shape | #Params |
|---|---|---|---|
| gammatone bank | 8 subbands, 150–2000 Hz, fixed FIR | 0 | |
| half-wave rectify | 0 | ||
| low-pass + downsample | to 32 Hz, frames | 0 | |
| broadband mean | average over 8 bands | 0 | |
| Total | 0 |
| Layer | Op | Output shape | #Params |
|---|---|---|---|
| 3 fixed FIR bandpass | 129-tap, theta/alpha/beta (frozen) | ea. | 0 |
| spatial_filters[3 ] | Conv1d(64 4, , no bias) | ea. | 768 |
| square; Hann smooth (250 ms); | — | 0 | |
| adaptive_avg_pool1d | to | — | |
| proj | Linear(12 12); GELU | 156 | |
| Total | 924 |
| Module | Stage 1 trainable | Stage 2 trainable |
| (EEG common front-end, ) | 23 360 | frozen |
| (envelope-spatial front-end, 0.5–9 Hz) | 140 024 | frozen |
| Subject adapter (FiLM lag, frames) | 12 513 | frozen |
| (CSP BandPower spatial head; CSP buf frozen) | 30 071 | frozen |
| (residual dilated CNN linear mTRF) | 84 210 | frozen |
| Auxiliary BandPower features (3 bands 4 spatial) | 924 | frozen |
| Model | S1 P / S | S2 P / S | S3 P / S | S4 P / S | Mean P (%) | Mean S (%) | (pp) |
|---|---|---|---|---|---|---|---|
| DARNet | 94.0/39.6 | 84.0/68.8 | 97.0/68.8 | 83.0/43.8 | 89.5 | 55.2 | 34.3 |
| ListenNet | 89.9/33.3 | 92.0/52.1 | 100.0/95.5 | 89.9/12.5 | 93.0 | 48.4 | 44.6 |
| DBPNet | 83.0/10.4 | 86.4/50.0 | 96.6/92.1 | 87.5/40.6 | 88.4 | 48.3 | 40.1 |
| DenseNet-3D | 82.4/76.7 | 91.8/88.8 | 95.1/73.3 | 76.4/40.2 | 86.4 | 69.8 | 16.6 |
| SWIM | 62.5/4.2 | 61.1/49.4 | 76.4/87.5 | 57.6/0.0 | 64.4 | 35.3 | 29.2 |