Misophonia, a disorder of decreased tolerance to specific sounds, affects 5-20% of the population, yet sufferers have no good options: therapy helps a minority, and earplugs or noise cancellation silence everything. We present a study for neural trigger sound suppression for misophonia, selectively removing trigger sounds. We curate a dataset covering the 10 most common trigger classes. Using streaming dual-path networks operating on 6 ms audio chunks, we explore both one-hot and multi-hot-conditioned models that suppress 1-3 triggers from the acoustic scene. We validate our model outputs in a listening study with 30 adults with clinically elevated misophonia impairment. Participants reported significantly lower distress and arousal, and improved valence, for suppressed audio.
Figures & tables
Trigger
Chewing
Cough
Drinking
Gum pop
Heavy breath
Lip smack
Sneezing
Sniffling
Tapping
Throat clearing
Total
FSD50K
231
-
-
-
-
-
-
-
-
-
231
CoughVID
-
1,187
-
-
-
-
-
-
-
-
1,187
ESC-50
-
40
40
-
-
-
20
-
-
-
100
MATA
-
-
34
-
-
21
12
28
-
-
95
VocalSound
-
-
-
-
-
-
332
-
-
3,504
3,836
Freesound
-
-
-
4
-
-
38
8
149
-
199
Table 1 : Number of audio clips per trigger class in our dataset.
Trigger Class
SI-SNRi (dB)
SNRi (dB)
Chewing
11.88±4.36
13.18±4.26
Cough
15.79±5.47
15.91±5.30
Drinking
10.73±3.52
11.74±3.48
Gum popping
13.29±4.07
13.85±3.96
Heavy Breathing
12.48±3.79
13.31±3.81
Lip smacking
14.35±5.27
15.07±5.05
Table 2 : Performance of a single-trigger multi-class model trained across all 10 trigger classes, evaluated with single-trigger mixtures and a single output channel.
# selected trigger classes
Model
Metric
1
2
3
Ours
SI-SNRi
19.86±5.82
18.17±5.23
16.72±5.10
SNRi
19.99±5.47
18.69±4.82
17.78±4.43
Waveformer
SI-SNRi
12.92±5.06
12.30±4.17
11.90±3.88
SNRi
13.67±4.87
13.76±4.07
14.07±3.93
Table 3: Multi-trigger suppression model with 0.496M parameters, compared to 2.020M for Waveformer baseline.
( LC/LB/LF , ms)
Latency
SNRi (dB)
SI-SNRi (dB)
6/6/4
10 ms
19.99±5.47
19.86±5.82
8/4/4
12 ms
17.08±5.02
16.76±5.25
4/12/0
4 ms
13.38±4.84
12.55±5.08
Table 4 : Ablation study over STFT parameters for the multi-trigger model, evaluated with a single trigger queried.
Figure 1 : Time-domain waveforms for gum popping suppression: input mixture, ground truth, and model output.
Measure
Original
Output
Δ [95% CI]
Distress (1–10)
7.68±1.60
5.31±2.48
−2.37 [1.45, 3.28]
Arousal (1–9)
7.02±1.42
4.89±2.19
−2.13 [1.36, 2.91]
Valence (1–9)
4.42±2.75
6.01±1.65
+1.59 [1.05, 2.13]
Table 5 : Affective response to raw audio vs. model output (30 participants). Lower is better for distress and arousal; higher better for valence. All paired t -tests p<.001 .