Speech Emotion Recognition (SER) infers a speaker's emotional state from speech and is increasingly deployed in human-computer interaction, education, and healthcare. Because speech also carries sensitive personal information, speakers may ask that some of their recordings be deleted, which requires removing the influence of those samples from an already trained SER model. Most machine unlearning methods can meet this request only with access to the remaining training data alongside the samples to be forgotten; this is impractical when the remaining data cannot be redistributed or has itself been deleted, and it adds storage and computation as the data grows. To this end, we propose a retain-free unlearning method that updates a pre-trained SER model using only the forget set. Our key idea is to synthesise adversarial samples from the forget set as a surrogate for the unavailable remaining data, and to constrain each parameter update by its estimated importance so that forgetting does not erase general knowledge. The experiments over several emotional-speech corpora and self-supervised backbones show that our method drives forget-set performance down to near chance while retaining much of the model's utility on the remaining and unseen-speaker data, narrowing the gap to methods that rely on the remaining set.
Figures & tables
N=10
N=30
N=50
N=100
Df↓
Dr↑
Dt↑
Df↓
Dr↑
Dt↑
Df↓
Dr↑
Dt↑
Df↓
Dr↑
Dt↑
Original (all)
0.857
0.996
0.902
1.000
0.995
0.902
1.000
0.995
0.902
1.000
0.995
0.902
Original ( Dr )
0.857
0.997
0.907
0.964
0.998
0.905
0.982
0.995
0.901
0.963
0.988
0.891
Upper bound: Machine unlearning with forget set Df and remain set Dr
UNSIR [ 22 ]
1.000
0.978
0.893
0.964
0.968
0.890
0.966
0.954
0.855
0.920
0.941
0.861
Bad Teaching [ 23 ]
0.143
0.974
0.860
0.429
0.995
0.906
0.650
0.994
0.905
0.154
0.993
0.902
Table 1 : Model performance (UAR) on Df , Dr and the test set Dt , when forgetting N samples. The best performance of the proposed approaches are compared to baselines (*: p<0.001 in a one-tailed z-test).
Wav2Vec 2.0
HuBERT
Method
Df↓
Dr↑
Dt↑
Df↓
Dr↑
Dt↑
Original (all)
–
–
0.902
–
–
0.902
N=30 , 1 speaker removed
NegGrad+ [ 26 ]
0.129
0.981
0.894
0.171
0.974
0.863
RandLabel+ [ 11 ]
0.119
0.964
0.852
0.093
0.961
0.845
NegGrad [ 15 ]
0.143
0.218
0.228
0.171
0.207
0.191
Table 2 : Ablation study on DEMoS (UAR).
Wav2Vec 2.0
HuBERT
Method
Df↓
Dr↑
Dt↑
Df↓
Dr↑
Dt↑
Original (all)
–
–
0.659
–
–
0.677
N=30 , 2 speakers removed
NegGrad+ [ 26 ]
0.225
0.990
0.687
0.168
0.764
0.516
RandLabel+ [ 11 ]
0.239
0.948
0.633
0.210
0.847
0.557
NegGrad [ 15 ]
0.420
0.472
0.416
0.205
0.406
0.312
Table 3 : Model comparison on IEMOCAP (UAR).
Fig. 3 : T-SNE visualisation of representations from the linear layer following wav2vec 2.0 when forgetting 10 samples from DEMoS. (a) Original (all), (b) RandLabel, (c) RandLabel-Adv, (d) RandLabel-Adv-EWC.