AccentCL: Robust Accent Classification with Incremental Expansion
Organizations: Department of Computer Science and Engineering Texas A&M University College Station, TX 77840, USA
Abstract
Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.
Figures & tables
| Method | All Acc ( ) | All Bal Acc ( ) | All Macro-F1 ( ) | OOD Acc ( ) | OOD Bal Acc ( ) | OOD Macro-F1 ( ) |
|---|---|---|---|---|---|---|
| CommonAccent [ 4 ] | 56.0 | 48.3 | 48.6 | 79.0 | 53.6 | 58.2 |
| Voxlect [ 5 ] | 64.8 | 65.8 | 68.1 | 83.8 | 77.5 | 82.6 |
| AccentCL base | 76.0 | 77.1 | 76.9 | 89.7 | 79.6 | 83.0 |
| w/o logit adjustment | 75.8 | 76.0 | 76.9 | 87.6 | 77.6 | 80.1 |
| w/o DMA | 75.7 | 77.7 | 75.8 | 84.8 | 72.8 | 75.9 |
| w/o logit adjustment & DMA | 75.6 | 76.0 | 76.5 | 81.9 | 65.1 | 69.2 |
| Method | New F1 | Old BAcc | ||
|---|---|---|---|---|
| Base model: 5 regional classes | ||||
| Frozen Base | – | 77.1 | – | – |
| Step 1: , adding Spanish-accented English | ||||
| Voxlect | 52.7 | 66.0 | – | – |
| Replay only | 81.4 | 76.2 | -1.6 | -4.4 (BI) |
| Replay + retention | 82.8 | 77.1 | -0.7 | -3.6 (BI) |