Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.
Figures & tables
Fig. 1: Model architecture of AccentCL. A frozen Whisper-Large-v3 encoder produces multi-layer features that are fused into an accent embedding. Phase 1 trains the regional classifier with domain mean alignment, and Phase 2 expands it with replay, retention, and an old-to-new margin loss.
Method
All Acc ( ↑ )
All Bal Acc ( ↑ )
All Macro-F1 ( ↑ )
OOD Acc ( ↑ )
OOD Bal Acc ( ↑ )
OOD Macro-F1 ( ↑ )
CommonAccent [ 4 ]
56.0
48.3
48.6
79.0
53.6
58.2
Voxlect [ 5 ]
64.8
65.8
68.1
83.8
77.5
82.6
AccentCL base
76.0
77.1
76.9
89.7
79.6
83.0
w/o logit adjustment
75.8
76.0
76.9
87.6
77.6
80.1
w/o DMA
75.7
77.7
75.8
84.8
72.8
75.9
w/o logit adjustment & DMA
75.6
76.0
76.5
81.9
65.1
69.2
TABLE I: Regional accent classification results on the shared 5-class regional label space. Predictions from CommonAccent and Voxlect are mapped to the same regional label space for evaluation. For models with a larger output label space, predictions outside the evaluated 5-class label set are retained and counted as errors. All metrics are reported in percentages.
Fig. 2: Training dataset composition by accent group. Bars show source-dataset percentages, and right-side numbers indicate utterance counts.
Method
New F1 ↑
Old BAcc ↑
Δavg↑
Δworst↑
Base model: 5 regional classes
Frozen Base
–
77.1
–
–
Step 1: 5→6 , adding Spanish-accented English
Voxlect
52.7
66.0
–
–
Replay only
81.4
76.2
-1.6
-4.4 (BI)
Replay + retention
82.8
77.1
-0.7
-3.6 (BI)
TABLE II: Class-incremental accent learning results and ablations. Old BAcc. is computed over the classes learned before each incremental step. Δavg and Δworst denote old-class accuracy changes relative to the old model.
Fig. 3: Per-class recall matrices for Voxlect & AccentCL after adding Spanish-accented English. Out denotes predictions outside the evaluated labels.
Fig. 4: t-SNE of Voxlect (top) and AccentCL embeddings (bottom) after adding Spanish-accented English. Points are colored by accent label (left) and by domain (right).
ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon) as a lightweight, accent-invariant auxiliary objective for CTC fine-tuning. An utterance-level contrastive loss regularizes encoder representations without architectural modification or explicit accent supervision. Experiments on the L2-ARCTIC benchmark show consistent WER reductions across multiple pretrained encoders, with up to 25 -- 29% relative reduction under unseen-accent evaluation. Analysis using within-transcript cosine dispersion indicates that SupCon promotes more compact and stable representation geometry under accent variability. Overall, SupCon provides an effective and model-agnostic regularization strategy for improving accent robustness.
Van-Phat Thai, Aradhya Dhruv, Duc-Thinh Pham +1
Air Traffic Management Research Institute, Nanyang Technological University, Singapore · Center of AI Research, VinUniversity, Vietnam
Regional accent classification in Brazilian Portuguese (pt-BR) suffers from the need for reliable labeling. While large self-supervised learning (SSL) speech models are powerful, their training pipelines dilute sociophonetic information, since accent labels are generally not reliable or are not used in training objectives. This work introduces a novel workflow for feature extraction using only acoustic labels. By isolating explicit regional accent landmarks and using a phoneme-based forced aligner (ZIPA), our targeted feature set captures dialectal variance more effectively than utterance embeddings, demonstrating that localized features can outperform general-purpose architectures on accent-related tasks using minimal and objective data labels.
Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz W. P. Biscainho
PEE/COPPE, UFRJ, Rio de Janeiro-RJ · Faculdade de Engenharia Elétrica e Computação (FEEC), UNICAMP, Campinas-SP · DEL/Poli & PEE/COPPE, UFRJ, Rio de Janeiro-RJ
Accent text-to-speech (TTS) aims to synthesize speech with a target accent while preserving speaker identity, but faces two key challenges: disentangling accent from speaker characteristics and effectively conditioning speech generation on the two disentangled factors. In this paper, we propose Joycent, a diffusion-based accent TTS framework that addresses both challenges. Our key idea is to separate accent and speaker information in both representation learning and TTS conditioning. Joycent uses WhisAID, a Whisper-based accent encoder with gradient reversal to learn speaker-disentangled accent representations, and introduces layer-specific conditional layer normalization to inject accent and speaker information at different stages of the text encoder. We evaluate Joycent on the Mandarin Regional Accent Corpus (MRAC) with seen and unseen speakers, including a challenging cross-accent setting where the speaker and accent prompts come from different accents. Experimental results show that Joycent improves accent similarity over existing methods while maintaining strong speaker similarity, with consistent gains under the challenging cross-accent setting. Subjective evaluation further confirms improved naturalness, accent similarity, and speaker preservation. The audio samples are available at https://oshindow.github.io/joycent/.
Xintong Wang, Junchuan Zhao, Ye Wang
School of Computing, National University of Singapore, Singapore