cs.CLOct 5, 2026

AccentCL: Robust Accent Classification with Incremental Expansion

Authors: Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna

Organizations: Department of Computer Science and Engineering Texas A&M University College Station, TX 77840, USA

Abstract

Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available. Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora. We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift. AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora. The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes. On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score. We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English. When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes. When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes. These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.

Figures & tables

Explore similar work

May 5, 2026cs.SD

Contrastive Regularization for Accent-Robust ASR

ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon) as a lightweight, accent-invariant auxiliary objective for CTC fine-tuning. An utterance-level contrastive loss regularizes encoder representations without architectural modification or explicit accent supervision. Experiments on the L2-ARCTIC benchmark show consistent WER reductions across multiple pretrained encoders, with up to 25 -- 29% relative reduction under unseen-accent evaluation. Analysis using within-transcript cosine dispersion indicates that SupCon promotes more compact and stable representation geometry under accent variability. Overall, SupCon provides an effective and model-agnostic regularization strategy for improving accent robustness.
May 28, 2026eess.AS

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

Regional accent classification in Brazilian Portuguese (pt-BR) suffers from the need for reliable labeling. While large self-supervised learning (SSL) speech models are powerful, their training pipelines dilute sociophonetic information, since accent labels are generally not reliable or are not used in training objectives. This work introduces a novel workflow for feature extraction using only acoustic labels. By isolating explicit regional accent landmarks and using a phoneme-based forced aligner (ZIPA), our targeted feature set captures dialectal variance more effectively than utterance embeddings, demonstrating that localized features can outperform general-purpose architectures on accent-related tasks using minimal and objective data labels.
Jun 15, 2026cs.SD

Joycent: Multi-Accent TTS via Disentangled Accent Modeling and Layer-Specific Conditioning

Accent text-to-speech (TTS) aims to synthesize speech with a target accent while preserving speaker identity, but faces two key challenges: disentangling accent from speaker characteristics and effectively conditioning speech generation on the two disentangled factors. In this paper, we propose Joycent, a diffusion-based accent TTS framework that addresses both challenges. Our key idea is to separate accent and speaker information in both representation learning and TTS conditioning. Joycent uses WhisAID, a Whisper-based accent encoder with gradient reversal to learn speaker-disentangled accent representations, and introduces layer-specific conditional layer normalization to inject accent and speaker information at different stages of the text encoder. We evaluate Joycent on the Mandarin Regional Accent Corpus (MRAC) with seen and unseen speakers, including a challenging cross-accent setting where the speaker and accent prompts come from different accents. Experimental results show that Joycent improves accent similarity over existing methods while maintaining strong speaker similarity, with consistent gains under the challenging cross-accent setting. Subjective evaluation further confirms improved naturalness, accent similarity, and speaker preservation. The audio samples are available at https://oshindow.github.io/joycent/.