Regional accent cues can be captured under matched conditions, but it remains unclear whether they persist between read and spontaneous speech. We study RVG1, with 500 German speakers from nine regions, comparing ten speech representations on regional classification and continuous geolocation under matched conditions and speaker-independent read--spontaneous transfer. Whisper performs best under matched conditions, reaching 0.489 nine-way UAR and 148 km median geolocation error, but drops to 0.11/0.18 UAR across transfer directions and 363 km geolocation error. Self-supervised models show a similar degradation, whereas speaker embeddings are less discriminative in-domain but more robust under transfer. This contrast is consistent across classification and geolocation. Across representations, robustness is associated with how little a representation shifts between styles (style-invariance), for which crossstyle speaker retrieval is an interpretable proxy. Age, sex, sentence-overlap, and duration controls do not account for the gap, although channel characteristics contribute. These results show that strong matched-condition performance does not indicate robust regional information.
Figures & tables
Figure 1: Overview of the study. Ten speech representations are evaluated on German regional classification and geolocation under matched conditions and speaker-independent read–spontaneous transfer.
Family
Representation
9-way
6-way
3-way
2-way
ASR
Whisper
0.489 / 0.568
0.627 / 0.691
0.767 / 0.788
0.835 / 0.841
SSL
WavLM
0.397 / 0.480
0.504 / 0.565
0.702 / 0.720
0.775 / 0.792
wav2vec2-XLSR
0.391 / 0.488
0.486 / 0.548
0.698 / 0.719
0.800 / 0.809
HuBERT
0.385 / 0.469
0.492 / 0.561
0.664 / 0.688
0.762 / 0.783
Audio-LLM
Qwen2-Audio
0.323 / 0.381
0.419 / 0.465
0.616 / 0.647
0.798 / 0.780
Speaker
x-vector
0.270 / 0.352
0.390 / 0.455
0.619 / 0.640
0.741 / 0.749
Table 1: Regional classification UAR / ACC values on read speech. Prior published RVG1 results are listed for context only (different partitions/preprocessing; not a controlled comparison).
Figure 2: Regional heterogeneity of the Whisper representation. Recall ranges from 0.10 (South Franconian) to 0.65 (Bavarian) and localization error from 116 to 208 km.
Figure 3: Read–spontaneous transfer for classification (top) and geolocation (bottom). (a) Cross-condition UAR by granularity, with a 9-way ranking reversal. (b) 9-way transfer by direction; content encoders degrade most. (c) Speaker retrieval vs. 9-way UAR ( ρ=+0.74 ). (d) In-domain geolocation error (Whisper: 148 km). (e) Geolocation transfer by direction; speaker embeddings and Mimi remain near baseline. (f) Speaker retrieval vs. transfer error ( ρ=−0.68 ; inverted y -axis).
Spoken language identification (LID) aims to recognize the target language regardless of accent. In practice, however, LID models fine-tuned from self-supervised speech representations frequently confuse accents with languages, misclassifying non-native (L2) speech as the speaker's first language (L1). We show that non-native speech representations lie between native target-language and native L1 poles, causing systematic misclassification. To address this, we introduce a geometric projection that estimates an L1-bias direction solely from native speech and removes it before the frozen LID head. Across five MMS-LID models and non-native corpora, this projection substantially improves target language identification for L2-accented speech while preserving predictions for native speech. These results show that accent-induced L1 bias can be corrected directly within the representation space without L2 training data or model adaptation.
Minu Kim, Jihwan Lee, David R. Mortensen +1
Signal Analysis and Interpretation Lab (SAIL), University of Southern California, USA · Language Technologies Institute, Carnegie Mellon University, USA
ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon) as a lightweight, accent-invariant auxiliary objective for CTC fine-tuning. An utterance-level contrastive loss regularizes encoder representations without architectural modification or explicit accent supervision. Experiments on the L2-ARCTIC benchmark show consistent WER reductions across multiple pretrained encoders, with up to 25 -- 29% relative reduction under unseen-accent evaluation. Analysis using within-transcript cosine dispersion indicates that SupCon promotes more compact and stable representation geometry under accent variability. Overall, SupCon provides an effective and model-agnostic regularization strategy for improving accent robustness.
Van-Phat Thai, Aradhya Dhruv, Duc-Thinh Pham +1
Air Traffic Management Research Institute, Nanyang Technological University, Singapore · Center of AI Research, VinUniversity, Vietnam
Determining the differences between two speakers' accents is a fundamental task in linguistics and speech technology research. The methodology used to measure these differences depends on the specific research area. A phonetics researcher may demonstrate accent variation by comparing vowel formants in paired recordings of individual words. These results will be interpretable, but the recordings will be time-consuming to collect and may not be representative of connected speech. Accented Text-to-Speech (TTS) research has pushed towards using accent embeddings derived from accent classification tasks. These embeddings can be produced from any speech recording, but are not readily interpretable. In this paper, we demonstrate that articulatory representations created through articulatory inversion can be used as an interpretable basis for accent comparison and that optimal transport provides a framework for accent comparison across arbitrary recording types.
Charles McGhee, Mark J. F. Gales, Kate M. Knill
ALTA Institute/MIL, Department of Engineering, University of Cambridge, UK