eess.ASMay 28, 2026

Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels

Authors: Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz W. P. Biscainho

Abstract

Regional accent classification in Brazilian Portuguese (pt-BR) suffers from the need for reliable labeling. While large self-supervised learning (SSL) speech models are powerful, their training pipelines dilute sociophonetic information, since accent labels are generally not reliable or are not used in training objectives. This work introduces a novel workflow for feature extraction using only acoustic labels. By isolating explicit regional accent landmarks and using a phoneme-based forced aligner (ZIPA), our targeted feature set captures dialectal variance more effectively than utterance embeddings, demonstrating that localized features can outperform general-purpose architectures on accent-related tasks using minimal and objective data labels.

Explore similar work

CardsList
  1. Contrastive Regularization for Accent-Robust ASR

    May 5, 2026Van-Phat Thai, Aradhya Dhruv, Duc-Thinh Pham +1Accented SpeechContrastive Learning

  2. PHONOS: PHOnetic Neutralization for Online Streaming Applications

    Mar 27, 2026Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah +1Accented Speech