Prosodic stress is a crucial aspect of automatic pronunciation assessment (APA), encompassing both sentence stress detection (SSD) and word stress detection (WSD). SSD highlights semantically salient words that shape discourse meaning, while WSD identifies the primary stressed syllable within each word to ensure lexical clarity. However, most prior work treats SSD and WSD as independent tasks, overlooking their shared reliance on prosodic cues such as pitch, duration, and intensity. To address this gap, we propose an effective SSD approach combining SSD with auxiliary WSD via a novel modeling paradigm. In addition, we introduce a word-span stress regularizer (WSR) that concentrates token-level SSD probabilities within each stressed word span. Experiments on the TinyStress-15K benchmark show that the proposed method outperforms strong baselines, with the complete configuration achieving the best SSD result.
Figures & tables
Figure 1: Illustration of stress detection at two linguistic levels. (a) SSD highlights the emphasized word within a sentence by shifting prominence across different lexical items. (b) WSD identifies the primary stressed syllable within a lexical item, where primary stress placement is indicated by the digit “1”.
Figure 2: Overall architecture of the proposed framework. A frozen Whisper backbone provides encoder representations to both heads and decoder representations to the SSD head. Sentence stress detection (SSD) operates at the token level with an additional word-span stress regularizer (WSR), while word stress detection (WSD) operates at the phone level. The SSD and WSD heads are task-specific and do not directly exchange hidden states or predictions.
#Audios
#Tokens
#Stress
#Unstress
Train
13,500
171,206
22,537
148,669
Valid
1,500
18,676
2,541
16,135
Test
1,000
12,548
1,705
10,843
Table 1: Statistics of the TinyStress-15K corpus.
Model
Precision
Recall
F1
GT alignment [ 14 ]
0.862
0.853
0.858
MFA [ 14 ]
0.776
0.859
0.815
WhiStress [ 14 ]
0.912
0.906
0.909
STRAW
0.945
0.924
0.934
- WSD
0.938
0.906
0.922
- WSR
0.942
0.917
0.929
Table 2: SSD performance on TinyStress-15K. STRAW denotes our proposed framework with auxiliary WSD and the WSR. Ablations remove the corresponding components.
Model
Precision
Recall
F1
STRAW
0.924
0.916
0.920
- WSR
0.935
0.907
0.921
Table 3: WSD performance of STRAW on TinyStress-15K.
Figure 3: Error analysis by part-of-speech (POS) on the test set. Bars show false negative rate (FNR) and false positive rate (FPR). The black line indicates the sample count per POS.
Speech-to-speech translation (S2ST) systems have achieved impressive progress in semantic accuracy and speech naturalness. However, the cross-lingual transfer of lexical stress, a vital cue for emphasis and speaker intent, remains heavily underexplored, compounded by a lack of reliable automatic evaluation metrics for tonal languages like Chinese. We investigate English-to-Chinese S2ST stress transfer by constructing a stress-annotated Chinese dataset and an XLS-R-based Mandarin stress detector. Integrating this with the English EmphAssess system, we propose a novel objective metric for cross-lingual stress evaluation. Furthermore, we fine-tune CosyVoice3 to build a stress-aware S2ST system. Experiments demonstrate that our proposed S2ST architecture significantly outperforms existing systems in stress translation capability while maintaining competitive translation quality. Furthermore, our evaluation metric exhibits a strong correlation with human subjective judgments.
Yuchen Song, Xi Chen, Mingze Li +1
The Chinese University of Hong Kong, Shenzhen, China · Shenzhen Loop Area Institute, China
Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly measured. The ABX discrimination task has been used to measure phonemic contrast in S3M representations via minimal pairs. We introduce prosodic ABX, an extension of this framework to evaluate prosodic contrast with only a handful of examples and no explicit labels. Also, we build and release a dataset of English and Japanese minimal pairs and use it along with a Mandarin dataset to evaluate contrast in English stress, Japanese pitch accent, and Mandarin tone. Finally, we show that model and layer rankings are often preserved across several experimental conditions, making it practical for low-resource settings.
Haitong Sun, Stephen McIntosh, Kwanghee Choi +3
The University of Tokyo, Japan · University of Texas at Austin, USA
Automatically detecting stress in speech provides an unobtrusive way to gain insights relevant to behavioral research or clinical assessment. This study investigates the automatic differentiation between a stressful and non-stressful situation, and the prediction of physiological and affective stress responses. Speech data was collected from 50 participants who either completed the Trier Social Stress Test (TSST) or a non-stressful control condition. With a processing pipeline that included speaker diarization and machine learning models, we achieved stress detection performance significantly above a mean baseline. Moreover, relevant physiological and affective stress responses were partially predictable from acoustic-prosodic features. Feature-importance analyses identified the most informative predictors contributing to model performance. The findings demonstrate that speech can serve as a meaningful and unobtrusive indicator of multiple dimensions of the human stress response.
Hanna Drimalla, Wieland R. Cremer, Christine Kraus +1
Human-Centered Artificial Intelligence Group, Faculty of Technology, Bielefeld University, Bielefeld, Germany. · Department of Cognitive Psychology, Faculty of Psychology, Ruhr University Bochum, Bochum, Germany.