A Novel Sentence Stress Detection Framework Leveraging Auxiliary Word-Stress Modeling and Loss Optimization
Organizations: National Taiwan Normal University, Taiwan
Abstract
Prosodic stress is a crucial aspect of automatic pronunciation assessment (APA), encompassing both sentence stress detection (SSD) and word stress detection (WSD). SSD highlights semantically salient words that shape discourse meaning, while WSD identifies the primary stressed syllable within each word to ensure lexical clarity. However, most prior work treats SSD and WSD as independent tasks, overlooking their shared reliance on prosodic cues such as pitch, duration, and intensity. To address this gap, we propose an effective SSD approach combining SSD with auxiliary WSD via a novel modeling paradigm. In addition, we introduce a word-span stress regularizer (WSR) that concentrates token-level SSD probabilities within each stressed word span. Experiments on the TinyStress-15K benchmark show that the proposed method outperforms strong baselines, with the complete configuration achieving the best SSD result.
Figures & tables
| #Audios | #Tokens | #Stress | #Unstress | |
| Train | 13,500 | 171,206 | 22,537 | 148,669 |
| Valid | 1,500 | 18,676 | 2,541 | 16,135 |
| Test | 1,000 | 12,548 | 1,705 | 10,843 |
| Model | Precision | Recall | F1 |
|---|---|---|---|
| GT alignment [ 14 ] | 0.862 | 0.853 | 0.858 |
| MFA [ 14 ] | 0.776 | 0.859 | 0.815 |
| WhiStress [ 14 ] | 0.912 | 0.906 | 0.909 |
| STRAW | 0.945 | 0.924 | 0.934 |
| - WSD | 0.938 | 0.906 | 0.922 |
| - WSR | 0.942 | 0.917 | 0.929 |
| Model | Precision | Recall | F1 |
|---|---|---|---|
| STRAW | 0.924 | 0.916 | 0.920 |
| - WSR | 0.935 | 0.907 | 0.921 |