Automatic speaking assessment systems can provide holistic proficiency scores, but often lack interpretable measures that characterize pronunciation quality. We propose a native-reference phone-class geometry for measuring second language (L2) pronunciation deviation without requiring pronunciation labels, read-aloud prompts, or matched recordings of the same text from native and L2 speakers. Given a native speech corpus, we average frame-level self-supervised representations for each context-dependent phone-class and use singular value decomposition (SVD) to derive a compact native-reference coordinate system. For each L2 utterance, we compute the corresponding averages and project them into the native-reference space. We then demonstrate that the distances between L2 and native-reference coordinates for matched phone-classes show consistent negative correlations with holistic speaking proficiency on the Dev subset of the Speak and Improve Corpus 2025 (Spearman's ρ=−0.53) and with pronunciation quality on the learner subset of the English Read by Japanese Students dataset (ρ=−0.34). These findings suggest that the proposed geometry captures acoustic-phonetic information relevant for proficiency rating while remaining applicable to spontaneous L2 speech without matched native recordings.
Figures & tables
Source
Feature
S&I
UME-ERJ
Train
Dev
Raw
# Words
0.55
0.57
−0.20
Speaking rate
0.54
0.52
−0.13
P⋆
Diphones
0.60
0.62
−0.23
Triphones
0.60
0.62
−0.22
Table 1 : Non-acoustic baselines. We report Spearman correlation ρ for S&I Train, S&I Dev, and UME-ERJ learner using LBS-4h native-reference corpus. Higher positive correlations indicate better agreement with proficiency level. P⋆ denotes the matched context-dependent phone-classes between native and L2 coordinates in our proposed method.
Source
Feature
S&I
UME-ERJ
Train
Dev
Silence
VAD
−0.36
−0.41
−0.10
Alignment
−0.60
−0.57
−0.07
ASR
WER
−0.22
−0.20
−0.24
Raw
Mean
N/A
−0.20
Encoder
DTW
−0.50
Table 2 : Acoustic baselines for similar evaluations and datasets as in Table 1 . Here, lower negative correlations indicate better agreement with proficiency. The silence ratio is calculated by considering both voice activity detection (VAD) and silence segments from the alignment. Encoder averaging and DTW apply only to UME-ERJ , which provides parallel learner–native recordings.
Class
Avg.
Random Boundary Shift [ms]
0
20
40
60
80
120
Diphone
Center
−0.46
−0.45
−0.42
−0.36
−0.30
−0.17
Full
−0.51
−0.50
−0.47
−0.41
−0.33
−0.19
Triphone
Center
−0.55
−0.55
−0.52
−0.48
−0.42
−0.29
Full
−0.56
−0.56
−0.54
−0.51
−0.46
−0.34
Table 3 : Effect of random boundary shift in milliseconds (ms) for different phone-classes and averaging strategies (Avg.) using the LBS-4h native corpus. Values are Spearman ρ of the cosine distance on the S&I Train, averaged over three random seeds.
Native-Ref
S&I Train
UME-ERJ
Corpus
Euc
Maha
Cosine
Euc
Maha
Cosine
TIMIT
−0.44
−0.43
−0.51
−0.25
−0.26
−0.13
LBS-4h
−0.52
−0.51
−0.56
−0.27
−0.28
−0.13
UME-ERJ
−0.35
−0.32
−0.48
−
Table 4 : Spearman ρ for three native-reference (Native-Ref) corpora evaluated on S&I Train and UME-ERJ learner subset using three distance metrics: Euclidean (Euc), Mahalanobis (Maha), and cosine distance. UME-ERJ native is held out for testing, see Table 6 .
L2 Task
Distance
Spearman ρ
Measure
Full
Rnd.
SVD
S&I Train
Euc
−0.50
−0.46
−0.52
Maha
−0.44
−0.46
−0.51
Cosine
−0.58
−0.53
−0.56
UME-ERJ
Euc
−0.28
−0.26
−0.27
Maha
−0.29
−0.26
−0.28
Table 5 : Spearman’s ρ between distances and S&I Train proficiency scores or UME-ERJ learner pronunciation scores. We report Euclidean (Euc), Mahalanobis (Maha), and cosine distance metrics under three native-reference coordinate constructions: full (no projection), random (Rnd.) and SVD projections, both of rank 32 . All experiments use triphone classes with full-unit averaging. More negative values indicate better agreement with proficiency. The best correlation in each row is highlighted.
L2 Task
Native-Ref
\boldmath∣P⋆∣
Correlation
Corpus
min
max
r
ρ
Raw
Partial
Conf. Int.
min
max
S&I Dev
LBS-4h
9
394
−0.53
−0.53
−0.19
−0.56
−0.49
UME-ERJ
UME-ERJ
2
54
−0.37
−0.34
−0.27
−0.37
−0.31
Table 6 : Final evaluation on S&I Dev and UME-ERJ learner using LBS-4h and UME-ERJ native-reference (Native-Ref) corpora. We report the range of matched phone-classes ∣P⋆∣ , Pearson’s r , raw and partial Spearman’s ρ , for ∣P⋆∣ and silence ratio as control variables, and the 95% bootstrap confidence interval (Conf. Int.) for raw ρ . More negative values indicate stronger agreement with proficiency.
Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for second-language (L2) pronunciation assessment in spontaneous speech. We propose a native-reference coordinate geometry in which phone-class averages from native speech define a low-dimensional reference subspace, and L2 speech is evaluated by its distance to matching native phone-class coordinates. Unlike prior distance-based approaches, our method does not require parallel recordings with matched linguistic content or dedicated pronunciation labels. Across different self-supervised encoders and modeling choices, the resulting native-reference distances show negative Spearman correlations up to -0.5 with speaking proficiency, indicating that higher-proficiency speakers tend to lie closer to the native-reference space.
Tina Raissi, Nhan Phan, Mikko Kurimo
Department of Information and Communications Engineering, Aalto University, Espoo, Finland
L2 speech assessment has traditionally focused on phonetic assessment, leaving the scoring of suprasegmental features such as rhythm and intonation underexplored. Moreover, assessment methods often require training with labeled L2 speech data, making them difficult to apply in low-resource settings. We investigate whether DTW over self-supervised WavLM representations can provide a text-free framework for assessing phonetic accuracy, rhythm, and intonation in English and Japanese L2 speech. Results show that a basic DTW-based approach that compares learner speech to native templates exceeds human agreement on holistic and sentence-level phonetic scoring. For rhythm, we introduce methods that measure the degree of warping in the DTW alignment path; our best method approaches human-level performance. For intonation, we combine DTW distance over prosodic residuals with pitch and intensity features, but performance remains more modest on some tasks. Our results point to self-supervised representations as a promising, text-free basis for multi-aspect pronunciation assessment.
Stephen McIntosh, Reuben Smit, Daisuke Saito +2
The University of Tokyo, Japan · Stellenbosch University, South Africa
Training automated pronunciation assessment often relies on labeled learner errors or non-native corpora that are costly to collect. We propose a lightweight framework trained only on native speech resources, operating unsupervised or lightly calibrated with a small set of scored utterances. At inference, learner speech is discretized with an SSL encoder and a K-means codebook. A token language model trained on native sequences computes surprisal where higher surprisal indicates phonotactic deviation. We add a transcript-guided Text2DUnit--DTW module that predicts native token sequences from reference text and aligns them to acoustic tokens to derive error-sensitive features. Surprisal and alignment features are fused via simple regression. On SpeechOcean762, PCC improves from 0.60 to 0.66 with transcript guidance, near supervised baselines. Cross-dataset evaluation on L2-ARCTIC shows consistent gains.