Automatic speaking assessment systems can provide holistic proficiency scores, but often lack interpretable measures that characterize pronunciation quality. We propose a native-reference phone-class geometry for measuring second language (L2) pronunciation deviation without requiring pronunciation labels, read-aloud prompts, or matched recordings of the same text from native and L2 speakers. Given a native speech corpus, we average frame-level self-supervised representations for each context-dependent phone-class and use singular value decomposition (SVD) to derive a compact native-reference coordinate system. For each L2 utterance, we compute the corresponding averages and project them into the native-reference space. We then demonstrate that the distances between L2 and native-reference coordinates for matched phone-classes show consistent negative correlations with holistic speaking proficiency on the Dev subset of the Speak and Improve Corpus 2025 (Spearman's ρ=−0.53) and with pronunciation quality on the learner subset of the English Read by Japanese Students dataset (ρ=−0.34). These findings suggest that the proposed geometry captures acoustic-phonetic information relevant for proficiency rating while remaining applicable to spontaneous L2 speech without matched native recordings.
Figures & tables
Source
Feature
S&I
UME-ERJ
Train
Dev
Raw
# Words
0.55
0.57
−0.20
Speaking rate
0.54
0.52
−0.13
P⋆
Diphones
0.60
0.62
−0.23
Triphones
0.60
0.62
−0.22
Table 1 : Non-acoustic baselines. We report Spearman correlation ρ for S&I Train, S&I Dev, and UME-ERJ learner using LBS-4h native-reference corpus. Higher positive correlations indicate better agreement with proficiency level. P⋆ denotes the matched context-dependent phone-classes between native and L2 coordinates in our proposed method.
Source
Feature
S&I
UME-ERJ
Train
Dev
Silence
VAD
−0.36
−0.41
−0.10
Alignment
−0.60
−0.57
−0.07
ASR
WER
−0.22
−0.20
−0.24
Raw
Mean
N/A
−0.20
Encoder
DTW
−0.50
Table 2 : Acoustic baselines for similar evaluations and datasets as in Table 1 . Here, lower negative correlations indicate better agreement with proficiency. The silence ratio is calculated by considering both voice activity detection (VAD) and silence segments from the alignment. Encoder averaging and DTW apply only to UME-ERJ , which provides parallel learner–native recordings.
Class
Avg.
Random Boundary Shift [ms]
0
20
40
60
80
120
Diphone
Center
−0.46
−0.45
−0.42
−0.36
−0.30
−0.17
Full
−0.51
−0.50
−0.47
−0.41
−0.33
−0.19
Triphone
Center
−0.55
−0.55
−0.52
−0.48
−0.42
−0.29
Full
−0.56
−0.56
−0.54
−0.51
−0.46
−0.34
Table 3 : Effect of random boundary shift in milliseconds (ms) for different phone-classes and averaging strategies (Avg.) using the LBS-4h native corpus. Values are Spearman ρ of the cosine distance on the S&I Train, averaged over three random seeds.
Native-Ref
S&I Train
UME-ERJ
Corpus
Euc
Maha
Cosine
Euc
Maha
Cosine
TIMIT
−0.44
−0.43
−0.51
−0.25
−0.26
−0.13
LBS-4h
−0.52
−0.51
−0.56
−0.27
−0.28
−0.13
UME-ERJ
−0.35
−0.32
−0.48
−
Table 4 : Spearman ρ for three native-reference (Native-Ref) corpora evaluated on S&I Train and UME-ERJ learner subset using three distance metrics: Euclidean (Euc), Mahalanobis (Maha), and cosine distance. UME-ERJ native is held out for testing, see Table 6 .
L2 Task
Distance
Spearman ρ
Measure
Full
Rnd.
SVD
S&I Train
Euc
−0.50
−0.46
−0.52
Maha
−0.44
−0.46
−0.51
Cosine
−0.58
−0.53
−0.56
UME-ERJ
Euc
−0.28
−0.26
−0.27
Maha
−0.29
−0.26
−0.28
Table 5 : Spearman’s ρ between distances and S&I Train proficiency scores or UME-ERJ learner pronunciation scores. We report Euclidean (Euc), Mahalanobis (Maha), and cosine distance metrics under three native-reference coordinate constructions: full (no projection), random (Rnd.) and SVD projections, both of rank 32 . All experiments use triphone classes with full-unit averaging. More negative values indicate better agreement with proficiency. The best correlation in each row is highlighted.
L2 Task
Native-Ref
\boldmath∣P⋆∣
Correlation
Corpus
min
max
r
ρ
Raw
Partial
Conf. Int.
min
max
S&I Dev
LBS-4h
9
394
−0.53
−0.53
−0.19
−0.56
−0.49
UME-ERJ
UME-ERJ
2
54
−0.37
−0.34
−0.27
−0.37
−0.31
Table 6 : Final evaluation on S&I Dev and UME-ERJ learner using LBS-4h and UME-ERJ native-reference (Native-Ref) corpora. We report the range of matched phone-classes ∣P⋆∣ , Pearson’s r , raw and partial Spearman’s ρ , for ∣P⋆∣ and silence ratio as control variables, and the 95% bootstrap confidence interval (Conf. Int.) for raw ρ . More negative values indicate stronger agreement with proficiency.