cs.CLSep 24, 2026

A Native-Reference Phone-Class Geometry for Second-Language Pronunciation Analysis

Authors: Tina Raissi, Nhan Phan, Chenxiao Wang, Mikko Kurimo

Organizations: Department of Information and Communications Engineering, Aalto University, Espoo, Finland

Abstract

Automatic speaking assessment systems can provide holistic proficiency scores, but often lack interpretable measures that characterize pronunciation quality. We propose a native-reference phone-class geometry for measuring second language (L2) pronunciation deviation without requiring pronunciation labels, read-aloud prompts, or matched recordings of the same text from native and L2 speakers. Given a native speech corpus, we average frame-level self-supervised representations for each context-dependent phone-class and use singular value decomposition (SVD) to derive a compact native-reference coordinate system. For each L2 utterance, we compute the corresponding averages and project them into the native-reference space. We then demonstrate that the distances between L2 and native-reference coordinates for matched phone-classes show consistent negative correlations with holistic speaking proficiency on the Dev subset of the Speak and Improve Corpus 2025 (Spearman's ρ ⁣= ⁣−0.53ρ\!=\!-0.53) and with pronunciation quality on the learner subset of the English Read by Japanese Students dataset (ρ ⁣= ⁣−0.34ρ\!=\!-0.34). These findings suggest that the proposed geometry captures acoustic-phonetic information relevant for proficiency rating while remaining applicable to spontaneous L2 speech without matched native recordings.

Figures & tables

Explore similar work

CardsList
  1. A Native-Reference Coordinate Geometry for L2 Pronunciation Deviation Using Self-Supervised Speech Models

    Sep 23, 2026Tina Raissi, Nhan Phan, Mikko KurimoPronunciation AssessmentGrapheme-To-Phoneme

  2. Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring

    Jul 15, 2026Stephen McIntosh, Reuben Smit, Daisuke Saito +2Grapheme-To-PhonemeProsody

  3. Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

    Jun 18, 2026Syeda Faiza Ahmed Sara, Shammur Absar ChowdhuryPronunciation AssessmentSurprisal