We introduce BranchShine-CR, a 25M-parameter model for multilingual transcription into the International Phonetic Alphabet (IPA). It combines log-mel features, a rotary-position E-Branchformer encoder, intermediate self-conditioned connectionist temporal classification (CTC), and consistency regularization across augmented views. On 16,646 shared IPApack++ test utterances, it achieves 4.47% IPA character error rate, a 22.3% relative reduction from ZIPA-CTC-NS, with approximately one-twelfth as many parameters while being trained from scratch. BranchShine-CR also outperforms a similarly sized NeMo Conformer baseline across all 41 dataset language labels. Ablation studies indicate the individual components synergetically acting in model performance contribution. These findings support compact IPA recognition capabilities under limited compute budget, for applications in low-resource on-device pronunciation assessment.
Figures & tables
Figure 1: BranchShine-CR architecture. Shared (W,b) produce the logits; shared A projects posteriors into the + nodes during training and inference. The inset uses two views of network Fθ : C combines final and auxiliary CTC; D is symmetric stopped-target KL.
Component
Original
CR
Front end
Raw waveform
80-bin log-mel
Encoder blocks
19
12
Hidden size / heads
288 / 8
256 / 4
Feed-forward width
480
1,024
Intermediate CTC layers
None
6, 10
Vocabulary size
112
112
Table 1: Configurations of the two BranchShine systems.
Model
Params (M)
IPA-CER (%)
Exact match (%)
PFER (%)
BranchShine-CR
25.39
4.47
41.37
2.09
ZIPA-CTC-NS
299.97
5.76
45.15
2.14
ZIPA-CTC
299.97
6.51
38.43
2.48
Original BranchShine
33.38
7.13
26.03
3.07
NeMo Conformer-CTC Medium
30.53
8.67
19.55
3.84
PhoneticXEUS
575.00
9.76
20.19
3.20
Table 2: IPApack++ results on the same 16,646 test utterances with identical references and normalization. Lower is better for IPA-CER and PFER ; higher is better for exact match. Training regimes differ.
Variant
Seed 17
Seed 29
Mean
Δ (pp)
A0: Full CR
16.87
16.83
16.85
—
A1: No consistency, two views
18.92
18.88
18.90
+2.05
A2: Single view, no consistency
20.70
20.74
20.72
+3.87
A3: No auxiliary supervision
17.23
17.80
17.51
+0.66
A4: No prediction feedback
16.66
16.48
16.57
-0.28
A5: No intermediate CTC
17.41
17.48
17.44
+0.59
Table 3: Ablation studies: final-step development CER (%) at 18,000 updates. Δ is the mean paired difference from A0 in percentage points where positive is worse. All variants use seeds 17 and 29, the dash denotes the baseline comparison.
Speech-to-IPA transcription is useful when the desired output is pronunciation rather than orthographic text, but competitive multilingual systems are often large and evaluation is sensitive to normalization choices. This paper presents BranchShine, a 33M-parameter raw-audio CTC recognizer with a lightweight convolutional front end and a 19-block RoPE E-Branchformer encoder. We find that BranchShine provides a compact and competitive operating point for IPA transcription under matched normalization and scoring. On a 16,660-utterance multilingual test set covering 41 language labels, BranchShine obtains 9.19% whitespace-insensitive IPA character error rate, compared with 9.78% for the 575.00M-parameter PhoneticXEUS baseline. A secondary child speech reading analysis shows a complementary operating profile: BranchShine is more conservative on incorrect readings, while Whisper-Medium is stronger on exact acceptance of correct readings. Overall, the results indicate that a compact raw-audio-to-IPA model can approach much larger baselines on character-level IPA transcription.
Multilingual language models often exhibit performance disparities across languages that can arise as early as the tokenization stage. Widely-used subword tokenization approaches favor high-resource languages, and tokenizer-free methods still yield longer sequences for scripts with a higher bytes-per-character ratio. To address these shortcomings, we propose to use the International Phonetic Alphabet (IPA) as a language-agnostic input representation for multilingual tokenizers. IPA provides a compact symbol inventory, greater cross-lingual character overlap, and a more balanced byte-per-character distribution across languages. We train matched pairs of text vs. IPA subword tokenizers across 24 languages and 14 scripts and demonstrate that IPA tokenizers consistently improve tokenization quality, especially for non-Latin scripts, and generalize more effectively to unseen languages and scripts.
Recent Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription rely on Grapheme-to-Phoneme (G2P) labels, but the phoneme labels are not necessarily phonetically faithful. To investigate this issue, we evaluate zero-shot phonetic classification on Chinese aspiration and Japanese moraic nasals. A PFM trained on G2P-labeled data excluding these two languages yields poor accuracy on both tasks, showing that multilingual coverage with discrete IPA tokens is not sufficient for unseen settings. To overcome this limitation, we propose a classification method based on continuous Articulatory Feature (AF) vectors extracted from each frame. This AF-based approach outperforms discrete token-based methods, particularly for rare phones. We further show that it is crucial to adopt the optimal temporal aggregation of AF vectors for the target distinction: single-frame classification is best for aspiration, while segmental classification substantially improves nasal classification.
Ryo Magoshi, Jaeyoung Lee, Shinsuke Sakai +1
Graduate School of Informatics, Kyoto University, Japan · NTT, Inc., Japan