We introduce BranchShine-CR, a 25M-parameter model for multilingual transcription into the International Phonetic Alphabet (IPA). It combines log-mel features, a rotary-position E-Branchformer encoder, intermediate self-conditioned connectionist temporal classification (CTC), and consistency regularization across augmented views. On 16,646 shared IPApack++ test utterances, it achieves 4.47% IPA character error rate, a 22.3% relative reduction from ZIPA-CTC-NS, with approximately one-twelfth as many parameters while being trained from scratch. BranchShine-CR also outperforms a similarly sized NeMo Conformer baseline across all 41 dataset language labels. Ablation studies indicate the individual components synergetically acting in model performance contribution. These findings support compact IPA recognition capabilities under limited compute budget, for applications in low-resource on-device pronunciation assessment.
Figures & tables
Figure 1: BranchShine-CR architecture. Shared (W,b) produce the logits; shared A projects posteriors into the + nodes during training and inference. The inset uses two views of network Fθ : C combines final and auxiliary CTC; D is symmetric stopped-target KL.
Component
Original
CR
Front end
Raw waveform
80-bin log-mel
Encoder blocks
19
12
Hidden size / heads
288 / 8
256 / 4
Feed-forward width
480
1,024
Intermediate CTC layers
None
6, 10
Vocabulary size
112
112
Table 1: Configurations of the two BranchShine systems.
Model
Params (M)
IPA-CER (%)
Exact match (%)
PFER (%)
BranchShine-CR
25.39
4.47
41.37
2.09
ZIPA-CTC-NS
299.97
5.76
45.15
2.14
ZIPA-CTC
299.97
6.51
38.43
2.48
Original BranchShine
33.38
7.13
26.03
3.07
NeMo Conformer-CTC Medium
30.53
8.67
19.55
3.84
PhoneticXEUS
575.00
9.76
20.19
3.20
Table 2: IPApack++ results on the same 16,646 test utterances with identical references and normalization. Lower is better for IPA-CER and PFER ; higher is better for exact match. Training regimes differ.
Variant
Seed 17
Seed 29
Mean
Δ (pp)
A0: Full CR
16.87
16.83
16.85
—
A1: No consistency, two views
18.92
18.88
18.90
+2.05
A2: Single view, no consistency
20.70
20.74
20.72
+3.87
A3: No auxiliary supervision
17.23
17.80
17.51
+0.66
A4: No prediction feedback
16.66
16.48
16.57
-0.28
A5: No intermediate CTC
17.41
17.48
17.44
+0.59
Table 3: Ablation studies: final-step development CER (%) at 18,000 updates. Δ is the mean paired difference from A0 in percentage points where positive is worse. All variants use seeds 17 and 29, the dash denotes the baseline comparison.