Phoneme-level models transcribe or generate speech as a sequence of phonemes, the smallest sound units that distinguish words. These models enable fine-grained pronunciation control and understanding, yet often fail on input that does not match any single training language, such as speech alternating between two languages, known as code-switching, or low-resource languages absent from training. We identify a systematic failure mode behind this, phonological interference: models assume the input is in a single language and impose its phonology, overriding local phoneme-level decisions that conflict with the assumed language. We measure interference by how often a model retains phonemes that one language has but the other lacks. On code-switched input, two phone recognizers (speech-to-phoneme models) and a phoneme-conditioned text-to-speech model lose 32% to 79% of these phonemes, but lose far fewer of the phonemes both languages share. On unseen languages, we find that phone recognizers impose the phonology of the training language they assign to the speech, and the more confident the assignment, the more they lose phonemes the unseen language has but the assigned language lacks. We probe the models' language estimate from their internal activations, and trace interference to a low dimensional subspace. On monolingual speech, steering this subspace toward another language makes the model lose the phonemes that only the original language uses and produce phonemes that only the target language has. We introduce windowed language estimation (WLE), an inference time repair that replaces the model's language estimate in this subspace with one computed from a short window around each position. On code-switched input, WLE removes 34% to 69% of the interference in all three models, and in the recognizers it leaves monolingual performance essentially unchanged.
Figures & tables
Figure 1: Phonological interference. Multilingual phoneme-level speech models estimate the language from the whole input and override individual phoneme predictions. (a) In a French–English code-switched utterance, a phone recognizer replaces English [r] in right with French [K]. The bar shows the model’s language estimate in a low dimensional language subspace at each position, with time running left to right. The model computes this estimate from the whole utterance, so it blends French (orange) into English (blue), and the model still reads the English span partly as French. (b) Windowed language estimation (WLE) replaces the language subspace coordinates at each position with those the model computes from a short window of audio around it. The estimate is then computed from the local language context, and [r] is correctly transcribed.
Figure 2: Qualitative examples of interference. PhoneticXeus transcriptions aligned to the reference, with unshared phonemes in bold. Output at unshared phonemes is shaded blue when the phoneme is kept, orange with a frame and bold type when it is replaced by the other language’s phoneme, and grey for any other error ( ∅ marks a deletion). sep transcribes the second span without the other language in its context, and tog transcribes it with the other language in its context. In the French–English example, English “the score” takes French [d] and [K]. The unseen language rows (Section 2.4 ) have no sep / tog conditions, so we list the model-assigned language instead.
Figure 3: Phonological interference on code-switched input. Each bar runs from recall with each span processed separately ( sep , hollow) to recall with both languages together ( tog , filled); its length is the interference Δ . Red: unshared phonemes. Grey: shared phonemes. Unshared phoneme recall drops much more than shared phoneme recall in every model and dataset (Table 5 ).
Figure 4: Phonological interference in unseen languages. Recall on 43 DoReCo languages, split by the model’s language probe confidence: lowest third (hollow) vs. highest third (filled). Unshared phoneme recall (red) drops the more the model is confident in the estimated language, while shared phonemes (grey) remain stable, isolating interference from general difficulty of unseen languages.
Figure 5: The language estimate drops on code-switched input, and editing it reproduces interference. (a) Probe score on each span’s own language, processed separately ( sep , hollow) vs. together ( tog , filled), with the conditions of Section 2.2 . (b) Recall on monolingual utterances before (hollow) and after (filled) the language subspace is set to the mean of another language B . Red: unshared phonemes. Grey: shared phonemes. Δ is the drop in each panel.
Model
Condition
rmin\textsctog↑
Δ (interference) ↓
TTS
No intervention
0.410
0.302
WLE, fixed window
0.513
0.199
WLE , span window
0.619
0.093
Oracle
0.606
0.106
PhoneticXeus
No intervention
0.138
0.507
WLE
0.430
0.214
Table 1: Effect of WLE on interference. Best result in bold . WLE removes 69% of the interference in TTS with the span window and 34% with the fixed window, 58% in PhoneticXeus, and 36% in POWSM. The Oracle rows (gray) use the true language of every frame or span and serve as a reference (Appendix J ). POWSM rows use its CTC head. rmin\textsctog and Δ as in Section 2.2 .
Model
Benchmark
Δ PER (pp) ↓
Δ PFER (pp) ↓
Sliding window
WLE
Sliding window
WLE
Seen
POWSM
TIMIT
+0.96
+0.07
+0.47
+0.01
L2-ARCTIC
+1.55
+0.10
+0.31
−0.00
GMU accent
+2.14
−0.32
+0.73
−0.07
PhoneticXeus
TIMIT
+1.33
+0.08
+0.49
+0.03
L2-ARCTIC
+3.22
+0.14
+0.50
+0.03
Table 2: Performance in monolingual speech. Change in phone error rate ( Δ PER) and phone feature error rate ( Δ PFER), in percentage points, on monolingual benchmarks relative to unmodified decoding. Negative = improvement. The lower (better) value in each pair is in bold . The sliding window degrades PER and PFER much more than WLE, except for POWSM’s PFER on DoReCo.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Language pair
Replacement
Example
Language pair
Replacement
Example
English–French
→
K
r eally
German–Spanish
I
→
i
m i t
English–German
w
→
v
w as
English–German
D
→
d
th e
English–Spanish
A
→
o
pr o bably
English–Spanish
z
→
s
i s
German–Hindi
K
→
R
de r
English–Italian
æ
→
a
a nd
Appendix
Table 3: Examples of unshared phonemes from language pairs in our experiments. Each phoneme at the left of an arrow occurs only in the first language of the pair; the arrow shows its most common replacement when both languages are processed together, with an example word whose bold letters spell the phoneme. Table 4 lists the full sets for every pair and dataset.
Pair (A–B)
Only in A
Only in B
SwitchLingua
English–French
r 54 D 33 30 æ 29 2 21 h 18
K 141 y 45 œ 11 ø 4
English–German
r 53 w 28 æ 28 D 28 18 T 8 Z 6
K 124 ç 44 5 35 ø 13 x 11 y 10 œ 5 Y 5
English–Hindi
f 35 D 33 æ 31 27 2 24 w 21 A 17 v 15 T 7
R 61 V 19 k h 18 b h 12 t h 11 S h 9 p h 8 t 7 d 6 Z h 3
English–Italian
I 105 r 54 @ 38 U 30 2 27 D 18 N 15 æ 11 A 11 11 T 10 h 6
r 173 L 5
English–Korean
@ 63 r 62 U 38 35 f 30 D 27 æ 26 v 17 O 14 A 12 Z 5 T 3
W 119 R 77 57 C 53 47
Appendix
Table 4: The full unshared phoneme sets of every language pair in both code-switched datasets. Since selected phonemes depend on an occurrence threshold in observed data, the exact phonemes selected may differ between the datasets for the same pair of languages (e.g., comparing English–German in both datasets). Phonemes are ordered from most to least frequent, and the subscript is the number of occurrences in those reference spans.
Unshared phonemes
Shared phonemes
Model
Δ
Δ/r\textscsep
Δexact
Δ
Δ/r\textscsep
Ratio
TTS (CosyVoice)
.302 [.265, .340]
.424
.224 [.186, .262]
.077 [.065, .090]
.105
2.9 ×
Phone rec. (PhoneticXeus)
.507 [.462, .551]
.786
.345 [.304, .387]
.167 [.145, .188]
.244
2.1 ×
Phone rec. (POWSM, CTC)
.448 [.404, .490]
.754
.365 [.327, .402]
.136 [.110, .161]
.210
2.7 ×
Phone rec. (POWSM, attn.)
.414 [.371, .456]
.769
.268 [.233, .304]
.139 [.116, .164]
.202
1.9 ×
Appendix
Table 5: Phonological interference on code-switched input (full numbers behind Figure 3), with 95% bootstrap intervals. Δ : drop in worst span recall from sep to tog . Δ/rmin\textscsep : drop as a fraction of the sep level. Δexact : the same drop under exact match, the rule used for shared phonemes; unshared Δ also accepts substitutions by phonemes the other language lacks. Ratio: unshared Δexact divided by shared Δ .
rmin
Δ from
Model
Alone
Same lang.
Other lang.
Alone
Same lang.
TTS (CosyVoice)
.570
.571
.318
.252 [.217, .287]
.253 [.222, .284]
Phone rec. (PhoneticXeus)
.591
.582
.396
.196 [.165, .228]
.186 [.156, .218]
Phone rec. (POWSM, CTC)
.636
.639
.180
.456 [.421, .491]
.459 [.425, .494]
Phone rec. (POWSM, attn.)
.598
.610
.165
.433 [.394, .472]
.445 [.408, .482]
Appendix
Table 6: Two sep baselines on Spliced FLEURS. Worst span recall of unshared phonemes with each half alone, next to a same language half ( sep in the main text) and next to the other language half ( tog ), and the drop to tog from each baseline, with 95% bootstrap intervals.
Setting
Model
γ (log-odds)
Odds ratio eγ
Spliced FLEURS
POWSM, CTC head
0.77 [0.61, 0.97]
2.2 [1.8, 2.6]
POWSM, attn. decoder
0.84 [0.66, 1.04]
2.3 [1.9, 2.8]
PhoneticXeus
0.27 [0.10, 0.47]
1.3 [1.1, 1.6]
TTS (CosyVoice)
0.10 [ − 0.02, 0.22]
1.1 [0.98, 1.3]
Edit toward B
POWSM, CTC head
2.28 [2.07, 2.53]
9.8 [7.9, 12.6]
PhoneticXeus
2.44 [2.31, 2.64]
11.5 [10.1, 14.1]
Appendix
Table 7: The same phoneme next to languages that lack it and languages that have it. γ : additional drop in the log-odds of recalling a phoneme when the other language lacks it, with the phoneme and the language pair held fixed; 0 if a phoneme’s difficulty alone decided the drop. 95% bootstrap intervals. The TTS interval on Spliced FLEURS includes 0.
Dataset
Model
Added substitutions
Most similar phoneme
Share
Spliced FLEURS
POWSM, CTC head
2.64 [2.42, 2.87]
0.68
0.29
POWSM, attn. decoder
2.86 [2.61, 3.10]
0.68
0.25
PhoneticXeus
1.78 [1.46, 2.11]
0.59
0.42
TTS (CosyVoice)
2.47 [2.18, 2.75]
0.79
0.35
SwitchLingua
POWSM, attn. decoder
3.36 [3.03, 3.72]
0.77
0.19
PhoneticXeus
2.76 [2.45, 3.10]
0.70
0.29
Appendix
Table 8: Replacements are far from the phonemes they replace. Added substitutions: mean PanPhon feature distance between an unshared phoneme and the other language’s phoneme that replaces it, over the replacements the other language adds (95% bootstrap intervals). Most similar phoneme: the same mean if each had been replaced by the closest phoneme of the other language. Share: fraction of replacements that are that closest phoneme.
Figure 6: Language probe accuracy by layer , on held-out monolingual speech. Held-out speakers : 5,800 utterances from 1,112 unseen speakers. Held-out speakers and corpus : each point is from a probe trained without that language’s corpus, so neither the speakers nor the recording pipeline have been seen. The vertical line marks the first layer within 0.01 of the best. For the TTS, the probe reads the language model at each speech token position, with the utterance’s own speech tokens as input. We read language estimates at layer 8 (POWSM), layer 17 (PhoneticXeus) and block 7 of the TTS language model.
Figure 7: Subspace edit sweeps. Left: effect of varying the start block (rank 15, development speakers). Middle: effect of varying rank (at the selected start block, test speakers). Right: probe accuracy for the true language inside the rank- k subspace (chance dotted). Solid: cumulative edit from start block to last; dashed: single block only. Blue: unwhitened; orange: whitened. See legend for additional markers.
Δunshared
Δshared
Δsh/Δunsh
Language subspace →B
.777 [.757, .791]
.229 [.220, .236]
.29
Language subspace →C (against C )
.777 [.756, .792]
—
—
One direction ( A→B )
.753 [.730, .767]
.228 [.220, .235]
.30
Erased to global mean
.417 [.399, .429]
.109 [.103, .114]
.26
Random shift (seed 1)
.029 [.025, .032]
.006 [.004, .009]
—
Random shift (seed 2)
.030 [.025, .035]
.006 [.005, .008]
—
Appendix
Table 9: Manipulating the language subspace reproduces interference (full numbers behind Figure 5(b) ). Δ : drop in recall from unedited to edited output, paired per utterance. Δsh/Δunsh : selectivity ratio; lower means the edit targets unshared phonemes more precisely. CIs are 95% bootstrap over speaker groups within language. The →C row gives the drop in recall of A ’s phonemes unshared with C ; the other two columns are defined against B and are left empty for it, and for the TTS model it uses 1,515 utterances with one decode each. Edited transcripts retain 96–98% of unedited length in all conditions.
SwitchLingua
Spliced FLEURS
WER, sep
.160
.152
WER, tog
.259
.540
Δ WER
+.099 [.071, .126]
+.388 [.356, .423]
Δ WER, span in the picked language
−.010 [ − .027, .006]
+.045 [.007, .097]
Δ WER, the other span
+.223 [.168, .282]
+.741 [.695, .784]
Appendix
Table 10: Whisper loses the span that is not in the language it picks. WER rises when the other language is present. The span in the picked language is transcribed about as well as it is alone, and almost all of the rise comes from the other span, which Whisper translates or drops. WER is the number of word errors divided by the number of reference words, over both spans. Δ is tog minus sep on the same recordings, with 95% bootstrap intervals over recordings. As in Section 2.2 , sep is the span alone on SwitchLingua, and the span next to another speaker of the same language on Spliced FLEURS. The bottom rows use the 274 recordings and 286 clips where Whisper picks one of the two languages.
Edit
ΔP(B)
Δ transcript in B
ΔP(A)
Toward B
+.201 [.184, .212]
+.221 [.195, .238]
−.447 [ − .465, − .427]
Toward C
+.007 [.005, .008]
+.002 [.000, .003]
−.449 [ − .468, − .429]
Mean of all centroids
+.008 [.006, .010]
+.003 [.001, .005]
−.276 [ − .298, − .261]
Random subspace
+.001 [.000, .001]
+.000 [.000, .000]
−.013 [ − .019, − .010]
Appendix
Table 11: Whisper transcribes in the language its encoder is steered toward. Every edit of the language subspace lowers the probability of A , and only the edit toward B raises the probability of B and moves the transcript to B . A is the language of the utterance, B is the target language and C is a third language. P(A) and P(B) are the probabilities Whisper gives to the language tokens of A and B . “Transcript in B ” is the text language identifier’s confidence that the transcript is written in B . Each entry is the change from the unedited model, where P(A) is 0.963 . Entries are averaged within each source language and then over the 14 languages, with 95% bootstrap intervals over speaker groups. The edit toward C uses the 2,027 utterances whose C has a language token.
Model
Condition
rmin\textsctog↑
Δ (interference) ↓
TTS
No intervention
.410
.302 [.265, .340]
In chunks, no edit
.351
.361 [.328, .394]
wle , fixed window
.513
.199 [.175, .223]
wle , fixed window, random subspace
.332
.380 [.350, .411]
Span by span, no edit
.383
.329 [.291, .368]
wle , span window
.619
.093 [.068, .120]
Appendix
Table 12: Interference on SwitchLingua with 95% bootstrap intervals over recordings. In all three models the random subspace leaves interference at the level of the same decoding without the edit.
Model
Benchmark
Sliding window
wle
wle , random subspace
Δ PER (pp)
POWSM
TIMIT
+0.96 [ +0.85 , +1.07 ]
+0.07 [ +0.02 , +0.11 ]
+0.01 [ −0.02 , +0.04 ]
L2-ARCTIC
+1.55 [ +1.37 , +1.74 ]
+0.10 [ +0.01 , +0.19 ]
+0.02 [ −0.02 , +0.07 ]
GMU accent
+2.14 [ +2.00 , +2.29 ]
−0.32 [ −0.39 , −0.26 ]
−0.04 [ −0.07 , −0.01 ]
DoReCo
+1.97 [ +1.77 , +2.19 ]
−0.09 [ −0.19 , +0.02 ]
−0.13 [ −0.18 , −0.09 ]
PhoneticXeus
TIMIT
+1.33 [ +1.22 , +1.44 ]
+0.08 [ +0.05 , +0.11 ]
+0.01 [ −0.01 , +0.03 ]
Appendix
Table 13: Δ PER (top) and Δ PFER (bottom), in percentage points, behind Table 2 , with 95% bootstrap intervals over clips.
Model
Word
Reference
sep
tog
WLE
(a) English words in a French–English sentence: Le Président américain a annoncé une décision importante aujourd’hui. Can you believe he revoked Biden’s security clearance?
TTS
revoked
IvoUkt
ivoUkt
K ivOk
voUkt
security
sIkjUIti
sIkjUIdi
sikju K iti
sIkjUIdi
PhoneticXeus
revoked
IvoUkt
IvoUkt
K ivOkt
ivoUkt
security
sIkjUIti
sIkjU@ti
sk yK iti
skjUiti
POWSM
revoked
IvoUkt
IvoUktu
K @vOkt
@vokt
Appendix
Table 14: Examples of WLE. Each model’s transcript of the words in bold when it processes each span alone ( sep ), the whole utterance ( tog ), and the whole utterance with WLE. For the TTS model, WLE uses the span window, and the transcript is PhoneticXeus’s reading of the synthesized span (first sampling seed). In the reference, the unshared phonemes of the word’s own language are in bold. In the transcripts, phonemes that only the other language uses are in red .
Condition
1 s
2 s
3 s
4 s
PhoneticXeus, SwitchLingua (no intervention: .507)
Sliding window
.223 [.188, .260]
.202 [.160, .243]
.207 [.163, .250]
.246 [.203, .290]
wle
.185 [.146, .222]
.214 [.171, .257]
.259 [.216, .304]
.306 [.263, .349]
POWSM, SwitchLingua (no intervention: .448)
Sliding window
.245 [.203, .287]
.153 [.120, .187]
.184 [.151, .218]
.211 [.171, .251]
wle
.321 [.282, .359]
.288 [.249, .327]
.294 [.255, .333]
.330 [.289, .371]
Appendix
Table 15: Interference Δ ( ↓ ) by window width, with 95% bootstrap intervals over recordings or clips. The unrepaired Δ is in parentheses, and the lowest value in each row is shown in bold .
SwitchLingua
Spliced FLEURS
Basis
PhoneticXeus
POWSM
PhoneticXeus
POWSM
No repair, Δ
.507
.448
.186
.459
Sliding window (2 s), Δ
.202
.153
.037
.081
Share of the window’s reduction
All 16 languages, rank 15 ( wle )
.96
.54
.93
.60
7 code-switched languages, rank 6
.88
.52
.92
.55
Appendix
Table 16: Share of the sliding window’s interference reduction recovered by wle with each basis. Each value is (Δno repair−Δ\textscwle)/(Δno repair−Δwindow) . The full 16-language basis recovers most of the window’s effect. Bases missing one language of the evaluated pair show partial repair; on SwitchLingua, whose pairs all contain English, bases missing both languages recover a quarter to a third of the window’s reduction in PhoneticXeus and no more than a random basis in POWSM.
SwitchLingua
Spliced FLEURS
Condition
PhoneticXeus
POWSM
PhoneticXeus
POWSM
No intervention
.507 [.462, .551]
.448 [.404, .490]
.186 [.156, .218]
.459 [.425, .494]
Sliding window (2 s)
.202 [.160, .243]
.153 [.120, .187]
.037 [.016, .059]
.081 [.060, .102]
wle
.214 [.171, .257]
.288 [.249, .327]
.047 [.023, .073]
.234 [.205, .264]
wle , random subspace
.494 [.449, .539]
.451 [.407, .494]
.178 [.148, .211]
.440 [.406, .475]
Oracle
.108 [.073, .144]
.121 [.080, .162]
.028 [.005, .050]
.143 [.118, .169]
Appendix
Table 17: Interference Δ (lower is better) with 95% intervals. The oracle reduces interference beyond wle on all four cells, and beyond the sliding window on three. The same edits through a random subspace have no effect, confirming the reduction depends on the language subspace.
Language tag
SwitchLingua
Spliced FLEURS
<unk>
.414 [.371, .456]
.445 [.408, .482]
Each span’s own language
.102 [.070, .134]
.248 [.213, .284]
Difference
.312 [.263, .360]
.197 [.160, .235]
Appendix
Table 18: POWSM’s attention decoder with each span’s own language tag. Interference Δ (lower is better), the drop in worst span recall of unshared phonemes from sep to tog , with 95% bootstrap intervals, on the recordings (266) and clips (274) that both decodes can score. Difference: Δ under <unk> minus Δ under the span’s own tag, paired per recording or clip.
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing. Using a controlled English/French HuBERT setting, we test two interventions which strengthen language discrimination: an auxiliary language classifier and per-language k-means targets. Across interventions, continuous-feature phone discrimination error (phone-ABX, lower is better) decreases from 11.6% in the bilingual baseline to 10.4% (monolingual: 10.8%), while lexical performance (sWUGGY, higher is better) increases from 52.1% to 56.7% (monolingual: 58.5%) and prosodic performance (ProsAudit, lexical subtask, higher is better) from 68.9% to 72.9% (monolingual: 72.6%). Across HuBERT training stages, the strongest gains on most linguistic measures occur when language discrimination is introduced in the first iteration, whereas later or repeated interventions yield smaller improvements and are accompanied by increased language-wise segregation. These results support a causal role for language discrimination in reducing the additional cost of multilingual learning.
Phone recognition (PR) serves as the atomic interface for language-agnostic modeling for cross-lingual speech processing and phonetic analysis. Despite prolonged efforts in developing PR systems, current evaluations only measure surface-level transcription accuracy. We introduce PRiSM, the first open-source benchmark designed to expose blind spots in phonetic perception through intrinsic and extrinsic evaluation of PR systems. PRiSM standardizes transcription-based evaluation and assesses downstream utility in clinical, educational, and multilingual settings with transcription and representation probes. We find that diverse language exposure during training is key to PR performance, encoder-CTC models are the most stable, and specialized PR models still outperform Large Audio Language Models. PRiSM releases code, recipes, and datasets to move the field toward multilingual speech models with robust phonetic ability: https://github.com/changelinglab/prism.
Non-speech interference can change a speech representation without causing comparable task loss. We test eight frozen encoders on four tasks, adding non-speech sounds throughout recordings, during speech, or in pauses. Under whole-recording interference, embedding drift tracks task loss across seven sounds, with mean Spearman correlations of 0.81-0.88. Moving the same sound between speech and pauses changes this pattern. At quiet to moderate levels, pause interference produces larger drift, while speech interference usually causes greater loss on intent recognition, speaker verification and speech recognition. Emotion recognition shows a weaker placement effect. Pause interference also changes speech-frame representations beyond the injected region. Even below the estimated recording background, interference can change embeddings as much as repeated speech takes do. Drift helps rank the effects of different sounds, but larger drift does not consistently indicate greater task loss.
Vsevolod Kovalev, Pranay Manocha
Boston University, Boston, MA, USA · Princeton University, Princeton, NJ, USA