Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams. We formalize this challenge as Word-Level Text Unmixing: given an interleaved lexical stream and source count K, recover the original source sequences while preserving every word occurrence and its within-source order exactly. Directly generating separated texts with LLMs can omit, duplicate, or hallucinate words, violating this exact-reconstruction objective. We therefore propose Evidence-Preserving Ownership Routing (EPOR), which decouples source-ownership prediction from reconstruction. EPOR adapts a causal LLM to predict canonical ownership routes conditioned on the mixed stream and prior routing decisions. At inference, completion-safe constrained decoding is combined with deterministic indexed reconstruction, yielding structurally valid K-source partitions that preserve every observed occurrence exactly once. We also introduce UNMIXBENCH, covering controlled synthetic mixtures, timestamp-derived speech from AMI and ICSI, layout-derived document streams from ReadingBank, and simulated concurrent digital outputs. Across five evaluation tracks, a 4B EPOR model achieves the lowest mean minimum-permutation word error rate among finetuned baselines, reducing the five-track mean by 22.3% relative to compact source-array generation and remaining competitive with zero-shot frontier LLMs. These results show that when lexical evidence is fully observed, separating ownership inference from lexical regeneration provides a reliable alternative to direct generation.
Figures & tables
Figure 1: The Word-Level Text Unmixing problem. Temporal, spatial, and concurrent serialization motivate recovery from an interleaved lexical stream. Given the observed words and known source count K , recover the latent sources while preserving every occurrence and its within-source order.
Figure 2: Evidence-Preserving Ownership Routing ( EPoR ). Training: the causal LLM learns canonical ownership from the complete mixed input and gold prefixes. Inference: completion-safe beam search scores candidate ownership routes, and then indexed reconstruction gathers the original occurrences into sources.
Track
Corpus
Examples / groups
K
T
Interleaving
Controlled
MultiNLI
2,640 / 240
2–4
16/24/32
Controlled intensity
Temporal
AMI
1,000 / 59
2/3
64
Word timestamps
Temporal
ICSI dev
400 / 15
2/3
64
Word timestamps
Spatial
ReadingBank
833 / 828
2–4
32/64
Reading-flow metadata
Concurrent
WikiText
1,200 / 400
2–4
32/64
Virtual event times
Table 1: The UnMixBench evaluation tracks. Counts give examples / provenance groups. K is the source count and T the number of observed word occurrences.
MNLI
AMI
ICSI-dev
RB
DIG
System
MP ↓
Own. ↑
Ev. ↑
MP ↓
Own. ↑
Ev. ↑
MP ↓
Own. ↑
Ev. ↑
MP ↓
Own. ↑
Ev. ↑
MP ↓
Own. ↑
Ev. ↑
Generative baselines
Compact source-array generation (4B)
0.6696
–
0.0845
.7275
–
.0423
.7439
–
.0400
.8415
–
.0240
.7310
–
.0464
Evidence-explicit source-array generation (4B)
0.6747
–
0.0814
.7418
–
.0343
.7530
–
.0317
.8547
–
.0272
.7314
–
.0544
Turn-marker generation (4B)
0.4271
0.6991
0.3957
0.6928
0.4441
0.6377
0.6690
0.4805
0.6600
0.6862
0.4959
0.3325
0.6449
0.5426
0.7769
Structured baselines
Table 2: Main results across five evaluation tracks on UnMixBench . Each block reports MP-WER (MP), permutation-aligned ownership accuracy (Own.), and evidence preservation (Ev.). Turn-marker ownership uses lexical alignment (Appendix C.2 ). Lower MP-WER is better; higher values are better for the other metrics. Best and second-best results are bold and underlined. Turn-marker uncertainty is given in Table 12 , and Appendix D.5 reports hosted-model settings.
Condition
Adapt.
Legal mask
MP-WER ↓
Evidence ↑
Complete K -way ↑
Ownership acc. ↑
Switch F1 ↑
EPoR
✓
✓
0.3757
1.0000
1.0000
0.7457
0.8514
w/o task adaptation
×
✓
0.6842
1.0000
1.0000
0.5278
0.5333
w/o legality masking
✓
×
0.3819
1.0000
0.9721
0.7429
0.8492
w/o adaptation and legality masking
×
×
1.2008
1.0000
0.1473
0.3830
0.0894
Route-independent fixed- template history
✓
✓
0.4486
1.0000
1.0000
0.7113
0.7603
Representation control
Table 3: EPOR component and representation ablations on MultiNLI. Adapt. denotes task adaptation; Legal mask denotes completion-safe decoding. PI-Route is a matched representation control with noncanonical legal support, rather than a component-removal condition.
Figure 3: Recovery across controlled mixing intensities. The same source groups are evaluated at eleven target values of M . Points average three training runs and 80 source groups per (K,T) condition; shaded regions show source-group bootstrap 95% confidence intervals.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
System
Scoring / state
Trainable head
Loss
Free generation
causal source-text likelihood
none beyond QLoRA
assistant-token LM
Turn-marker generation
causal marked-stream likelihood
none beyond QLoRA
assistant-token LM
Global Transformer
bidirectional contextual emissions
256-d projection; 2-layer, 8-head encoder; linear emission
occurrence CE
EMA prototype
emission plus cosine to per-source EMA ( λ=.75 , scale 4)
256-d projection/query; shared biases
occurrence CE
RGS-CRF
global emissions plus 4×4 transitions
Global head plus transition matrix
exact legal-route CRF NLL
Legal projection
edit cost to generated source arrays; EMA tie guide
none in projection
none
Appendix
Table 4: Implementation summary and trainable parameter counts for the comparison baselines.
Corpus
Role
Examples / groups
Source unit and preprocessing
SNLI 1.0
Train
10,000 / 10,000
Premise and hypothesis spans; deduplication across held-out splits, NFKC normalization, whitespace collapse, and deterministic regex unitization.
SNLI 1.0
Development
1,000 / 1,000
Same preprocessing as training; source/text-disjoint from training and test data.
MultiNLI 1.0
Controlled evaluation
2,640 / 240
Premises only; NFKC normalization, whitespace collapse, and eight whitespace units per source. Prompt and document identifiers define provenance.
MultiNLI 1.0
Paired construction
480 / 240
Same fixed-length preprocessing; identical source word occurrences under controlled and independent-clock mixing.
MultiNLI 1.0
Unequal-length stress
540 / 180
Variable source spans; independent-clock, long-run, and bursty mixing. Prompt and document identifiers define provenance.
WikiText-103
Domain transfer
1,448 / 1,448
Length-filtered sentence spans materialized with the deterministic regex unitizer.
Appendix
Table 5: Corpora, roles, and preprocessing used in the experiments.
Split
Corpus
Source unitization
Examples / groups
K
T /length
Interleaving
Disjointness
Seed
Train
SNLI
NFKC + regex words
10,000 / 10,000
2–4
variable
controlled IID
held-out source/text
2027
Development
SNLI
NFKC + regex words
1,000 / 1,000
2–4
variable
controlled IID
train/test source/text
2027
Controlled evaluation
MultiNLI
first 8 whitespace units/source
2,640 / 240
2–4
16/24/32
11 controlled intensities
source + document
2701
Paired construction
MultiNLI
first 8 whitespace units/source
480 / 240
2–4
16/24/32
controlled + renewal clock
source + document
13013
Unequal-length stress
MultiNLI
variable source spans
540 / 180
2–4
variable
clock + long-run + bursty
source + document
20270905
Appendix
Table 6: Core UnMixBench specification. All held-out groups are provenance-disjoint from training/development; the paired and stress sets are also disjoint from the controlled MultiNLI evaluation.
Track
Serialization mechanism
Examples
Groups
K counts
T counts
Mean M
AMI
annotated word-time merge
1,000
59
2:667; 3:333
64:1,000
0.275
ICSI dev
annotated/repaired word-time merge
400
15
2:200; 3:200
64:400
0.290
ReadingBank
metadata-derived reading flows
833
828
2:300; 3:300; 4:233
32:383; 64:450
0.176
Digital–smooth
virtual event-time merge
400
400
2:134; 3:133; 4:133
32:201; 64:199
0.582
Digital–bursty
virtual event-time merge
400
400
2:134; 3:133; 4:133
32:201; 64:199
0.266
Digital–unequal-rate
virtual event-time merge
400
400
2:134; 3:133; 4:133
32:201; 64:199
0.531
Appendix
Table 7: Serialization-transfer extension. Temporal and spatial tracks use real metadata; digital profiles simulate a concurrent-production mechanism with virtual event-time schedules and reuse the same source groups. Counts describe frozen evaluation views before model inference.
Condition
Beam
Score normalization
Generated length
EOS
Codec validation
Adapted, constrained (EPOR)
4
LM beam score
exactly T
excluded as route action
atomic; reject unsupported
Frozen, constrained
4
LM beam score
exactly T
excluded as route action
atomic; reject unsupported
Adapted, unconstrained
4
LM beam score
exactly T
suppressed before T
retain as invalid ( −1 )
Frozen, unconstrained
4
LM beam score
exactly T
suppressed before T
retain as invalid ( −1 )
Appendix
Table 8: Search settings for the adaptation-by-support ablations. “LM beam” uses the decoder’s original beam-score convention; no legal-subset renormalization is added.
Method
MP-WER ↓
Ownership acc. ↑
Switch F1 ↑
Evidence ↑
Compact source-array generation
0.6696 ± 0.0608
–
–
0.0845
Evidence-explicit source-array generation
0.6747 ± 0.0242
–
–
0.0814
Turn-marker generation
0.4271 ± 0.0096
0.6991
0.8081
0.3957
Global Transformer
0.5368 ± 0.0192
0.6658
0.6847
1.0000
EMA prototype
0.5165 ± 0.0140
0.6703
0.7390
1.0000
RGS-CRF
0.4799 ± 0.0063
0.6917
0.7001
1.0000
Appendix
Table 9: Full results on the controlled MultiNLI UnMixBench test set. MP-WER is mean ± sample SD over three training runs; remaining entries are three-run means, with corresponding SDs provided in the supplementary results.
Prompt
Parse
Limit hit
Conditional MP-WER
Conditional Evidence
Compact
0.6972 ± 0.1147
n/a
0.5272 ± 0.0108
0.1239 ± 0.0259
Evidence-explicit
0.6852 ± 0.0274
0.0034 ± 0.0023
0.5256 ± 0.0225
0.1189 ± 0.0029
Appendix
Table 10: Generation-pipeline diagnostics, mean ± sample SD over three runs. Conditional metrics use each method’s successfully parsed subset.
Metric
Turn-marker generation
EPoR − Turn-marker
Turn-marker − source-array generation
MP-WER ↓
0.4271 ± 0.0096
−0.0515[−0.0636,−0.0395]
−0.2476[−0.2720,−0.2238]
Ownership acc. ↑
0.6991 ± 0.0068
+0.0466[0.0381,0.0554]
–
Switch F1 ↑
0.8081 ± 0.0032
+0.0432[0.0338,0.0530]
–
Evidence ↑
0.3957 ± 0.0131
+0.6043[0.5617,0.6461]
+0.3143[0.2778,0.3524]
Appendix
Table 11: Task-matched turn-marker generation on the 2,640-example MultiNLI evaluation. The turn-marker column is mean ± sample SD over three training runs. Difference columns give paired point estimates with 95% source-group bootstrap intervals.
Track
EPoR MP
Turn-marker MP
Paired difference [95% CI]
Ownership acc.
Evidence
MultiNLI
0.3757
0.4271 ± 0.0096
−0.0515[−0.0636,−0.0395]
0.7457 / 0.6991
1.0000 / 0.3957
AMI
0.6320
0.6928 ± 0.0312
−0.0608[−0.0724,−0.0481]
0.6205 / 0.4441
1.0000 / 0.6377
ICSI development
0.6161
0.6690 ± 0.0185
−0.0529[−0.0704,−0.0353]
0.6228 / 0.4805
1.0000 / 0.6600
ReadingBank
0.6156
0.6862 ± 0.0154
−0.0706[−0.0859,−0.0556]
0.6224 / 0.4959
1.0000 / 0.3325
Digital
0.6474
0.6449 ± 0.0088
+0.0026[−0.0067,0.0120]
0.5912 / 0.5426
1.0000 / 0.7769
Appendix
Table 12: EPOR and turn-marker generation across five tracks. Turn-marker MP-WER is mean ± sample SD over three runs. Paired differences are EPoR minus turn-marker, with 95% group-bootstrap intervals; negative values favor EPoR . Ownership accuracy and Evidence columns give EPoR / turn-marker means.
Track
Ownership difference [95% CI]
Switch F1: EPoR / turn-marker
Switch F1 difference [95% CI]
MultiNLI
+0.0466 [0.0381, 0.0554]
0.8514 / 0.8081
+0.0432 [0.0338, 0.0530]
AMI
+0.1765 [0.1632, 0.1891]
0.2567 / 0.4469
−0.1902 [-0.2090, -0.1714]
ICSI development
+0.1423 [0.1258, 0.1572]
0.3018 / 0.5168
−0.2150 [-0.2379, -0.1885]
ReadingBank
+0.1264 [0.1121, 0.1412]
0.4459 / 0.4161
+0.0298 [0.0148, 0.0449]
Digital
+0.0486 [0.0406, 0.0568]
0.6764 / 0.7136
−0.0372 [-0.0500, -0.0246]
Appendix
Table 13: Ownership and boundary recovery after lexical alignment. Differences are EPoR minus turn-marker, with paired 95% group-bootstrap intervals. Positive differences favor EPoR . All examples, including strict parsing failures, contribute to the reported scores.
Table row / scope
Provider
Requested model ID
Thinking
Temp.
Completion-token budget
Qwen3.8-Flash / all
Alibaba Cloud Model Studio
qwen3.8-flash
off
0
max(8192,64T)
Qwen3.8-Max / all
Alibaba Cloud Model Studio
qwen3.8-max
off
0
max(8192,64T)
DeepSeek-V4-Flash / MNLI
DeepSeek API
deepseek-v4-flash
off
0
max(256,16T)
DeepSeek-V4-Flash / AMI, ICSI-dev, RB, DIG
Alibaba Cloud Model Studio
deepseek-v4-flash
off
0
max(8192,64T)
DeepSeek-V4-Pro / MNLI
DeepSeek API
deepseek-v4-pro
off
0
max(256,16T)
DeepSeek-V4-Pro / AMI, ICSI-dev, RB, DIG
Alibaba Cloud Model Studio
deepseek-v4-pro
off
0
max(8192,64T)
Appendix
Table 14: Hosted-model execution configuration for Table 2 . “Scope” identifies the cells supplied by each run: MNLI, AMI, ICSI-dev, ReadingBank (RB), or the three-profile digital aggregate (DIG).
System
Temp.
JSON parse ↑
Limit hit ↓
Refusal ↓
DeepSeek-V4.1-Flash
0
0.9284
0.0000
0.0000
DeepSeek-V4-Pro
0
0.8811
0.0000
0.0000
Qwen3.8-Flash
0
0.8023
0.0000
0.0004
Qwen3.8-Max
0
0.9826
0.0000
0.0000
Appendix
Table 15: MultiNLI compliance diagnostics for four hosted evaluations under the generous max(8192,64T) budget. Metrics include all examples. Parse failures and provider refusals receive no repair.
Condition
MP-WER ↓
Ownership acc. ↑
Switch F1 ↑
Evidence ↑
Complete K -way ↑
Unmodified history
0.3729
0.7457
0.8514
1.0000
1.0000
Route-independent template
0.9126
0.4722
0.5057
1.0000
1.0000
25% eligible-label rotation
0.4761
0.6702
0.7763
1.0000
1.0000
50% eligible-label rotation
0.5335
0.6263
0.7080
1.0000
1.0000
100% eligible-label rotation
0.5941
0.5796
0.5692
1.0000
1.0000
Appendix
Table 16: Inference-time ownership-history interventions on MultiNLI. The checkpoint is fixed; only the prefix shown to the scorer changes. The controller uses the actual candidate prefix, leaving output legality intact.
Figure 4: Sensitivity to model-visible ownership history. Points report one adapted model; bands are 95% source-group bootstrap intervals over 240 groups, retaining all eleven interleavings within each group. Corruption begins only after all K labels have appeared; the decoder’s true legal prefix is unchanged.
Representation
MP-WER ↓
LM train sequences
Rel. train time ↓
Canonical ( EPoR )
0.3757
60,000
1.00 ×
Arbitrary labels (PI-Route)
0.3782
602,608
4.08 ×
Appendix
Table 17: Matched representation comparison. Recovery values are three-run means; both models have 64,929,792 trainable parameters.
Method
MP-WER ↓
Ownership acc. ↑
Switch F1 ↑
Evidence ↑
Structure ↑
EPoR
0.3757 ± 0.0056
0.7457 ± 0.0039
0.8514 ± 0.0022
1.0000
1.0000
PI-Route
0.3782 ± 0.0012
0.7438 ± 0.0037
0.8521 ± 0.0023
1.0000
1.0000
Appendix
Table 18: Full matched representation comparison on MultiNLI, mean ± sample SD over three training runs.
Similarity band
MP-WER ↓
Ownership acc. ↑
Switch F1 ↑
Low
0.3419 [0.3227, 0.3619]
0.7747 [0.7610, 0.7882]
0.8657 [0.8515, 0.8790]
Medium
0.3775 [0.3513, 0.4035]
0.7494 [0.7317, 0.7673]
0.8451 [0.8259, 0.8641]
High
0.3888 [0.3665, 0.4115]
0.7323 [0.7151, 0.7491]
0.8487 [0.8358, 0.8613]
Extreme
0.3944 [0.3730, 0.4163]
0.7266 [0.7100, 0.7428]
0.8461 [0.8298, 0.8616]
Appendix
Table 19: Recovery across matched source-similarity strata. Entries are matched means with 95% source-group bootstrap intervals. Lower MP-WER is better; higher values are better otherwise.
Construction
Mixing M
Switches
Mean run
Run var.
Ambiguity
Controlled
0.4953
12.4542
2.6454
0.4096
0.1208
Independent clock
0.5651
14.4542
1.7185
1.0254
0.1583
Appendix
Table 20: Route statistics for paired controlled and independent-clock interleavings.
Backbone
Output
Controlled MP-WER ↓
Clock MP-WER ↓
Evidence C/I ↑
Qwen3.5-4B
EPoR
0.4163
0.5189
1.0000/1.0000
RGS-CRF
0.5344
0.6075
1.0000/1.0000
Free generation
0.6012
0.6480
0.1708/0.1792
Qwen3.5-9B
EPoR
0.3625
0.4459
1.0000/1.0000
Free generation
0.5368
0.5886
0.2125/0.2333
Gemma-7B-IT
EPoR
0.4630
0.5488
1.0000/1.0000
Appendix
Table 21: Full cross-backbone comparison on 240 paired source groups. Each backbone uses one training run. Evidence reports controlled/clock; route methods have structural validity 1.0 in both conditions.
Method
Clock MP-WER
Long-run MP-WER
Bursty MP-WER
Min. evidence ↑
EPoR
0.500
0.288
0.469
1.000
RGS-CRF
0.568
0.347
0.612
1.000
Task-finetuned source-array generation
0.623
0.658
0.592
0.139
Random legal route
0.726
0.775
0.697
1.000
Equal contiguous blocks
0.683
0.278
0.759
1.000
Gold-length contiguous blocks †
0.668
0.000
0.740
1.000
Appendix
Table 22: Unequal-length construction comparison on 180 paired source groups. † The final row receives gold source lengths and is a privileged reference rather than a deployable baseline.
Domain
RGS-CRF MP-WER
EPoR MP-WER
Δ MP-WER [95% CI]
WikiText
0.3583 ± 0.0081
0.2506 ± 0.0136
-0.1077 [-0.1174, -0.0982]
AG News
0.6187 ± 0.0220
0.7322 ± 0.0228
0.1135 [0.1056, 0.1213]
AMI
0.6483 ± 0.0103
0.6320 ± 0.0086
-0.0163 [-0.0257, -0.0062]
ICSI reserve
0.6581 ± 0.0036
0.6283 ± 0.0057
-0.0297 [-0.0457, -0.0113]
Appendix
Table 23: Domain transfer, mean ± SD over three training runs. WikiText and AG News intervals resample examples; AMI and ICSI-reserve intervals resample meetings. AMI is the same frozen 1,000-window view used in the temporal-serialization analysis; ICSI here denotes the separate 300-window consumed reserve. Δ is EPoR minus RGS-CRF.
View
Free-c
Free-e
Global
EMA
RGS-CRF
Projection
EPoR
Δ to best alternative [95% CI]
AMI
0.7275
0.7418
0.6406
0.6596
0.6483
0.6431
0.6320
−0.0086 [ −0.0175 , +0.0003 ]
ICSI development
0.7439
0.7530
0.6389
0.6674
0.6495
0.6483
0.6161
−0.0228 [ −0.0377 , −0.0078 ]
ReadingBank
0.8415
0.8547
0.6820
0.7427
0.6555
0.6843
0.6156
−0.0399 [ −0.0517 , −0.0281 ]
Digital–smooth
0.7193
0.7283
0.7718
0.7326
0.7467
0.6741
0.6750
+0.0009 [ −0.0128 , +0.0152 ]
Digital–bursty
0.7431
0.7307
0.6763
0.6794
0.6406
0.6764
0.6156
−0.0251 [ −0.0409 , −0.0090 ]
Digital–unequal-rate
0.7306
0.7352
0.7482
0.7114
0.7223
0.6788
0.6517
−0.0270 [ −0.0413 , −0.0130 ]
Appendix
Table 24: Source-array and structured baselines by serialization view. Entries are mean MP-WER over three training seeds. “Free-c” and “Free-e” denote compact and evidence-explicit task-finetuned generation. Δ compares EPoR with the lowest-mean source-array or structured baseline shown in that view; 95% CIs resample provenance groups after averaging seeds. Negative differences favor EPoR . Best and second-best means are bold and underlined.
Figure 5: Candidate quality, selection, and cost. Selected and candidate-oracle MP-WER are shown across legal beam sizes, together with gold-route coverage, runtime, and peak allocated memory. Points average three trained models; intervals resample source groups.
Unlearning aims to remove the influence of specific training data sources, but this has proved challenging because the contributions of different sources are entangled within the model. Isolating source contributions to disjoint parameters makes removal easier, though it obstructs joint learning across sources. We propose NULLs (Natively Unlearnable LLMs), a model class that satisfies the two opposing goals of isolating source-specific contributions and learning jointly across sources, by training a set of shared backbone neurons alongside a pool of sparsely activated sinks. During training, information specific to a source naturally concentrates in its sinks while information shared across sources accumulates in the backbone. A source is then unlearned at deployment by disabling its corresponding sinks, with no gradient updates and no access to the retained data. We show that NULLs scales to Wikipedia's ~6M articles, isolating each as an independent source. Unlearning a single article removes knowledge specific to it while preserving facts shared with semantically related articles, closely matching retraining from scratch. We note that unlearning with NULLs is also robust: in a case study of unlearning the Harry Potter books, NULLs resists both adversarial extraction and relearning that reverses post-hoc unlearning. Finally, NULLs preserves general language capabilities, matching a standard transformer on downstream benchmarks. Together, these results suggest that source-level unlearning need not be an afterthought. It can be built natively into LLM training while retaining the benefits of shared representation learning.
Gaurav R. Ghosal, Pratyush Maini, Aditi Raghunathan
We introduce TextSeal, a state-of-the-art watermark for large language models. Building on Gumbel-max sampling, TextSeal introduces dual-key generation to restore output diversity, along with entropy-weighted scoring and multi-region localization for improved detection. It supports serving optimizations such as speculative decoding and multi-token prediction, and does not add any inference overhead. TextSeal strictly dominates baselines like SynthID-text in detection strength and is robust to dilution, maintaining confident localized detection even in heavily mixed human/AI documents. The scheme is theoretically distortion-free, and evaluation across reasoning benchmarks confirms that it preserves downstream performance; while a multilingual human evaluation (6000 A/B comparisons, 5 languages) shows no perceptible quality difference. Beyond its use for provenance detection, TextSeal is also ``radioactive'': its watermark signal transfers through model distillation, enabling detection of unauthorized use.
Speech language models (SLMs) have been extensively studied, with the common paradigm incorporating text data and pre-trained text LMs. A leading approach is speech-text interleaving in which models are trained over sequences containing both speech and text tokens, aiming to boost even speech-only capabilities. Yet the way these two modalities interact in the model latent space remains unclear. In this work, we analyze interleaved speech-text LMs from different model families and sizes through the scope of the logit lens to provide such insight. We reveal that these models go through an implicit transcription phase in which the text token of the spoken word becomes decodable in intermediate layers, despite not being trained for speech recognition. The transcription of the word appears as one of the top candidate words for as much as 77% of the data. Following this stage, the models proceed to predict the next word in the text space before transforming back to the speech domain. We finally analyze the role of interleaving data, and initializing from text LMs in eliciting this behavior, as well as seeing how this correlates with spoken knowledge abilities. Our analysis sheds light on the internal mechanisms underlying the relationship between speech and text modalities and could shape SLM optimization.