Cross-linguistic effects are a central topic in bilingual first-language acquisition. Artificial learners can help investigate L1-L2 interactions by enabling controlled comparisons across language combinations and learning conditions. Recent work explores this direction by training bilingual language models under developmentally plausible constraints. However, human and model learners still diverge in fundamental ways, with one major difference being input modality: children learn primarily from spoken input, whereas language models are typically trained on orthographic text. To reduce this gap, researchers have trained models on phonemic representations of speech. In this work, we combine these research directions to train bilingual BabyLMs with phonemic input. We keep English fixed as the L2 and vary the L1 across German, Swedish, Persian, and Basque, selected to represent contrasting combinations of syntactic and phoneme-inventory distance from English. Our results show stronger L1-related variation in grammatical learning trajectories under phonemic than orthographic input, while early lexical differences align with phoneme-inventory similarity.
Figures & tables
Figure 1: We train matched bilingual orthographic and phoneme LMs with four different L1s and English as the L2. We investigate L1-related differences in English grammatical performance with minimal-pair evaluations. Here, we illustrate the phoneme Syntactic and Lexical metrics from BabySLM Lavechin et al. (2023) .
Figure 2: Syntactic and phoneme-inventory distances from English for the eligible language set, computed with URIEL+ Khan et al. (2025) . Languages selected for training the bilingual models are marked with ◊ . We opt for languages that occupy contrasting regions of the syntactic and phoneme-inventory distance space.
Figure 3: English evaluation accuracy as a function of L1–English syntactic distance for phoneme (top) and orthographic (bottom) LMs. For each model we report an average across 3 seeds with SD error bars. Pearson correlation ( r ) over averages shown for each benchmark. BabySLM Lexical is available only in phonemic form.
mBLiMP
BabySLM
L1
BLiMP
L1
eng
Lexical
Syntactic
Orthographic
eng
69.1
–
83.4
–
88.2
deu
66.2
84.7
78.8
–
83.8
eus
63.2
87.2
76.8
–
79.7
pes
65.2
84.6
78.8
–
83.1
swe
65.4
–
77.5
–
83.5
Table 1: Minimal-pair accuracy for the English monolingual references and bilingual LMs trained on orthographic or phonemic input. Results are averaged across 3 different seeds. BLiMP and MultiBLiMP (mBLiMP) are evaluated in their original orthographic form for the orthographic LMs and in phonemic form for the phoneme LMs. BabySLM Lexical is available only for the phoneme condition.
Figure 4: Evaluation accuracy over training for the bilingual Phoneme LMs . The first vertical dotted line marks the onset of joint L1–English training; horizontal dotted lines indicate chance performance. Lines correspond to average performance across 3 random seeds with ± 1 SD shading. For the first 1,000 steps after L2 introduction (10k–11k) we save logarithmically spaced checkpoints, highlighted with violet.
Figure 5: Evaluation accuracy over training for the bilingual Orthographic LMs .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Training Steps
Phase one (L1)
10k
Phase two (L1+L2)
20k
Monolingual
30k
Embedding params
(orthographic / phoneme)
23M / 500k
Non-embedding params
85M
Appendix
Table 2: Training hyperparameters.
Figure 6: Data category distribution across languages.
Figure 7: BLiMP accuracy per category for Orthographic LMs throughout training. Dashed lines indicate L2 onset and random change accuracy (50%). Lines correspond to average performance across 3 random seeds with ± 1 SD shading. Plots in each row are grouped according to the broader linguistic field: Morphology (first row), Syntax (second row), Syntax/Semantics (third row left), Semantics (third row right).
Figure 8: BLiMP accuracy per category for Phoneme LMs throughout training.
Recent studies suggest that child-directed speech is not conducive to language learning in BabyLMs. However, current evaluations focus predominantly on comprehension and not production, which is central to usage-based theories of language acquisition which argue how CDS facilitates early language use through constructional ''frames'' (frequent lexical patterns with open slots). We introduce a novel generation-based evaluation inspired by such theories in form of a frame-completion task, and compare Llama models trained with CDS, the BabyLM corpus, and web-crawl data (FineWeb-edu) on comprehension benchmarks and our novel framework. Our results reveal a clear dissociation between models' comprehension and production capabilities: while FineWeb-trained models excel at minimal pairs, CDS-trained models produce grammatical completions substantially earlier in training and concentrate probability mass on appropriate slot-fillers. These findings show that comprehension benchmarks underestimate what CDS affords to BabyLMs.
Bastian Bunzeck, Sina Zarrieß
Computational Linguistics, Department of Linguistics Bielefeld University, Germany
Large language models (LLMs) have increasingly been used to investigate how children acquire syntax at an early stage of development. This is notably the central scientific goal of the BabyLM challenge, a community-wide effort to develop models that achieve human-level syntactic performance while being trained on developmentally realistic corpora. In this paper, we reflect on the use of LLMs in the study of infant syntax learning by providing an epistemological assessment of several studies from this research program. We discuss how datasets are built, which models are implemented, how they are trained and syntactically evaluated. We observe significant assumptions in the methodology of BabyLM and related studies, thus mitigating their theoretical scope. We additionally observe that using developmentally-realistic corpora have limited effects on models performance on commonly-used benchmarks, which suggest important computational differences between LLMs and the infant syntax learner.
Hélie Bazin, Anouk Barberousse, François Yvon
SCAI, SND, ISIR · Sorbonne Université, Sorbonne Center for Artificial Intelligence (SCAI) · SND +3
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on some measures, closes this multilingual gap on continuous phonetic and higher-level linguistic measures, while preserving substantial cross-language sharing. Using a controlled English/French HuBERT setting, we test two interventions which strengthen language discrimination: an auxiliary language classifier and per-language k-means targets. Across interventions, continuous-feature phone discrimination error (phone-ABX, lower is better) decreases from 11.6% in the bilingual baseline to 10.4% (monolingual: 10.8%), while lexical performance (sWUGGY, higher is better) increases from 52.1% to 56.7% (monolingual: 58.5%) and prosodic performance (ProsAudit, lexical subtask, higher is better) from 68.9% to 72.9% (monolingual: 72.6%). Across HuBERT training stages, the strongest gains on most linguistic measures occur when language discrimination is introduced in the first iteration, whereas later or repeated interventions yield smaller improvements and are accompanied by increased language-wise segregation. These results support a causal role for language discrimination in reducing the additional cost of multilingual learning.