Pre-trained language models (PLMs) with encoder-based architectures have shown impressive capabilities in zero-shot cross-lingual transfer for various language understanding tasks. However, applying this technique to dependency parsing remains a significant challenge due to its syntactic nature. To boost model generalizability across linguistic typologies, we propose a cross-lingual unsupervised bootstrapping method to improve syntactic knowledge within the PLM. We show that our method achieves a significant improvement in zero-shot parsing performance in low-resource languages. Analysis of these bootstrapped models uncovers increased robustness in recognizing syntactic structures, evidenced by higher scores in parameter-free tree probing tests.
Figures & tables
Figure 1: Layer-wise CKA similarity between original and modified English sentences obtained from pre-trained mBERT representations. A higher score means more robustness for word swapping. We use the English EWT test treebank (UD v2.8) in this experiment.
Figure 2: Original sentence (above) and OV reordered sentence (below). Dashed arcs indicate relations related to OV reordering, while solid arcs indicate relations unrelated to OV reordering.
Model
All (n=28)
Unseen /
Seen from mBERT (grouped by pre-training data size in MB)
Avg
Avg
(0, 50]
(50, 100]
(200, )
Avg
UDapter (Reproduced)
51.29/36.65
42.01/27.52
67.49/51.09
65.73/49.75
75.99/62.63
70.88/55.92
UDapter-B-mBERT L12
-0.82 † /-0.48 †
-0.79 † /-0.33
-2.07 † /-1.52
-0.89 † /-1.83 †
+0.05/+0.21
-0.86 † /-0.82 †
with MLM
+0.16/ +0.28 †
-0.31 † /-0.18
+2.75/+3.49
+0.69/+0.35
+0.22/-0.00 †
+1.16/+1.24
with SynAug OV
-0.30 † /+0.16
-1.05 † /-0.38
+3.70 † /+3.63 †
-1.41 † /-1.59
+0.78/+0.98
+1.27 / +1.29 †
with SynAug OV + MLM
+0.16 † /+0.25 †
-0.03/ +0.06 †
+2.63 † /+2.15 †
-1.44/-0.94 †
+0.01/+0.31
+0.56/+0.65 †
Table 1: Zero-shot evaluation results on 28 target languages grouped by inclusion in mBERT pre-training data and the amount of pre-training data (low-resource, middle-resource, and high-resource languages). The scores reported are the average UAS/LAS change compared to UDapter. UDapter-B-mBERT denotes bootstrapped mBERT with two settings: L12 (last layer only) and L12+L6 (last and sixth layers). † indicates a statistically significant result from the McNemar’s test ( p<0.05 ).
Figure 3: Zero-shot LAS parsing performance grouped by the target languages’ canonical verb and object ordering. Results are from two UDapter models: Bootstrapped-mBERT L12 with MLM and VO augmentation, and Bootstrapped-mBERT L12+L6 with MLM loss. * denotes languages included in the mBERT pre-training data.
Treebank
mBERT
Bootstrapped mBERT
(test set)
L12
L12+L6
Unseen
50.85/55.76
-0.25/+0.95
+0.12/+1.05
Seen
53.74/59.57
+0.88/+1.85
+1.08/+1.88
(0, 50]
53.76/58.45
+0.21/+0.92
+1.03/+1.34
(50, 100]
53.55/60.01
+0.53/+0.83
+0.46/+0.77
(200, )
53.81/60.20
+1.56/+3.05
+1.43/+2.83
Table 2: Comparison of average UAS/UUAS obtained from parameter-free tree probing between vanilla mBERT and bootstrapped models trained with MLM ( λ = 0.1) and no sentence augmentation. The scores shown are from the layer with the highest UAS.
Encoder
All
Unseen
Seen
XLM-100
+0.35/+0.34
+0.17/-0.07
+0.78/+1.26
(Wiki)
(+0.26/+0.09)
(+0.36/+0.11)
(+0.06/+0.04)
Distil-mBERT
+0.53 /+0.37
+1.40/+0.80
-0.30/-0.13
(Wiki, Distill )
(+1.07/+1.82)
(+0.80/+1.60)
(+1.65/+2.31)
XLMR
-0.03/-0.39
-0.11/-0.36
-0.04/-0.52
(CC-100)
(-0.90/-1.48)
(-0.27/-0.93)
(-2.05/-2.48)
Table 3: Comparison of average UAS/LAS change and probed UAS/UUAS scores (in parentheses). The results come from a bootstrapped model that uses the last layer representation, MLM auxiliary loss, and VO augmentation. The results are selected based on the highest average UAS for the seen languages.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Sentences with Noun-Adj reordering (a) and Adposition (b) reordering. Dashed arcs indicate relations related to reordering.
Model
All (n=28)
Unseen /
Seen from mBERT (grouped by pre-training data size in MB)
Avg
Avg
(0, 50]
(50, 100]
(200, )
Avg
UDapter (Reproduced)
51.29/36.65
42.01/27.52
67.49/51.09
65.73/49.75
75.99/62.63
70.88/55.92
UDapter-B-mBERT L12
-0.82 † /-0.48 †
-0.79 † /-0.33
-2.07 † /-1.52
-0.89 † /-1.83 †
+0.05/+0.21
-0.86 † /-0.82 †
MLM ( λ=0.1 )
+0.16/+0.28 †
-0.31 † /-0.18
+2.75/+3.49
+0.69/+0.35
+0.22/-0.00 †
+1.16/+1.24
MLM ( λ=0.01 )
-0.38/+0.05
-0.81 † /-0.14
+1.71 † /+2.03 †
+0.21/-1.13
-0.14 † /+0.08 †
+0.55 † /+0.46 †
MLM ( λ=0.001 )
-1.07 † /-0.93 †
-1.50 † /-0.81 †
-0.84 † /-2.22 †
+1.64/-1.03 †
-0.58 † /-0.51 †
-0.17 † /-1.20 †
Appendix
Table 5: Zero-shot evaluation result on 28 target languages grouped by inclusion in mBERT pre-training data and the amount of pre-training data. The scores reported are the average UAS/LAS change compared to UDapter. UDapter-B-mBERT denotes bootstrapped mBERT with two settings: L12 (last layer only) and L12+L6 (last and sixth layers). The λ denotes the weight assigned to the MLM auxiliary loss during the model bootstrapping phase. † indicates a statistically significant result from McNemar’s test ( p<0.05 ). The pre-training data sizes (in MB) for seen languages are as follows: Breton (be), Tagalog (tl), Yoruba (yo) (0-50 MB); Marathi (mr), Welsh (cy) (50-100 MB); Belarusian (be), Kazakh (kk), Tamil (ta), Telugu (te) (>200 MB). Data sizes are calculated from the Wikipedia dump plaintext (snapshot: 2018-02-25).
Model
All (n=28)
Unseen /
Seen from the encoder (grouped by pre-training data size in MB)
Avg
Avg
(0, 50]
(50, 100]
(200, )
Avg
UDapter-Distil-MBert
45.73/30.74
39.31/25.29
50.13/31.34
56.24/37.45
67.65/52.84
59.27/42.25
UDapter XLM-100
49.87/36.18
43.67/29.92
55.25/41.61
62.33/46.76
69.08/56.51
62.97/49.38
UDapter-Distil-MBert L6
+0.30 † /+0.20 †
+1.41 † /+0.82 †
-0.94 † /+0.05
-1.92 † /-2.51 †
-0.41/-0.45
-0.92/-0.67 †
MLM ( λ=0.01 )
+0.74 † /+0.49 †
+1.68 † /+0.62 †
-1.53 † /+0.81 †
+1.18/+0.33
-0.24/-0.02
-0.47 † /+0.38 †
SynAug
+0.58 † /+0.31
+1.40 † /+0.64 †
+0.55/+0.45 †
-0.18 † /-1.62 †
+0.56 † /+0.55 †
+0.40 † /+0.08 †
Appendix
Table 6: Zero-shot evaluation result on 28 target languages grouped by inclusion in mBERT pre-training data and the amount of pre-training data. The scores reported are the average UAS/LAS change compared to UDapter. UDapter-B-Distil-mBERT denotes bootstrapped mBERT. The λ denotes the weight assigned to the MLM auxiliary loss during the model bootstrapping phase. The row marked with "-" indicates that the model failed to converge. † indicates a statistically significant result from McNemar’s test ( p<0.05 ).
Model
All (n=28)
Unseen /
Seen from XLMR (grouped by pre-training data size in GB)
Avg
Avg
(0, 4]
(4, 10]
(10, )
Avg
UDapter-mMiniLM
51.50/36.85
40.52/26.53
56.86/35.80
75.94/61.09
80.01/72.22
68.01/51.62
UDapter-XLMR
53.96/40.20
41.03/27.80
66.71/47.05
78.57/64.03
81.58/73.84
73.73/58.10
UDapter-B-XLMR L12
-1.23/-0.86
-1.67/-0.98
-0.64/-0.57
-0.40/-0.61
-0.62/-0.45
-0.55/-0.56
MLM ( λ=0.01 )
+0.14/-0.27
+0.10/-0.38
-0.36/-0.58
+0.29/+0.30
+0.29/+0.26
-0.01/-0.11
SynAug
-0.44/-0.48
-0.18/-0.21
-0.69/-0.94
-0.48/-1.31
+0.48/+0.40
-0.40/-0.83
Appendix
Table 7: Zero-shot evaluation result on 28 target languages grouped by inclusion in mBERT pre-training data and the amount of pre-training data. The scores reported are the average UAS/LAS change compared to UDapter. UDapter-B-XLMR denotes bootstrapped XLMR. UDapter-B-mMiniLM denotes bootstrapped Multilingual MiniLM. The λ denotes the weight assigned to the MLM auxiliary loss during the model bootstrapping phase. † indicates a statistically significant result from McNemar’s test ( p<0.05 ).
UDapter
UDapter-Bootstrapped mBERT
(Baseline)
Single-Layer (L12)
Multi-Layer (L12+L6)
MLM
SynAug
MLM
MLM
SynAug
MLM
+SynAug
+SynAug
akk
42.38/49.11
-0.12/+0.40
-1.04/-0.12
+0.58/+1.15
-0.81/-0.86
-4.89/-6.10
+0.06/-0.29
be ∗
34.9/47.3
+2.58/+4.40
-0.11/-1.62
+1.67/+3.81
+2.61/+4.20
-1.44/-2.12
+1.44/+0.63
bho
40.09/48.86
-0.06/+1.27
-0.23/+1.82
-0.55/+0.80
-0.90/+0.55
-2.44/+1.15
-0.57/+0.57
Appendix
Table 8: Comparison of per-language probing UAS/UUAS between vanilla mBERT and bootstrapped models on test languages. The test languages include Akkadian (akk), Bambara (bm), Belarusian (be), Bhojpuri (bho), Breton (be), Buryat (bxr), Cantonese (yue), Erzya (myv), Faroese (fo), Karelian (krl), Kazakh (kk), Komi Permyak (koi), Komi Zyrian (kpv), Kurmanji (kmr), Livvi (olo), Marathi (Mr), Mbya Guarani (gun), Moksha (mdf), Nigerian Pidgin (pcm), Sanskrit (sa), Swiss German (gsw), Tagalog (tl), Tamil (ta), Telugu (te), Upper Sorbian (hsb), Warlpiri (wbp), Welsh (cy), and Yoruba (yo). Language codes with asterisks (*) indicate languages included in the mBERT pre-training data.
UDapter
UDapter-Bootstrapped mBERT
(Baseline)
Single-Layer (L12)
Multi-Layer (L12+L6)
MLM
SynAug
MLM
MLM
SynAug
MLM
+SynAug
+SynAug
akk
25.76/5.56
-0.22/+1.24 †
-9.29 † /-2.00 †
-1.94/+0.32
-1.35/+1.40 †
-7.13 † /-1.46 †
-0.54/+1.73 †
be ∗
84.31/79.54
+0.28/+0.52
+0.52/+0.66 †
+0.11/+0.46
+0.37/+0.70 †
+0.07/+0.48
+0.11/+0.42
bho
51.75/36.43
-0.10/+0.08
-0.49/-0.31
-0.84 † /-0.76 †
-0.14/+0.29
-0.14/-0.20
-0.47/-0.18
Appendix
Table 9: Comparison of zero-shot dependency parsing in UAS/LAS between vanilla and bootstrapped mBERT on 28 target languages. † indicates statistically significant results from McNemar’s test ( p<0.05 ). Language codes with asterisks (*) are languages included in the mBERT pre-training data.
Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood. We evaluate four parsers -- the Biaffine LSTM, Stack-Pointer Network, AfroXLMR-large, and RemBERT -- across ten typologically diverse languages, with a focus on low-resource African languages. We find that the Biaffine LSTM consistently outperforms transformer models in low-resource regimes, with transformers recovering their advantage as training data increases. The crossover falls within a resource range typical of treebanks for under-resourced languages. Morphological complexity (measured via MATTR) emerges as a significant secondary predictor of transformers' relative disadvantage after controlling for corpus size. These results indicate that the Biaffine LSTM may be better suited for syntactic tool development in low-resource regimes until sufficient annotated data is available to leverage the representational capacity of pre-trained transformers.
Dependency parsing consists of finding a tree representation for a sequence. Unsupervised dependency parsing aims to develop parsing methods without a gold standard during model training. In human languages, an unsupervised parser can be evaluated because some gold standard is usually available or can be created. For other species, a gold standard is unknown. Thus one may conclude that it is impossible to determine the accuracy of an unsupervised parser and, consequently, dependency parsing is unfeasible in other species. However, here we apply recent advances in network science to demonstrate that the proportion of correct edges retrieved by a parser must be high for the sequences of vocalizations or gestures that non-human primates produce due to the fast decay of the sequence length distribution. In contrast, human language sequences lack that property. Therefore, evaluation without a gold standard is feasible in non-human primates but a hard problem in humans.
Ramon Ferrer-i-Cancho, Catherine Hobaiter, Thore Bergman +1
Universitat Politècnica de Catalunya, Department of Computer Science · University of St Andrews, School of Psychology and Neuroscience · University of Michigan, Departments of Psychology and Ecology and Evolutionary Biology +1
Low-resource language varieties used by specific groups remain neglected in the development of Multilingual Language Models. A great deal of cross-lingual research focuses on inter-lingual language transfer which strives to align allied varieties and minimize differences between them. However, for low-resource varieties, linguistic dissimilarity is also an important cue allowing generalization to unseen varieties. Unlike prior approaches, we propose a two-stage Language Generalization framework that focuses on capturing variety-specific cues while also exploiting rich overlap offered by high-resource source variety. First, we propose TOPPing, a source-selection method specifically designed for low-resource varieties. Second, we suggest a lightweight VACAI-Bowl architecture that learns variety-specific attributes with one branch while a parallel branch captures variety-invariant attributes using adversarial training. We evaluate our framework on structural prediction tasks, which are among the few tasks available, as proxy for performance on other downstream tasks. Using VACAI-Bowl with TOPPing yields an average 54.62% improvement in the dependency parsing task, which serves as a proxy for performance on other downstream tasks across 10 low-resource varieties.
Jinju Kim, Haeji Jung, Youjeong Roh +2
Department of Electrical and Computer Engineering, Sungkyunkwan University · Language Technologies Institute, Carnegie Mellon University · Department of Computer Science, University of British Columbia +1