Pre-trained language models (PLMs) with encoder-based architectures have shown impressive capabilities in zero-shot cross-lingual transfer for various language understanding tasks. However, applying this technique to dependency parsing remains a significant challenge due to its syntactic nature. To boost model generalizability across linguistic typologies, we propose a cross-lingual unsupervised bootstrapping method to improve syntactic knowledge within the PLM. We show that our method achieves a significant improvement in zero-shot parsing performance in low-resource languages. Analysis of these bootstrapped models uncovers increased robustness in recognizing syntactic structures, evidenced by higher scores in parameter-free tree probing tests.
Figures & tables
Figure 1: Layer-wise CKA similarity between original and modified English sentences obtained from pre-trained mBERT representations. A higher score means more robustness for word swapping. We use the English EWT test treebank (UD v2.8) in this experiment.
Figure 2: Original sentence (above) and OV reordered sentence (below). Dashed arcs indicate relations related to OV reordering, while solid arcs indicate relations unrelated to OV reordering.
Model
All (n=28)
Unseen /
Seen from mBERT (grouped by pre-training data size in MB)
Avg
Avg
(0, 50]
(50, 100]
(200, )
Avg
UDapter (Reproduced)
51.29/36.65
42.01/27.52
67.49/51.09
65.73/49.75
75.99/62.63
70.88/55.92
UDapter-B-mBERT L12
-0.82 † /-0.48 †
-0.79 † /-0.33
-2.07 † /-1.52
-0.89 † /-1.83 †
+0.05/+0.21
-0.86 † /-0.82 †
with MLM
+0.16/ +0.28 †
-0.31 † /-0.18
+2.75/+3.49
+0.69/+0.35
+0.22/-0.00 †
+1.16/+1.24
with SynAug OV
-0.30 † /+0.16
-1.05 † /-0.38
+3.70 † /+3.63 †
-1.41 † /-1.59
+0.78/+0.98
+1.27 / +1.29 †
with SynAug OV + MLM
+0.16 † /+0.25 †
-0.03/ +0.06 †
+2.63 † /+2.15 †
-1.44/-0.94 †
+0.01/+0.31
+0.56/+0.65 †
Table 1: Zero-shot evaluation results on 28 target languages grouped by inclusion in mBERT pre-training data and the amount of pre-training data (low-resource, middle-resource, and high-resource languages). The scores reported are the average UAS/LAS change compared to UDapter. UDapter-B-mBERT denotes bootstrapped mBERT with two settings: L12 (last layer only) and L12+L6 (last and sixth layers). † indicates a statistically significant result from the McNemar’s test ( p<0.05 ).
Figure 3: Zero-shot LAS parsing performance grouped by the target languages’ canonical verb and object ordering. Results are from two UDapter models: Bootstrapped-mBERT L12 with MLM and VO augmentation, and Bootstrapped-mBERT L12+L6 with MLM loss. * denotes languages included in the mBERT pre-training data.
Treebank
mBERT
Bootstrapped mBERT
(test set)
L12
L12+L6
Unseen
50.85/55.76
-0.25/+0.95
+0.12/+1.05
Seen
53.74/59.57
+0.88/+1.85
+1.08/+1.88
(0, 50]
53.76/58.45
+0.21/+0.92
+1.03/+1.34
(50, 100]
53.55/60.01
+0.53/+0.83
+0.46/+0.77
(200, )
53.81/60.20
+1.56/+3.05
+1.43/+2.83
Table 2: Comparison of average UAS/UUAS obtained from parameter-free tree probing between vanilla mBERT and bootstrapped models trained with MLM ( λ = 0.1) and no sentence augmentation. The scores shown are from the layer with the highest UAS.
Encoder
All
Unseen
Seen
XLM-100
+0.35/+0.34
+0.17/-0.07
+0.78/+1.26
(Wiki)
(+0.26/+0.09)
(+0.36/+0.11)
(+0.06/+0.04)
Distil-mBERT
+0.53 /+0.37
+1.40/+0.80
-0.30/-0.13
(Wiki, Distill )
(+1.07/+1.82)
(+0.80/+1.60)
(+1.65/+2.31)
XLMR
-0.03/-0.39
-0.11/-0.36
-0.04/-0.52
(CC-100)
(-0.90/-1.48)
(-0.27/-0.93)
(-2.05/-2.48)
Table 3: Comparison of average UAS/LAS change and probed UAS/UUAS scores (in parentheses). The results come from a bootstrapped model that uses the last layer representation, MLM auxiliary loss, and VO augmentation. The results are selected based on the highest average UAS for the seen languages.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Sentences with Noun-Adj reordering (a) and Adposition (b) reordering. Dashed arcs indicate relations related to reordering.
Model
All (n=28)
Unseen /
Seen from mBERT (grouped by pre-training data size in MB)
Avg
Avg
(0, 50]
(50, 100]
(200, )
Avg
UDapter (Reproduced)
51.29/36.65
42.01/27.52
67.49/51.09
65.73/49.75
75.99/62.63
70.88/55.92
UDapter-B-mBERT L12
-0.82 † /-0.48 †
-0.79 † /-0.33
-2.07 † /-1.52
-0.89 † /-1.83 †
+0.05/+0.21
-0.86 † /-0.82 †
MLM ( λ=0.1 )
+0.16/+0.28 †
-0.31 † /-0.18
+2.75/+3.49
+0.69/+0.35
+0.22/-0.00 †
+1.16/+1.24
MLM ( λ=0.01 )
-0.38/+0.05
-0.81 † /-0.14
+1.71 † /+2.03 †
+0.21/-1.13
-0.14 † /+0.08 †
+0.55 † /+0.46 †
MLM ( λ=0.001 )
-1.07 † /-0.93 †
-1.50 † /-0.81 †
-0.84 † /-2.22 †
+1.64/-1.03 †
-0.58 † /-0.51 †
-0.17 † /-1.20 †
Appendix
Table 5: Zero-shot evaluation result on 28 target languages grouped by inclusion in mBERT pre-training data and the amount of pre-training data. The scores reported are the average UAS/LAS change compared to UDapter. UDapter-B-mBERT denotes bootstrapped mBERT with two settings: L12 (last layer only) and L12+L6 (last and sixth layers). The λ denotes the weight assigned to the MLM auxiliary loss during the model bootstrapping phase. † indicates a statistically significant result from McNemar’s test ( p<0.05 ). The pre-training data sizes (in MB) for seen languages are as follows: Breton (be), Tagalog (tl), Yoruba (yo) (0-50 MB); Marathi (mr), Welsh (cy) (50-100 MB); Belarusian (be), Kazakh (kk), Tamil (ta), Telugu (te) (>200 MB). Data sizes are calculated from the Wikipedia dump plaintext (snapshot: 2018-02-25).
Model
All (n=28)
Unseen /
Seen from the encoder (grouped by pre-training data size in MB)
Avg
Avg
(0, 50]
(50, 100]
(200, )
Avg
UDapter-Distil-MBert
45.73/30.74
39.31/25.29
50.13/31.34
56.24/37.45
67.65/52.84
59.27/42.25
UDapter XLM-100
49.87/36.18
43.67/29.92
55.25/41.61
62.33/46.76
69.08/56.51
62.97/49.38
UDapter-Distil-MBert L6
+0.30 † /+0.20 †
+1.41 † /+0.82 †
-0.94 † /+0.05
-1.92 † /-2.51 †
-0.41/-0.45
-0.92/-0.67 †
MLM ( λ=0.01 )
+0.74 † /+0.49 †
+1.68 † /+0.62 †
-1.53 † /+0.81 †
+1.18/+0.33
-0.24/-0.02
-0.47 † /+0.38 †
SynAug
+0.58 † /+0.31
+1.40 † /+0.64 †
+0.55/+0.45 †
-0.18 † /-1.62 †
+0.56 † /+0.55 †
+0.40 † /+0.08 †
Appendix
Table 6: Zero-shot evaluation result on 28 target languages grouped by inclusion in mBERT pre-training data and the amount of pre-training data. The scores reported are the average UAS/LAS change compared to UDapter. UDapter-B-Distil-mBERT denotes bootstrapped mBERT. The λ denotes the weight assigned to the MLM auxiliary loss during the model bootstrapping phase. The row marked with "-" indicates that the model failed to converge. † indicates a statistically significant result from McNemar’s test ( p<0.05 ).
Model
All (n=28)
Unseen /
Seen from XLMR (grouped by pre-training data size in GB)
Avg
Avg
(0, 4]
(4, 10]
(10, )
Avg
UDapter-mMiniLM
51.50/36.85
40.52/26.53
56.86/35.80
75.94/61.09
80.01/72.22
68.01/51.62
UDapter-XLMR
53.96/40.20
41.03/27.80
66.71/47.05
78.57/64.03
81.58/73.84
73.73/58.10
UDapter-B-XLMR L12
-1.23/-0.86
-1.67/-0.98
-0.64/-0.57
-0.40/-0.61
-0.62/-0.45
-0.55/-0.56
MLM ( λ=0.01 )
+0.14/-0.27
+0.10/-0.38
-0.36/-0.58
+0.29/+0.30
+0.29/+0.26
-0.01/-0.11
SynAug
-0.44/-0.48
-0.18/-0.21
-0.69/-0.94
-0.48/-1.31
+0.48/+0.40
-0.40/-0.83
Appendix
Table 7: Zero-shot evaluation result on 28 target languages grouped by inclusion in mBERT pre-training data and the amount of pre-training data. The scores reported are the average UAS/LAS change compared to UDapter. UDapter-B-XLMR denotes bootstrapped XLMR. UDapter-B-mMiniLM denotes bootstrapped Multilingual MiniLM. The λ denotes the weight assigned to the MLM auxiliary loss during the model bootstrapping phase. † indicates a statistically significant result from McNemar’s test ( p<0.05 ).
UDapter
UDapter-Bootstrapped mBERT
(Baseline)
Single-Layer (L12)
Multi-Layer (L12+L6)
MLM
SynAug
MLM
MLM
SynAug
MLM
+SynAug
+SynAug
akk
42.38/49.11
-0.12/+0.40
-1.04/-0.12
+0.58/+1.15
-0.81/-0.86
-4.89/-6.10
+0.06/-0.29
be ∗
34.9/47.3
+2.58/+4.40
-0.11/-1.62
+1.67/+3.81
+2.61/+4.20
-1.44/-2.12
+1.44/+0.63
bho
40.09/48.86
-0.06/+1.27
-0.23/+1.82
-0.55/+0.80
-0.90/+0.55
-2.44/+1.15
-0.57/+0.57
Appendix
Table 8: Comparison of per-language probing UAS/UUAS between vanilla mBERT and bootstrapped models on test languages. The test languages include Akkadian (akk), Bambara (bm), Belarusian (be), Bhojpuri (bho), Breton (be), Buryat (bxr), Cantonese (yue), Erzya (myv), Faroese (fo), Karelian (krl), Kazakh (kk), Komi Permyak (koi), Komi Zyrian (kpv), Kurmanji (kmr), Livvi (olo), Marathi (Mr), Mbya Guarani (gun), Moksha (mdf), Nigerian Pidgin (pcm), Sanskrit (sa), Swiss German (gsw), Tagalog (tl), Tamil (ta), Telugu (te), Upper Sorbian (hsb), Warlpiri (wbp), Welsh (cy), and Yoruba (yo). Language codes with asterisks (*) indicate languages included in the mBERT pre-training data.
UDapter
UDapter-Bootstrapped mBERT
(Baseline)
Single-Layer (L12)
Multi-Layer (L12+L6)
MLM
SynAug
MLM
MLM
SynAug
MLM
+SynAug
+SynAug
akk
25.76/5.56
-0.22/+1.24 †
-9.29 † /-2.00 †
-1.94/+0.32
-1.35/+1.40 †
-7.13 † /-1.46 †
-0.54/+1.73 †
be ∗
84.31/79.54
+0.28/+0.52
+0.52/+0.66 †
+0.11/+0.46
+0.37/+0.70 †
+0.07/+0.48
+0.11/+0.42
bho
51.75/36.43
-0.10/+0.08
-0.49/-0.31
-0.84 † /-0.76 †
-0.14/+0.29
-0.14/-0.20
-0.47/-0.18
Appendix
Table 9: Comparison of zero-shot dependency parsing in UAS/LAS between vanilla and bootstrapped mBERT on 28 target languages. † indicates statistically significant results from McNemar’s test ( p<0.05 ). Language codes with asterisks (*) are languages included in the mBERT pre-training data.
Universitat Politècnica de Catalunya, Department of Computer Science · University of St Andrews, School of Psychology and Neuroscience · University of Michigan, Departments of Psychology and Ecology and Evolutionary Biology +1
Department of Electrical and Computer Engineering, Sungkyunkwan University · Language Technologies Institute, Carnegie Mellon University · Department of Computer Science, University of British Columbia +1