End-to-end speech parsing, a task recently proposed, consists in predicting both the transcription and the syntactic tree for a spoken utterance. Existing architectures for speech parsing often utilise intermediate neural networks. In this work, we examine the effectiveness of intermediate neural networks (NN) for parsing, and, specifically, what role do they play. We introduce a simpler end-to-end architecture for speech parsing, where we remove these intermediate NN units, reducing the parameters by 12%, while achieving comparable or better performance than prior method on both automatic speech recognition (ASR) and parsing. We demonstrate that intermediate NN units help reduce the representational gap when the pre-trained encoder is frozen. We do a comprehensive evaluation of speech parsing on French, and medium-low resource languages Slovenian and Naija. We further investigate the impact of the training data size and intermediate layers of the pretrained speech encoder on speech parsing.
Figures & tables
Figure 1: Dependency parsing tree from the NaijaNSC treebank. en : Ala-… Alaska Pepper was shocked. The two lines under the wordform line are the sequence of POS and lemma tags respectively.
Dataset
Section
Duration
Trees
Speakers
Train
6.78h
7,268
67
Naija
Dev
0.87h
990
10
Test
0.91h
972
11
Train
4.88h
3,474
464
Slovenian
Dev
0.67h
400
80
Test
0.99h
358
36
Table 1: Data statistics for the Naija-NSC treebank, Slovenian SST, ParisStories, and Orféo
Figure 2: Comparison of the original architecture ( Pupier et al., 2024 ) and our proposed architecture.
Figure 3: Text parsing with frozen text encoder - with and without intermediate NN unit.
Table 2: Evaluation on test dataset with settings described in Section 5.5 . Parameter counts: w Wav2vec, g Speech_Large_fr_114K , f feedforward + LSTM, p parsing module, r Roberta module. W2T refers to Wav2tree. Best scores across end-to-end speech parsing highlighted in bold. Underlined refers to statistical significant results
Text parsing architecture
Intermediate NN
LAS
LAS
LAS
# of parameters
(LSTM)
(Naija)
(Slovenian)
(ParisStories)
( fine-tuned)
Joint fine-tuning
no
83.9
71.3
72.0
559M r +5M p
Joint fine-tuning + LSTM
yes
81.7
66.1
69.8
559M r +5M p +16M i
Frozen pre-trained text model
no
51.5
40.9
44.8
5M r
Frozen pre-trained text model + LSTM
yes
74.6
55.4
62.9
5M p +16M i
Table 3: Comparative analysis for RQ1 - LAS on the test dataset with and without intermediate NN unit for all 3 treebanks with experiments described in Section 5.1.2 . Best scores across sections highlighted in bold .
Similarity between
CKA
CKA
CKA
fine-tuned encoder &
Naija
Slovenian
French
Frozen encoder without
17%
21%
23%
intermediate NN unit
Frozen encoder with
44%
32%
34%
intermediate NN unit
Table 4: Similarity between representations from fine-tuned and frozen text encoder described in Section 6.1
Figure 4: Scaling training data across 2 setups - End-to-End (Baseline) and Frozen ASR - as described in Section 5.2
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Architecture
Baseline Wav2tree ( Pupier et al., 2024 )
Simplified Wav2tree (ours)
Epoch
80
80
Batch size
4
4
Tuning parameters
Learning rate (ASR)
0.00005
0.00005
Learning rate (Remaining)
0.0003
0.0003
Optimizer
AdamW
AdamW
Appendix
Table 5: Hyperparameters across Wav2tree architecture ( Pupier et al., 2024 ) and Simplified Wav2tree (ours).
Figure 5: Predicted word boundary and the corresponding uncollapsed sequence of the predicted transcription (in French) "J AI MON COLLÈGUE" ( en: I have my colleague). Symbols ϵ and _ represents respectively the blank and space labels. Indices {1,2,3,4} refer to the word position in the output sequence.
Index
Punctuation presence
WER ↓
UPOS ↑
UAS ↑
LAS ↑
1
Train & Eval data
34.4
76.5
62.8
57.6
2
Training data
32
78.5
68.8
61.8
3
None
31.8
79.1
69.1
61.8
Appendix
Table 6: Simplified W2T results across different variations of Naija dataset. Best scores highlighted in bold .
Figure 6: Gold and predicted transcription and trees for a Naija utterance (sent_id: WAZL_15_MC_Abi-MG__180). The predicted annotation has all UAS, UPOS and LAS as 72.7. The scores are low because not all markups are predicted by ASR - ‘(’, ‘)’ and ‘ < ’(marked as ‘<’) are not predicted by ASR.
English Vocab
Number of words
WER
1
8,991
24.5%
0
1,269
52.5%
Appendix
Table 7: WER for tokens in Naija-NSC treebank.
# of markups
Trees
Avg markups in ground_truth
Avg markups in prediction
LAS with markups
LAS without markups
LAS decrease
[0, 2]
423
1.5
1.9
66.0
66.0
0.0
[3, 4]
232
3.4
3.5
63.0
64.5
-1.4
[5, 6]
137
5.4
4.8
57.1
61.1
-3.9
[7, 12]
139
8.7
7.1
54.2
60.3
-6.1
>=13
41
16.5
12.6
44.0
54.0
-10.0
Appendix
Table 8: Parsing results comparison with and without syntax-prosody segmentation annotations.
Index
Hypothesis
p value
p value
p value
Naija
Slovenian
ParisStories
1
Null
21.6%
0.2%
3.6%
2
Alternative
-
0.1%
1.8%
Appendix
Table 9: Statistical significance testing between Simplified and Baseline architecture
Index
Hypothesis
p value
p value
p value
Naija
Slovenian
ParisStories
1
Null
21.6%
0.2%
3.6%
2
Alternative
-
0.1%
1.8%
Appendix
Table 10: Statistical significance testing between Simplified and Baseline architecture